Learned Initializations for Optimizing Coordinate-Based Neural Representations

Matthew Tancik, Ben Mildenhall, Terrance Wang, Divi Schmidt, Pratul P. Srinivasan, Jonathan T. Barron, Ren Ng

Introduction

Recent work has demonstrated the potential of representing complex low-dimensional signals using deep fully-connected neural networks (typically referred to as multilayer perceptrons, or MLPs). A coordinate-based neural representation fθf_{\theta} for a given signal is an MLP (with weights θ\theta) that is optimized to map from an input coordinate x\mathbf{x} to the signal’s value at that coordinate. For example, fθf_{\theta} could map from 2D pixel coordinates to RGB color values to encode an image. Unlike a signal stored as a discretely sampled array of values, a coordinate-based neural representation is continuous and is not constrained to have a fixed spatial resolution. This fact has recently been exploited to design representations for 3D shapes (which typically occupy a small 2D subset of 3D space) that do not require cubic storage complexity, in contrast to 3D voxel grids .

However, one limitation of these neural representations is that computing network weights θ\theta that reproduce a given signal typically requires solving an optimization problem by running many steps of gradient descent. This can take between seconds (when encoding a small image) and hours (when solving an inverse problem to recover a high resolution radiance field, as in NeRF ). Common approaches to address this issue include concatenating a latent vector to the input coordinate and supervising a single neural network to represent an entire class of signals , or training a hypernetwork to map from signal observations (or a latent code) to MLP weights . However, each of these strategies is restricted to representing only signals within its learned latent space, potentially limiting its ability to express previously unseen target signals.

Recent work has shown that optimization-based meta-learning can dramatically reduce the number of gradient descent steps required to optimize a neural representation to encode a new signal in the case of signed distance fields of 2D and 3D shapes. In this work, we propose learning the weight initialization for neural representations across a wide variety of underlying signal types, such as images, volumetric data, and 3D scenes. We show that compared to a standard random initialization, using fixed, learned values for the initial network weights acts as a strong prior that enables both faster convergence during optimization and better generalization when only partial observations of the target signal are available. In the context of using neural representations for 3D reconstruction from images, a learned initialization specialized to a particular ShapeNet class allows the network to recover 3D shape from a single image over the course of optimization, whereas a standard randomly initialized network fails unless provided with multiple input views. Given a meta-training set consisting of observations of different signals sampled from a fixed underlying class, our setup applies an optimization-based meta-learning algorithm (MAML or Reptile ) in order to produce initial weights better suited for representing that specific signal class (e.g., face images from CelebA or 3D chairs from ShapeNet ).

The biggest advantage of our approach is its simplicity. Given an existing framework for test-time optimization of a neural representation, implementing an outer loop with MAML or Reptile update steps only requires a few extra lines of code and a dataset of training examples. Once the meta-learning phase is complete, the learned initial weights can be stored and later reloaded in place of a standard network initialization whenever a new signal needs to be encoded. This minor implementation change can significantly alter the behavior of the network during optimization.

Related Work

Neural representations have recently risen to prominence as compact representations for 3D shapes. These methods represent shapes as implicit surfaces defined as a level set of an MLP network and enable full object reconstruction from incomplete 3D point cloud data or depth scans . Later work combined this idea with various formulations of differentiable rendering to recover neural representations of 3D shape using only 2D image observations .

Coordinate-based neural networks have also been used to represent other low-dimensional signals, such as 2D images, where such networks (when trained via genetic algorithms) have been referred to as compositional pattern–producing networks . Recent works have shown that standard ReLU MLPs fail to adequately represent fine details in these complex low-dimensional signals due to a spectral bias and address this issue by either replacing the ReLU activations with sine functions or by lifting the input coordinates into a Fourier feature space . Our work makes use of these observations and presents a technique that enables a coordinate-based MLP to learn from the process of fitting many signals within a category so that it can quickly optimize to fit any new signal using fewer steps and fewer observations.

Meta-learning

Meta-learning typically addresses the problem of few-shot learning, where some examples of a given task (including training and test data) are used to learn an algorithm that achieves better performance on new, previously unseen instances of the same task. A prototypical example from computer vision is few-shot image classification, where a network must learn to differentiate between new classes at test time based on only a small number of labeled instances of each class.

Most relevant to this work are optimization-based meta-learning algorithms such as Model-Agnostic Meta Learning (MAML) and Reptile , as well as various extensions . Given a network architecture for performing a task, these methods use an outer loop of gradient-based learning to find a weight initialization that allows the network to more efficiently optimize for new instances of the underlying task at test time. These methods assume the use of a standard gradient-based optimization method such as stochastic gradient descent or Adam at test time, making them easy to layer on top of existing implementations, as opposed to more complex methods such as Ravi et al. , which trains a “meta-learner” LSTM network to perform gradient updates for the underlying task. An exhaustive review of meta-learning algorithms is provided in the survey paper by Hospedales et al. .

MetaSDF specifically applies this idea of learning a weight initialization to the task of fitting neural representations to represent signed distance fields, and shows that this strategy achieves much more rapid convergence than standard approaches such as DeepSDF . Our work applies meta-learning to neural representations for a wider variety of underlying signal types and further explores the power of using initial weight settings as a prior.

Overview

If direct pointwise observations {(xi,T(xi)}i\{(\mathbf{x}_{i},T(\mathbf{x}_{i})\}_{i} of the signal TT are available, fθf_{\theta} can be supervised by gradient descent using a simple L2 loss:

Let θ0\theta_{0} denote the initial network weights before any gradient steps are taken, and let θi\theta_{i} denote the weights after ii steps of optimization. Basic gradient descent applies the rule:

with a learning rate parameter α\alpha, whereas more sophisticated optimizers such as Adam keep track of gradient moments over time to redirect the optimization trajectory. Given a fixed budget of mm optimization steps, different initial weight values θ0\theta_{0} will result in different final weights θm\theta_{m} and signal approximation error L(θm)L(\theta_{m}). When emphasizing the functional dependence of θm\theta_{m} on the initial weights and a particular signal, we will write θm(θ0,T)\theta_{m}(\theta_{0},T).

It is often the case that only indirect observations of TT are available, taken through some forward measurement model M(T,p)M(T,\mathbf{p}). For example, if TT is a 3D object, M(T,p)M(T,\mathbf{p}) could be a 2D image captured of the object from camera pose p\mathbf{p}. In this case, recovering a neural representation for TT from observations {pi,M(T,pi)}i\{\mathbf{p}_{i},M(T,\mathbf{p}_{i})\}_{i} requires solving an inverse problem by taking gradient steps on a loss that incorporates the forward model MM:

If MM discards too much information about TT or the set of provided observations is too small, the resulting network fθf_{\theta} may not match TT closely. For example, accurately recovering a 3D object from a single 2D view may not be possible without strong a priori knowledge of the object’s shape.

We assume that we are given a dataset of observations of signals TT from a particular distribution T\mathcal{T} (e.g., 2D face images or 3D chairs) and our goal is to find initial weights θ0∗\theta_{0}^{*} that will result in the lowest possible final loss L(θm)L(\theta_{m}) when optimizing a network fθf_{\theta} to represent a new, previously unseen signal from the same distribution:

This problem of trying to learn the initial weights of a network to serve as a good starting point for gradient descent across a distribution of tasks is addressed by a variety of optimization-based meta-learning algorithms, such as MAML and Reptile .

Given a task TT, calculating the weight values θm(θ0,T)\theta_{m}(\theta_{0},T) requires taking mm optimization steps, which are collectively referred to as the inner loop. MAML wraps an outer loop of meta-learning around this inner loop in order to learn the initial weights θ0\theta_{0}. Each outer loop samples a signal TjT_{j} from T\mathcal{T} and applies the update rule:

with meta-learning step size β\beta. This update rule applies gradient descent to the loss on the weights θm(θ0j,Tj)\theta_{m}(\theta_{0}^{j},T_{j}) resulting from the inner loop optimization.

Reptile [26]

Reptile uses the same meta-learning setup as MAML but applies a simpler update rule that does not require calculating second-order gradients:

This rule moves the previous weight initialization θ0j\theta_{0}^{j} in the direction of the task-optimized weights θm(θ0j,Tj)\theta_{m}(\theta_{0}^{j},T_{j}).

2 Experimental setup

The meta-learning algorithms described previously are conceptually simple, requiring no changes to the architecture or optimization procedure of a coordinate-based neural representation when given a new signal to encode at “test time” (after meta-learning is complete). These algorithms produce only a set of initial network weights θ0∗\theta_{0}^{*} that are then used as a starting point for gradient descent. Test-time optimization on new signals is not limited to the same number of steps mm as were used in the inner loop during meta-learning; indeed, at test time we often observe benefits from optimizing for significantly more iterations than were used during the inner loop of the meta-learning algorithm.

MAML is typically able to produce a better initialization than Reptile given a fixed number of inner loop steps mm, but Reptile can be unrolled for more inner loop steps because it is less memory-intensive than MAML. For some tasks, MAML’s limited number of inner loop steps means that it can only observe a small percentage of the observations of a target signal. In these cases, we use Reptile to maximize the number of different observations seen over the course of the inner loop. Experimentally we find it beneficial to unroll more steps for more complex tasks.

Each of our experiments involves two phases:

Meta-learning, where we use MAML or Reptile in combination with a training dataset of example tasks (observations of different signal instances) to optimize initial network weights for that class of signals, and

Test-time optimization, where we use standard gradient-based optimization to fit the weights of a network to observations of a previously unseen signal from the same class.

We aim to answer the following question: how do different initial network weight settings influence the ability of a neural representation to fit to a new signal during test-time optimization?

Results

We present results on 2D image regression, 2D computed tomography (CT) reconstruction, 3D object reconstruction, and 3D scene reconstruction. For each task, we demonstrate the benefits of using meta-learned initial weights optimized to reconstruct a specific class of signals.

For 2D image regression, a meta-learned weight initialization leads to faster convergence and better performance during test-time optimization. For CT reconstruction, it allows for better reconstruction quality from fewer supervision views during test-time optimization. For 3D shape reconstruction from images, it allows for faster convergence at test time and makes single view reconstruction possible. For Phototourism landmark reconstruction, it can be optimized at test time to transfer the appearance of a single input image onto the whole landmark, which can then be rendered from novel camera views.

Here we provide the basic setup for each task. Please see the supplement for full implementation details.

A prototypical example of a coordinate-based neural representation is an MLP optimized to represent a 2D image by taking in 2D pixel coordinates and outputting RGB color values. We consider four different distributions T\mathcal{T}: images of faces (CelebA ), natural images (Imagenette ), images of text (Text), and 2D signed distance fields of simple curves (SDF). Each category contains around ten thousand examples. Given a sampled image T∼TT\sim\mathcal{T}, we provide all 178×178178\times 178 pixels as observations for optimizing the network weights θ\theta in the inner loop. Since this task is not memory constrained, we use MAML to meta-learn the weights over 2 unrolled gradient steps (separately for each category T\mathcal{T}). In each of these inner loop steps, the entire image is reconstructed and used to calculate the loss. For the MLP fθf_{\theta}, we use 5 layers with 256 channels each and sine function nonlinearities, as in SIREN .

CT reconstruction

Computed tomography (CT) is a widely used medical imaging technique that captures projective measurements of the volumetric density of a target object. Tancik et al. use a coordinate-based neural representation to reconstruct a 2D signal from 1D integral projections; the underlying MLP fθf_{\theta} takes in a 2D coordinate and outputs a scalar volume density at that location. Here T\mathcal{T} is a dataset of 2048 randomly generated 256×256256\times 256 pixel Shepp-Logan phantoms , where we provide 2D integral projections of a bundle of 256 parallel rays from a random angle as the measurement for each sampled signal TT during meta-learning. We use Reptile to meta-learn the initial weights over 12 unrolled gradient steps. We found this to outperform MAML, which was limited to 3 unrolled steps due to memory constraints. For the MLP fθf_{\theta}, we use 5 layers with 256 channels each and ReLU nonlinearities, and we apply random Fourier features to the input coordinates .

View synthesis for ShapeNet [3] objects

The goal of view synthesis is to generate a novel view of a scene from a set of reference images. Recently, neural radiance fields (NeRF) proposed a method to accomplish this task by using a neural representation that predicts a color and density for any input 3D location and 2D viewing direction within the scene, along with a differentiable volumetric rendering model to generate new views from that representation. This network is optimized to minimize the residual of re-rendering each of the input reference images from their respective camera poses. In our view synthesis experiments, we use a simplified NeRF model (simple-NeRF) that maintains the same image supervision and volume rendering context. Unlike the original NeRF model, we do not feed in the viewing direction and we use a single model instead of the two “coarse” and “fine” models used by NeRF.

For view synthesis on objects from the ShapeNet dataset, we consider three categories T\mathcal{T}: Chairs, Cars, and Lamps. We provide 25 128×128128\times 128 pixel reference images during meta-learning for each 3D object TT. The reference viewpoints are randomly distributed on a sphere and are oriented towards the target object, and each object is oriented in the canonical coordinate frame. The scenes are lit by a randomly selected environment map and rendered using ray tracing. We use Reptile to meta-learn the initial weights (for each shape category) over 32 unrolled gradient steps. For the MLP fθf_{\theta}, we use 6 layers with 256 channels each and ReLU nonlinearities, and apply a positional encoding to the input coordinates .

View synthesis for Phototourism [16] scenes

This dataset consists of thousands of posed tourist photographs of famous landmarks. Our objective is to use these images to create an underlying representation that can be explored and rendered from novel viewpoints with varying lighting conditions. The primary challenge is the diversity of the capture conditions: the photos are taken with different lighting conditions, camera hardware, camera viewpoint, and varying transient objects like people and cars. Each underlying dataset T\mathcal{T} for meta-learning θ0∗\theta_{0}^{*} consists of images of a single landmark (Trevi, Sacre Couer, or Brandenburg); the category is the overall 3D structure of the landmark itself, and the signal is its particular appearance (resulting from the time of day, lighting, weather conditions, etc) within a single photo. If a standard NeRF model is trained directly on this data, it learns a blurry representation of the scene that roughly corresponds to the mean of the environmental conditions. NeRF in the Wild explores these shortcomings and proposes extensive architectural modifications to account for the variations. We find that these shortcomings can be addressed to some degree solely with a better initialization and no architectural changes.

We apply meta-learning to the same simple-NeRF model from the ShapeNet experiment. The meta-training dataset for each landmark consists of thousands of images with varying resolution and intrinsic/extrinsic camera parameters. We use Reptile to meta-learn the initial weights (for each landmark) over 64 unrolled gradient steps. At test time, we optimize the simple-NeRF (starting from the initial weights θ0∗\theta_{0}^{*} for that landmark) to reproduce the appearance of a new image, and then render that simple-NeRF from other viewpoints. For the underlying MLP fθf_{\theta}, we use 6 layers with 256 channels each and ReLU nonlinearities, and apply positional encoding to the input coordinates .

2 Baselines

As well as a Standard randomly initialized network (Glorot et al. ), we compare to various other initialization schemes in several of our experimental settings:

Mean: we optimize a network from scratch such that its output matches the mean signal ET∼T[T]E_{T\sim\mathcal{T}}[T] from the current class T\mathcal{T}.

Matched: we optimize a network from scratch such that its output matches the output of a network using the meta-learned initialization for the current class T\mathcal{T}.

Shuffled: we randomly permute the weights (within each network layer) of the meta-learned initialization θ0∗\theta_{0}^{*} for the current class T\mathcal{T}.

Both the Mean and Matched baselines demonstrate the difference between having a good initialization in signal space versus weight space—despite Mean and Matched being initialized so that the loss against a randomly sampled signal will be low, they are a worse starting point for gradient descent than the actual meta-learned initial weights. The Shuffled baseline demonstrates that matching the statistical distribution of the meta-learned initial weights is not sufficient for better convergence or generalization. We find that using the Adam optimizer performs best for all of the baseline initializations, but that standard stochastic gradient descent works best for the meta-learned initializations (we choose the best optimizer and hyperparameters for each task and initialization using a held-out validation set, see supplement for details).

3 Faster convergence

In Figure 2, we visualize the network output for a variety of initial weight settings, showing the output images after 0, 1, and 2 gradient steps of test-time optimization. The meta-learned initial weights are optimized to represent face images (CelebA ). When using the learned initial weights θ0∗\theta_{0}^{*} (Meta), the target image is already clearly visible after the very first step. In contrast, the baseline initialization methods take an order of magnitude more iterations to represent the target image to the same accuracy (see Table 1). The Mean, Matched, and Shuffled baselines perform better than the completely random Standard initialization, but still take over ten times as many iterations to reach the same quality as the meta-initialized network can after 2 steps. In particular, this demonstrates that neither matching the image space output nor the statistical distribution of the meta-learned weights is sufficient for achieving a similar speedup.

View synthesis for ShapeNet [3] objects

In Figure 5, we plot the image reconstruction accuracy for a held-out test set of objects from the Chair category. During test-time optimization, 25 views are observed. We find that starting from the optimized weights θ0∗\theta_{0}^{*} allows the network to recover the chair more quickly compared to the Standard weight initialization. We note that after many steps, both methods end up at a similar quality.

4 Generalizing from partial observations

We perform meta-learning experiments across multiple datasets to determine the extent that the optimized weight initialization acts as a class-specific prior. We compare initializations trained on four different image datasets (CelebA, Imagenette, Text, and SDF). Table 2 presents a confusion matrix demonstrating that optimizing the network initialization does in fact induce a dataset-dependent prior, with each learned initialization generalizing best to the same dataset distribution it was trained on.

CT reconstruction from sparse views

We report the reconstruction quality over a test set of phantoms given varying numbers of views at test time in Table 3 and visualize one test example in Figure 3. We observe poor reconstructions from the Standard initialization when few views are provided. The meta-learned initializations are consistently able to match the PSNR of Standard with half as many views. The Mean initialization is generated by training a network to reconstruct the mean of the training phantoms. It is better able to preserve the structure of the phantom compared to Standard but still performs worse than the meta-learned initializations.

Single image view synthesis for ShapeNet [3]

A simple-NeRF model with a Standard random initialization relies on multi-view consistency to reconstruct the appearance of a 3D object. With only a single view, this naïve model is unable to recover any meaningful shape. We find that a learned initialization “bakes in” a class-specific shape prior that enables the recovery of 3D geometry (Figure 4, Table 4). We can meta-learn an effective weight initialization for single-view reconstruction by optimizing over a dataset with 25 training views of each object (MV Meta). We find that this prior persists even if the meta-training dataset only contains a single reference image per scene (SV Meta), meaning that the meta-learning phase has no access to multiview information for any particular object.

View synthesis with appearance transfer for Phototourism [16]

As described in §4.1, these images have different camera poses and visual appearance (lighting, sky, etc.) as they are taken by tourists at different times. Our goal at test time is to explore the landmark from varying camera viewpoints but rendered with the same appearance as in a target photograph. In every step of the meta-learning outer loop, we supervise the simple-NeRF model to match the appearance of a random photo of the landmark (with varying pose and appearance). We find that performing test-time optimization using a single new photograph allows us to render convincing unobserved viewpoints of the scene with the same environmental conditions.

In Figure 6, we show results for two landmarks. We test-time optimize the meta-learned weights for five target images (shown on the left side of the grid), taking 150 gradient steps for each image. We then render each of the resulting simple-NeRF networks from the five different viewpoints (shown in the row above the grid). The result is an image from the camera position of the corresponding top row image and matching the appearance of the left column image.

Quantitative evaluation on the Phototourism dataset is difficult as multiple views with the same environmental conditions do not exist. To overcome this, for Table 5 we optimize and evaluate on the same image, by optimizing to match the appearance of the left half of the image and subsequently evaluating metrics on the right half. For comparison, we train a simple-NeRF model with a standard random initialization from scratch on each landmark, then test-time optimize it to match the left half of each new view before evaluating it on the right half. This is algorithmically equivalent to Reptile with one inner optimization step. We find that unrolling Reptile for 64 inner steps performs better, producing significantly clearer renderings of the landmark.

Conclusion

Our results show that simply modifying a coordinate-based neural representation’s initial weight values can guide the network along a significantly better optimization trajectory, without changing the underlying architecture or test-time optimization procedure. These meta-learned initial weights can result in faster convergence or act as a strong prior for representing signals from a given distribution. This partially ameliorates a major shortcoming of neural representations (separately optimizing a network for each new signal) without limiting their representational power.

There are many additional directions to explore, such as applying more sophisticated meta-learning algorithms or more precisely characterizing the geometry of weight space for these networks. One limitation of our current approach is that it requires a sizable dataset of example signals from a target distribution in order to derive beneficial initial weights. Another shortcoming is that our method still requires some amount of test-time optimization.

As the number of use cases for neural representations continues to rapidly expand, we believe this work takes an important step toward understanding the importance of their initial weights and optimization behavior.

Acknowledgements

MT is funded by an NSF fellowship and a Berkeley DeepDrive grant. BM is funded by Google through the BAIR Commons Program. Google University Relations provided a generous donation of GCP compute credits. We thank Ruichao Ren from NVIDIA for donating GPU hardware.

References

Appendix A Implementation details

We found that modifying the weight initialization for these coordinate-based networks drastically changed their convergence behavior during test-time optimization. As a result, we tuned the optimization method and hyperparameters for each part of each experiment (using held-out validation sets) in order to provide the fairest possible comparison and to not bias the results against the non-meta-learned initializations. For example, we often found that SGD outperformed Adam when doing test-time optimization using meta-learned initializations, but that Adam was significantly better than SGD with a standard random initialization.

All experiments are implemented in JAX . Each experiment is trained on either a single NVIDIA V100, 2080 Ti, or 3080 Ti. In all cases where the Adam optimizer is used, we keep the standard parameter choices for β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999, ϵ=10−8\epsilon=10^{-8}.

For this task we use a SIREN architecture (ω0=200\omega_{0}=200) with 5 layers of 256 channels each. For the randomly initialized Standard baseline, we use the specific initialization procedure as proposed in the SIREN paper.

MAML is trained for 150K iterations. Each iteration has an outer batch size of 3 target images. The inner batch contains all pixels of the target image. The outer loop uses the Adam optimizer with learning rate of 10−510^{-5}. The inner loop performs two steps of gradient descent with a learning rate of 10−210^{-2}.

We additionally meta-learn another initialization using Reptile . We use the same learning rates as in MAML but with an outer batch size of 10 target images. We report the Reptile reconstruction accuracy in Table 6. We note that Reptile also outperforms the non-meta-learned weights.

During test-time optimization, we use gradient descent with learning rate of 10−210^{-2} when starting from the MAML initial weights. For the baseline methods (Standard, Mean, Matched, Shuffled) we used Adam with learning rate of 10−410^{-4}, which performed significantly better than than gradient descent.

A.2 CT reconstruction

For this task we use an MLP with 5 layers of 256 channels each. The network uses a ReLU activation after each layer with the exception of the last layer, which has a sigmoid activation. Prior to inputting the coordinates into the network, we encode them using random Fourier features sampled from a normal distribution with σ=30\sigma=30, as was done in Tancik et al. .

Reptile is trained for 100K iterations. Each iteration has an outer batch size of 1. The inner batch contains 20 CT projections, each with 256 measurements, taken from a randomly sampled direction. The outer loop uses the Adam optimizer with learning rate of 5×10−55\times 10^{-5}. The inner loop performs 12 inner loop steps of gradient descent with a learning rate of 10110^{1}.

We perform test-time optimization experiments with different numbers of supervision views to compare reconstruction quality. We found that the models are more prone to overfitting when fewer views are provided. We tune the learning rate and number of gradient steps for each initialization method according to a held-out set of 16 validation images. We report all of the test-time optimization hyper-parameters in Table 7.

A.3 ShapeNet [3] view synthesis

We use a simplified NeRF model for our view synthesis tasks. This model uses a single network rather than two networks (coarse and fine), and we do not provide view directions as input. The network is an MLP with 6 layers, each with 256 channels and ReLU activations. As in NeRF , we apply a positional encoding to each input coordinate with the form

with N=20N=20 encodings and log-max frequency f=8f=8. We accumulate 128 samples per ray for rendering.

Reptile is trained for 100K iterations with an outer batch size of 1. The inner loop step optimizes over a batch of 128 rays. We perform 32 inner loop steps for every outer loop step. The outer loop uses the Adam optimizer with learning rate 5×10−45\times 10^{-4} for the Chairs scenes and 5×10−55\times 10^{-5} for the Lamps and Cars scenes.

The test-time optimization parameters vary depending on the scene and the number of views available during meta-learning. Each experiment uses an inner batch of 64 rays. The Shuffled and Matched initializations are computed based on the MV Meta weights. For the 25 view chair reconstruction, we use stochastic gradient descent with a learning rate of 10−110^{-1} for the Reptile initialization; for the standard initialization, we use Adam with a learning rate of 10−410^{-4}. The test-time hyper parameters for the single view experiments are listed in Table 8.

A.4 Phototourism [16] view synthesis

We use the same architecture as described in §A.3. Reptile is trained for 150K iterations with an outer batch size of 1. The inner loop step optimizes over a batch of 64 rays, with 128 volume rendering samples per ray. The outer loop uses the Adam optimizer with a learning rate of 5−45^{-4}. We train with 64 inner loop steps using gradient descent with a learning rate of 1010. We compare to Basic NeRF which has the same setup, but only one inner step. For Basic NeRF we train Trevi for 60K iterations, Brandenburg for 100K iterations, and Sacre Coeur for 200K iterations. To transfer the appearance of a new photo during test-time optimization, we take 150 gradient steps with a learning rate of 1010.

Appendix B Weight space interpolation

We find that linearly interpolating between networks in weight space produces meaningful outputs when using meta-learned weights. Figure 7 shows interpolation between networks trained to represent images, and Figure 8 shows interpolations between networks that are trained to reconstruct a Phototourism landmark.