Unconstrained Scene Generation with Locally Conditioned Radiance Fields
Terrance DeVries, Miguel Angel Bautista, Nitish Srivastava, Graham W. Taylor, Joshua M. Susskind
Introduction
Spatial understanding entails the ability to infer the geometry and appearance of a scene when observed from any viewpoint or orientation given sparse or incomplete observations. Although a wide range of geometry and learning-based approaches for 3D view synthesis can interpolate between observed views of a scene, they cannot extrapolate to infer unobserved parts of the scene. The fundamental limitation of these models is their inability to learn a prior over scenes. As a result, learning-based approaches have limited performance in the extrapolation regime, whether it be inpainting disocclusions or synthesizing views beyond the boundaries of the given observations. For example, the popular NeRF approach represents a scene via its radiance field, enabling continuous view interpolation given a densely captured scene. However, since NeRF does not learn a scene prior, it cannot extrapolate views. On the other hand, conditional auto-encoder models for view synthesis are able to extrapolate views of simple objects from multiple viewpoints and orientations. Yet, they overfit to viewpoints seen during training (see for a detailed analysis). Moreover, conditional auto-encoders tend to favor point estimates (e.g. the distribution mean) and produce blurry renderings when extrapolating far from observed views as a result.
A learned prior for scenes may be used for unconditional or conditional inference. A compelling use case for unconditional inference is to generate realistic scenes and freely move through them in the absence of input observations, relying on the prior distribution over scenes (see Fig. 1 for examples of trajectories of a freely moving camera on scenes sampled from our model). Likewise, conditional inference lends itself to different types of problems. For instance, plausible scene completions may be sampled by inverting scene observations back to the learned scene prior . A generative model for scenes would be a practical tool for tackling a wide range of problems in machine learning and computer vision, including model-based reinforcement learning , SLAM , content creation , and adaptation for AR/VR or immersive 3D photography.
In this paper we introduce Generative Scene Networks (GSN), a generative model of scenes that allows view synthesis of a freely moving camera in an open environment. Our contributions are the following. We: (i) introduce the first generative model for unconstrained scene-level radiance fields; (ii) demonstrate that decomposing a latent code into a grid of locally conditioned radiance fields results in an expressive and robust scene representation, which outperforming strong baselines; (iii) infer observations from arbitrary cameras given a sparse set of observations by inverting GSN (i.e., fill in the scene); and (iv) show that GSN can be trained on multiple scenes to learn a rich scene prior, while rendered trajectories are smooth and consistent, maintaining scene coherence.
Related Work
An extensive body of literature has tackled view synthesis of simple synthetic objects against a homogeneous background . Relevant models typically rely on predicting 2D/3D flow fields to transform pixels in the input view to pixels in the predicted view. Moreover, in a continuous representation is introduced in the form of the weights of a parametric vector field, which is used to infer appearance of 3D points expressed in a world coordinate system. At inference time the model is optimized to find latent codes that are predictive of the weights. The above approaches rely on the existence of training data of the form input-transformation-output tuples, where at training time the model has access to two views (input and output) in a viewing sphere and the corresponding 3D transformation. Sidestepping this requirement showed that similar models can be trained in an adversarial setting without access to input-transformation-output tuples.
Recent approaches have been shown to encode a continuous radiance field in the parameters of a neural network by fitting it to a collection of high resolution views of realistic scenes. The encoded radiance field is rendered using the classical volume rendering equation and a reconstruction loss is used to optimize the model parameters. While successful at modelling the radiance field at high resolutions, these approaches require optimizing a new model for every scene (where optimization usually takes days on commodity hardware). Thus, the fundamental limitation of these works is that they are unable to represent multiple scenes within the same model. As a result, these models cannot learn a prior distribution over multiple scenes.
Some newer works have explore generative radiance field models in the adversarial setting . However, are restricted to modeling single objects where a camera is positioned on a viewing sphere oriented towards the object, and they use a 1D latent representation, which prevents them from efficiently scaling to full scenes (c.f. § 4 for experimental evidence). Concurrent to our work, Niemeyer et al. model a scene with multiple entities but report results on single-object datasets and simple synthetic scenes with 2-5 geometrical shapes in the CLEVR dataset . The object-level problem may be solved by learning a pixel-aligned representation for conditioning the radiance field . However, are designed as conditional auto-encoders as opposed to a generative model, thereby ignoring the fundamentally stochastic nature of the view synthesis problem. In a model is trained on long trajectories of nature scenarios. However, due to their iterative refinement approach the model has no persistent scene representation. While allowing for perpetual generation, the lack of a persistent scene representation limits its applicability for other downstream tasks that may require loop closure.
Approaches tackling view synthesis for a freely moving camera in a scene offer the capability to fully explore a scene . Irrespective of the output image resolution, this is strictly a more complex problem than a camera moving on a sphere oriented towards a single object . For scenes with freely moving cameras the model needs to learn to represent a full scene that consists of a multitude of objects rather than a single object. In this setup, it is often useful to equip models with a memory mechanism to aggregate the set of incoming observations. In particular, it has been useful to employ a dictionary-based memory to represent an environment where keys are camera poses and values are latent observation representations . Furthermore, the memory may be extended to a 2D feature map that represents at top-down view of the environment . At inference time, given a query viewpoint the model queries the memory and predicts the expected observation. This dependency on stored observations significantly limits these models’ ability to cope with unseen viewpoints. GSN can perform a similar task by matching observations to scenes in its latent space, but with the added benefit that the learned scene prior allows for extrapolation beyond the given observations. A summarized comparison of GSN with the most relevant related work is shown in Tab. 1.
Method
In GRAF and -GAN a 1D global latent code is used to condition an MLP which parameterizes a single radiance field. Although such 1D global latent code can be effective when representing individual object categories, such as cars or faces, it does not scale well to large, densely populated scenes. Simply increasing the capacity of a single radiance field network has diminishing returns with respect to render quality . It is more effective to distribute a scene among several smaller networks, such that each can specialize on representing a local region . Inspired by this insight we propose the addition of a global generator that learns to decompose large scenes into many smaller, more manageable pieces.
2 Locally Conditioned Radiance Field
3 Sampling Camera Poses
Unlike standard 2D generative models which have no explicit concept of viewpoint, radiance fields require a camera pose for rendering an image. Therefore, camera poses need to be sampled from pose distribution in addition to the latent code , which is challenging for the case of realistic scenes and a freely moving camera. GRAF and -GAN avoid this issue by training on datasets containing objects placed at the origin, where the camera is constrained to move on a viewing sphere around the object and oriented towards the origin.Camera poses T on a sphere oriented towards the origin are constrained to as opposed to free cameras in . For GSN, sampling camera poses becomes more challenging due to i) the absence of such constraints (i.e. need to sample ), ii) and the possibility of sampling invalid locations, such as inside walls or other solid objects that sporadically populate the scene.
To overcome the issue of sampling invalid locations we perform stochastic weighted sampling over a an empirical pose distribution composed by a set of candidate poses, where each pose is weighted by the occupancy (i.e., the value predicted by the model) at that location. When sampling the candidate poses, we query the generator at each candidate location to retrieve the corresponding occupancy. Then the occupancy values are normalized with a softmin to produce sampling weights for a multinomial distribution. This sampling strategy reduces the likelihood of sampling invalid camera locations while retaining stochasticity required for scene exploration and sample diversity.
4 Discriminator
Our model adopts the convolutional discriminator architecture from StyleGAN2 .The discriminator takes as input an image, concatenated with corresponding depth map normalized between $C$ to the discriminator and to enforce a reconstruction penalty on real images, similar to the self-supervised discriminator proposed by Lui et al. . The additional regularization term restricts the discriminator from overfitting, while encouraging it to learn relevant features about the input that will provide a useful training signal for the generator.
5 Training
Experiments
In this section we report both quantitative and qualitative experimental results on different settings. First, we compare the generative performance of GSN with recent state-of-the-art approaches. Second, we provide extensive ablation experiments that show the quantitative improvement obtained by the individual components of our model. Finally, we report results on view synthesis by inverting GSN and querying the model to predict views of a scene given a set of input observations.
We evaluate the generative performance of our model on three datasets: i) the VizDoom environment , a synthetic computer-generated world, ii) the Replica dataset containing 18 scans of realistic scenes that we render using the Habitat environment , and iii) the Active Vision Dataset (AVD) consisting of 20k images with noisy depth measurements from 9 real world scenes.We will release code for reproducibility. Images are resized to resolution for all generation experiments. To generate data for the VizDoom and Replica experiments we render sequences of steps each, where an interactive agent explores a scene collecting RGB and depth observations as well as camera poses. Sequences for AVD are defined by an agent performing a random walk through the scenes according to rules defined in . At training time we sample sub-sequences and express camera poses relative the middle frame of the sub-sequence. This normalization enforces an egocentric coordinate system whose origin is placed at the center of W (see supplementary material for details).
We compare GSN to two recent approaches for generative modelling of radiance fields: GRAF and -GAN . For fair comparison all models use the same base architecture and training pipeline, with two main differences between models. The first difference is that GSN uses locally conditioned radiance fields, while GRAF and -GAN use global conditioning. The second difference is the type of layer used in the radiance field generator network: GRAF utilizes linear layers, -GAN employs modulated sine layers, and GSN uses modulated linear layers. Quantitative performance is evaluated with the Fréchet Inception Distance (FID) and SwAV-FID metrics, which measure the distance between the distributions of real and generated images in pretrained image embedding spaces. We sample 5,000 real and 5,000 generated images when calculating either of these metrics.
Tab. 2 shows the results of our generative modelling comparison. Despite our best attempts to tune -GAN’s hyperparameters, we find that it struggles to learn detailed depth in this setting, which significantly impedes render quality and leads to early collapse of the model. GSN achieves much better performance than GRAF or -GAN across all datasets, obtaining an improvement of 10-14 absolute points on FID. We attribute GSN’s higher performance to the increased expressiveness afforded by the locally conditioned radiance field, and not the specific layer type used (compare GRAF in Tab. 2 to GSN + global latents in Tab. 3). As a measure of qualitative results we show scenes sampled from GSN trained on the Replica dataset at resolution in Fig. 1, and on the VizDoom, Replica, and AVD datasets at resolution in Fig. 4.
Latent Space Interpolation To confirm that GSN is learning a useful scene prior we demonstrate in Fig. 5 some examples of interpolation in the global latent space. The model aligns both geometry and appearance features such that traversing the latent space transitions smoothly between scenes without producing unrealistic off-manifold samples.
2 Ablation
To further analyze our contributions, we report extensive ablation experiments where we show how our design choices affect the generative performance of GSN. For ablation, we perform our analysis on the Replica dataset at 64 ×64 resolution and analyze the following factors: (i) the choice of latent code representation (global vs local); (ii) the effects of global and local coordinate systems on the learned representation; (iii) the effect of sampled trajectory length; (iv) the depth resolution needed in order to successfully learn a scene prior; (v) the regularization applied to the discriminator.
What are the benefits of a local coordinate system? Enabling feature sharing in local coordinate systems to (e.g. a latent can be used to represent the same scene part irrespective of its position on the grid). To empirically validate this hypothesis we train GSN with both global and local coordinate systems. We then sample from the prior obtaining 5k latent grids W for each model. Next, we perform a simple rigid re-arrangement of latent codes in W by applying a 2D rotation at different angles and measuring FID by using the resulting grids to predict a radiance field. A local coordinate system is significantly more robust to re-arranging local latent codes than a global one (Fig. 6; see supplementary material for qualitative results).
How long do camera trajectories need to be? The length of trajectories that collect the camera poses affects the representation of large scenes. Given that we normalize camera poses w.r.t. the middle step in a trajectory, if the trajectories are too short the model will only explore small local neighbourhoods and will not be able to deal with long displacements in camera poses during test time. Models trained on short trajectories fail to generalize to long trajectories during evaluation (Fig. 7). However, models trained with long trajectories do not struggle when evaluated on short trajectories, as evidenced by the stability of the FID when evaluated on short trajectories.
How much depth information is required during training? The amount of depth information used during training affects GSN. To test the sensitivity of GSN to depth information, we down-sample both real and generated depth maps to the target resolution, then up-sample back to the original resolution before concatenating with the corresponding image and feeding them to the discriminator. Without depth information GSN fails to train (Tab. 4). However, we demonstrate that the depth resolution can be reduced to a single pixel without finding statistically significant reduction in the generated image quality. This is a significant result, as it enables GSN to be trained in settings with sparse depth information in future work. We speculate that in early stages of training the added depth information guides the model to learn some aspect of depth, after which it can learn without additional supervision.
Does need to be regularized? Different forms of regularizing the discriminator affect the quality of generated samples. We discover that greatly benefits from regularization (Tab. 5). In particular, data augmentation and the discriminator reconstruction penalty are both crucial; training to rapidly diverge without either of these components. The R1 gradient penalty offers an additional improvement of training stability helping adversarial learning to converge.
3 View Synthesis
We now turn to the task of view synthesis, where we show how GSN performs in comparison with two approaches for scene-level view synthesis of free moving cameras: Generative Temporal models with Spatial Memory (GTM-SM) and Incremental Scene Synthesis (ISS) . Taking the definition in , the view synthesis task is defined as follows: for a given step in a sequence we want the model to predict the target views conditioned on the source views along the camera trajectory. Note that this view synthesis problem is unrelated to video prediction, since the scenes are static and camera poses for source and target views are given. To tackle this task both GTM-SM and ISS rely on auto-encoder architectures augmented with memory mechanisms that directly adapt to the conditional nature of the task. For GSN to deal with this task we follow standard practices for GAN inversion (see supplementary material for details on the encoder architecture and inversion algorithm). We invert source views into the prior and use the resulting latent to locally condition the radiance field and render observations using the camera poses of the target views .
Following we report results on the Vizdoom and AVD datasets in Tab. 6.There is no code or data release for . In private communication with the authors of we discussed the data splits and settings used for generating trajectories and follow them as closely as possible. We thank them for their help. We report two different aspects of view synthesis: the capacity to fit the source views or Memorization (e.g. reconstruction), and the ability to predict the target views or Hallucination (e.g. scene completion) using L1 and SSIM metrics.We note that while these metrics are a good proxy for reconstruction quality, they do not tell a complete picture in the hallucination setting due to the stochastic nature of the view synthesis problem. GSN outperforms both GTM-SM and ISS for nearly all tasks (Tab. 6), even though it was not trained to explicitly learn a mapping from to (see supplementary material for qualitative results). We attribute this success to the powerful scene prior learned by GSN.
Conclusions
We have made great strides towards generative modeling of unconstrained, complex and realistic indoor scenes. We introduced GSN, a generative model that decomposes a scene into many local radiance fields, which can be rendered by a free moving camera. Our decomposition scheme scales efficiently to big scenes while preserving details and distribution coverage. We show that GSN can be trained on multiple scenes to learn a rich scene prior, while rendered trajectories on scenes sampled from the prior are smooth and consistent, maintaining scene coherence. The prior distribution learned by GSN can also be used to infer observations from arbitrary cameras given a sparse set of observations. A multitude of avenues for future work are now enabled by GSN, from improving the rendering performance and training on large-scale datasets to exploring the wide range of down-stream tasks that benefit from a learned scene prior like model-based reinforcement learning , SLAM or 3D completion for immersive photography .
References
Appendix A Model Architectures and Training Details
In this section we summarize the model architectures, hyperparameter settings, and other training details used for producing the GSN models presented in this paper.
The mapping network maps the global latent code z to an intermediate non-linear latent space . All models in our experiments, including those that do not use the global generator, use a mapping network which consists of a normalization step followed by three linear layers with LeakyReLU activations , as shown in Tab. 7.
A.2 Global Generator
The purpose of the global generator (Tab. 8) is to map from a single global latent code z to a 2D grid of local latent codes W which represent the spatial layout of the scene. The global generator is composed of successive modulated convolutional layers that are conditioned on the global latent code. Following StyleGAN , the model learns a constant input for the first layer. The first layers in every pair of modulated convolutional layers thereafter upsamples the feature map resolution by , which is implemented as a transposed convolution with a stride of 2, followed by bilinear filtering .
The output resolution of the global generator (which is also the spatial resolution of W) is a hyperparameter which effectively controls the size of the spatial region represented by each individual local latent code. We set the global generator output resolution to for all experiments.
A.3 Local Generator
The local generator is composed of a locally conditioned radiance field network which maps coordinates and view direction to appearance and occupancy , a volumetric rendering step which accumulates along sampled rays to convert and values to feature vectors, and a refinement network which upsamples feature maps to higher resolution RGB images.
The locally conditioned radiance field network (Fig. 8) mimics the architecture of the the original NeRF network . To condition the network such that it can represent many different radiance fields we swap out the fixed linear layers for modulated linear layers similar to those used in CIPS , where each modulated linear layer is conditioned on . Each modulated linear layer has 128 channels.
When performing volumetric rendering we threshold occupancy values with a softplus as in D-NeRF as opposed to the standard ReLU, as we find it leads to more stable training. For all experiments we sample 64 samples per ray. When generating images we sample the radiance field network to produce feature maps at resolution, and when generating resolution images we sample feature maps at resolution.
Once volumetric rendering has been performed we upsample the resulting feature map with refinement blocks (Fig. 10) until the desired resolution is achieved, then apply a sigmoid to bound the final output, as in GIRAFFE . In general, we found that sampling higher resolution feature maps directly from the radiance field produced higher quality results compared to sampling at low resolution and applying many refinement blocks, but the computational cost is significantly higher.
A.4 Discriminator
The discriminator is based on the architecture used in StyleGAN2 (Tab. 9), including residual blocks (Fig. 10) and minibatch standard deviation layer . When including depth information as input to the discriminator we normalize it to $$. In the case that the radiance field network is sampled at a resolution lower than the final output resolution (such as when using the refinement network), then resulting depth maps will have lower resolution than real examples. To prevent the discriminator from using this difference in detail to differentiate real and fake samples we downsample all real depth maps to match the resolution of the generated depth maps, then upsample them both back to full resolution.
The decoder (Tab. 10) takes as input resolution feature maps from the discriminator (before the minibatch standard deviation layer), and applies successive transposed convolutions with a stride of 2 and bilinear filtering to upsample the input until the original resolution is recovered.
A.5 Sampling Camera Poses
We use poses from camera trajectories in the training set as candidate poses during sampling, as real camera poses better reflect the true distribution of occupable locations compared to uniformly sampling over the entire scene region. Sampled camera poses are normalized and expressed relative to the camera pose in the middle of the trajectory. This normalization enforces an egocentric coordinate system whose origin is placed at the center of W. Note that despite working with trajectories of multiple camera poses, we still only sample a single camera pose per generated scene during training.
A.6 Training Details
We use the RMSprop optimizer with , , and a learning rate of 0.002 for both the generator and discriminator. Following StyleGAN , we set the learning rate of the mapping network less than the rest of the network for improved training stability. Equalized learning rate is used for all learnable parameters, and an exponential moving average of the generator weights with a decay of 0.999 is used to stabilize test-time performance. Differentiable data augmentations such as random translation, colour jitter, and Cutout are applied to all inputs to the discriminator in order to combat overfitting. To save compute, the R1 gradient penalty is applied using a lazy regularization strategy by applying it every 16 iterations. We set to 0.01 and to 1000 for all experiments.
All resolution models used for the generation performance evaluation (GSN and otherwise) were trained for 500k iterations with a batch size of 32. Training takes 4 days on two NVIDIA A100 GPUs with 40GB of memory each. Mixed precision training is applied to the generator for a small reduction in memory cost and training time. We do not apply mixed precision training to the discriminator as training stability decreases in this case.
Appendix B Inverting GSN for View Synthesis
where the first term encourages the reconstruction of the local latent grid, and the second term encourages samples from locally conditioned radiance field to match the original input views.
At inference time, given a trained encoder , we feed the source views through our encoder to predict an initialization latent code grid that we then optimize via SGD for iterations to get . Given that scenes do not share a canonical orientation, we predict at multiple rotation angles of about the axis to find the generator’s preferred orientation and use this orientation during optimization (note that relative transformations between camera poses do not change with this global transformation). We define the preferred orientation as the one that minimizes an auto-encoding LPIPS loss . The optimization process is performed by freezing the weights of and computing a reconstruction loss w.r.t. . We then use in the locally conditioned radiance field and render observations using the camera poses of to produce (i.e. to auto-encode source views), while also rendering from the camera poses of the target views to produce (i.e. to predict unseen parts of the scene). Future work will explore in depth how to improve the quality and efficiency of the inversion approach for GSN-based models where the generative model tends to prefer a certain orientation. In Fig. 11 we show qualitative results for view synthesis on Vizdoom on held out sequences not seen during training. We can see how GSN learns a robust prior that is able to fill in the scene with plausible completions (e.g. row 5 and 8), even if those completions do not strictly minimize the L1 reconstruction loss.
In addition, Fig. 12 shows qualitative view synthesis results on the Replica dataset , showing the applicability of GSN for view synthesis on realistic data. In this experiment we follow the settings described for Vizdoom in terms of and . In Fig. 12 the top 3 rows show results on data from the training set (e.g. scenes that were observed during training) and the bottom 3 rows show test set results (e.g. results on unseen scenes). We can see how GSN successfully uses the prior learned from training data to find a plausible scene completion for unseen scenes that respects the global scene geometry.
Appendix C Qualitative Results on Local vs. Global Coordinate Systems
In this section we demonstrate the robustness of GSN w.r.t. re-arrangement of the latent codes in W. In order to do so we sample different scenes from our learned prior and apply a rigid transformation to their corresponding W (a 2D rotation). In principle, this rotation should amount to a rotation of the scene represented by W that does not change the radiance field prediction. To qualitatively evaluate this effect we sample different scenes and rotate their corresponding W by degrees while (i) rotating the camera by the same amount about the axis so that the rendered image should remain constant and (ii) keeping the camera fixed so that the scene should rotate. In Fig. 13-14 we show the result of the setting in (i) for a local and global coordinate system respectively. In these results we see how a local coordinate system is drastically more robust to re-arrangements of the latent codes than a global coordinate system. In addition, we show results for the setting in (ii) in Fig. 15-16 for local and global coordinate systems respectively. In this case, we see how given a fixed camera, a rotation of W amounts to rotating the scene. In this case we can also see how a local coordinate system results in higher rendering quality compared to that of a global coordinate system, which suffers from degradation as the rotation angle increases.
Appendix D Scene Editing
A nice property of the local latent grid produced by GSN is that it can be used to perform scene editing by directly altering W. This property allows us a degree of manual control for scene synthesis beyond what we get from randomly sampling the generator. While the low resolution of W used in current models currently limits us to high level scene modifications, training with larger local latent grids could allow for more fine-grained control over scenes, such as rearrangement of furniture.
We find that, as with most image composition operations, the results of scene editing appear most convincing when the inputs are well aligned in terms of appearance and geometry. We demonstrate editing operations by manipulating the codes from single scenes, since we don’t need to worry about matching appearance and geometry, but multiple scenes could be combined if they were similar enough. In Fig. 17-18 we manipulate the local latent codes by mirroring them along the horizontal axis to produce unique scenes.