GRAF: Generative Radiance Fields for 3D-Aware Image Synthesis

Katja Schwarz, Yiyi Liao, Michael Niemeyer, Andreas Geiger

Introduction

Convolutional generative adversarial networks have demonstrated impressive results in synthesizing high-resolution images from unstructured image collections. However, despite this success, state-of-the-art models struggle to properly disentangle the underlying generative factors including 3D shape and viewpoint. This is in stark contrast to humans who have the remarkable ability to reason about the 3D structure of the world and imagine objects from novel viewpoints.

As reasoning in 3D is fundamental for applications in robotics, virtual reality or data augmentation, several recent works consider the task of 3D-aware image synthesis , aiming at photorealistic image generation with explicit control over the camera pose.

In contrast to 2D generative adversarial networks, approaches for 3D-aware image synthesis learn a 3D scene representation which is explicitly mapped to an image using differentiable rendering techniques, hence providing control over both, scene content and viewpoint. Since 3D supervision or posed images are often hard to obtain in practice, recent works try to solve this task using 2D supervision only . Towards this goal, existing approaches generate discretized 3D representations, i.e., a voxel-grid representing either the full 3D object or intermediate 3D features as illustrated in Fig. 1. While modeling the 3D object in color space allows for exploiting differentiable rendering, the cubic memory growth of voxel-based representations limits to low resolution and results in visible artifacts. Intermediate 3D features are more compact and scale better with image resolution. However, this requires to learn a 3D-to-2D mapping for decoding the abstract features to RGB values, resulting in entangled representations which are not consistent across views at high resolutions.

In this paper, we demonstrate that the dilemma between coarse outputs and entangled latents can be resolved using conditional radiance fields, a conditional variant of a recently proposed continuous representation for novel view synthesis . More specifically, we make the following contributions: i) We propose GRAF, a generative model for radiance fields for high-resolution 3D-aware image synthesis from unposed images. In addition to viewpoint manipulations, our approach allows to modify shape and appearance of the generated objects. ii) We introduce a patch-based discriminator that samples the image at multiple scales and which is key to learn high-resolution generative radiance fields efficiently. iii) We systematically evaluate our approach on synthetic and real datasets. Our approach compares favorably to state-of-the-art methods in terms of visual fidelity and 3D consistency while generalizing to high spatial resolutions. We release our code and datasets at https://github.com/autonomousvision/graf.

Related Work

Image Synthesis: Generative Adversarial Networks (GANs) have significantly advanced the state-of-the-art in photorealistic image synthesis. In order to make the image synthesis process more controllable, several recent works have proposed to disentangle the underlying factors of variation . However, all of these methods ignore the fact that 2D images are obtained as projections of the 3D world. While some of the methods demonstrate that the disentangled factors capture 3D properties to some extend , modeling the image manifold using 2D convolutional networks remains a difficult task, in particular when seeking representations that faithfully disentangle viewpoint variations from object appearance and identity. Instead of directly modeling the 2D image manifold, we therefore follow a recent line of works which aims at generating 3D representations and explicitly models the image formation process.

3D-Aware Image Synthesis: Learning-based novel view synthesis has been intensively investigated in the literature . These methods generate unseen views from the same object and typically require camera viewpoints as supervision. While some works generalize across different objects without requiring to train an individual network per object, they do not yield a full probabilistic generative model for drawing unconditional random samples. In contrast, we are interested in generating novel objects from multiple views by learning a 3D-aware generative model from unposed 2D images.

Several recent works exploit generative 3D models for 3D-aware image synthesis . Many methods require 3D supervision or assume 3D information as input . E.g. Texture Fields synthesize novel textures conditioned on a particular 3D shape. Consequently, they require a 3D shape as input and colored surface points as supervision. Instead, we learn a generative model for both shape and texture from 2D images alone. This is a difficult task which only few works have attempted so far: platonicGAN learns a textured 3D voxel representation from 2D images using differentiable rendering techniques. However, such voxel-based representations are memory intensive, precluding image synthesis at high image resolutions. In this work, we avoid these memory limitations by using a continuous representation which allows for rendering images at arbitrary resolution. HoloGAN and some related works learn a low-dimensional 3D feature combined with a learnable 3D-to-2D projection. However, as evidenced by our experiments, learned projections can lead to entangled latents (e.g., object identity and viewpoint), particularly at high resolutions. While 3D consistency can be encouraged using additional constraints , we take advantage of differentiable volume rendering techniques which do not need to be learned and thus incorporate 3D consistency into the generative model by design.

Implicit Representations: Recently, implicit representations of 3D geometry have gained popularity in learning-based 3D reconstruction . Key advantages over voxel or mesh-based methods are that they do not discretize space and are not restriced in topology. Recent hybrid continuous grid representations extend implicit representations to complicated or large scale scenes but require 3D input and do not consider texture. Another line of works propose to learn continuous shape and texture representations from posed multi-view images only, by making the rendering process differentiable. As these models are limited to single objects or scenes of small geometric complexity, Mildenhall et al. propose to represent scenes as neural radiance fields which allow for multi-view consistent novel-view synthesis of more complex, real-world scenes from posed 2D images. They demonstrate compelling results on this task, however, their method requires many posed views, needs to be retrained for each scene, and cannot generate novel scenes. Inspired by this work, we exploit a conditional variant of this representation and show how a rich generative model can be learned from a collection of unposed 2D images as input.

Method

We consider the problem of 3D-aware image synthesis, i.e., the task of generating high-fidelity images while providing explicit control over camera rotation and translation. We argue for representing a scene by its radiance field as such a continuous representation scales well wrt. image resolution and memory consumption while allowing for a physically-based and parameter-free projective mapping. In the following, we first briefly review Neural Radiance Fields (NeRF) which forms the basis for the proposed Generative Radiance Field (GRAF) model.

As demonstrated in , the positional encoding γ(⋅)\gamma(\cdot) enables better fitting of high-frequency signals compared to directly using x\mathbf{x} and d\mathbf{d} as input to the multi-layer perceptron fθ(⋅)f_{\theta}(\cdot). We confirm this with an ablation study in our supp. material. As the volume color c\mathbf{c} varies more smoothly with the viewing direction than with the 3D location, the viewing direction is typically encoded using fewer components, i.e., Ld<LxL_{\mathbf{d}}<L_{\mathbf{x}}.

Volume Rendering: For rendering a 2D image from the radiance field fθ(⋅)f_{\theta}(\cdot), Mildenhall et al. approximate the intractable volumetric projection integral using numerical integration. More formally, let {(cri,σri)}i=1N\{(\mathbf{c}_{r}^{i},\sigma_{r}^{i})\}_{i=1}^{N} denote the color and volume density values of NN random samples along a camera ray rr. The rendering operator π(⋅)\pi(\cdot) maps these values to an RGB color value cr\mathbf{c}_{r}:

The RGB color value cr\mathbf{c}_{r} is obtained using alpha composition

where TriT_{r}^{i} and αri\alpha_{r}^{i} denote the transmittance and alpha value of sample point ii along ray rr and δri=∥xri+1−xri∥2\delta_{r}^{i}=\left\|\mathbf{x}_{r}^{i+1}-\mathbf{x}_{r}^{i}\right\|_{2} is the distance between neighboring sample points. Given a set of posed 2D images of a single static scene, Mildenhall et al. optimize the parameters θ\theta of the neural radiance field fθ(⋅)f_{\theta}(\cdot) by minimizing a reconstruction loss (sum of squared differences) between the observations and the predictions. Given θ\theta, novel views can be synthesized by invoking π(⋅)\pi(\cdot) for each pixel/ray.

2 Generative Radiance Fields

In this work, we are interested in radiance fields as a representation for 3D-aware image synthesis. In contrast to , we do not assume a large number of posed images of a single scene. Instead, we aim at learning a model for synthesizing novel scenes by training on unposed images. More specifically, we utilize an adversarial framework to train a generative model for radiance fields (GRAF).

We sample the camera pose ξ=[R∣t]\boldsymbol{\xi}=[\mathbf{R}|\mathbf{t}] from a pose distribution pξp_{\xi}. In our experiments, we use a uniform distribution on the upper hemisphere for the camera location with the camera facing towards the origin of the coordinate system. Depending on the dataset, we also vary the distance of the camera from the origin uniformly. We choose K\mathbf{K} such that the principle point is in the center of the image.

Ray Sampling: The K×KK\times K patch P(u,s)\mathcal{P}(\mathbf{u},s) is determined by a set of 2D image coordinates

which describe the location of every pixel of the patch in the image domain Ω\Omega as illustrated in Fig. 4. Note that these coordinates are real numbers, not discrete integers which allows us to continuously evaluate the radiance field. The corresponding 3D rays are uniquely determined by P(u,s)\mathcal{P}(\mathbf{u},s), the camera pose ξ\boldsymbol{\xi} and intrinsics K\mathbf{K}. We denote the pixel/ray index by rr, the normalized 3D rays by dr\mathbf{d}_{r} and the number of rays by RR where R=K2R=K^{2} during training and R=WHR=WH during inference.

3D Point Sampling: For numerical integration of the radiance field, we sample NN points {xri}\{\mathbf{x}_{r}^{i}\} along each ray rr. We use the stratified sampling approach of , see supp. material for details.

The network architecture of our conditional radiance field gθg_{\theta} is illustrated in Fig. 4. We first compute a shape encoding h\mathbf{h} from the positional encoding of x\mathbf{x} and the shape code zs\mathbf{z}_{s}. A density head σθ\sigma_{\theta} transforms this encoding to the volume density σ\sigma. For predicting the color c\mathbf{c} at 3D location x\mathbf{x}, we concatenate h\mathbf{h} with the positional encoding of d\mathbf{d} and the appearance code za\mathbf{z}_{a} and pass the resulting vector to a color head cθc_{\theta}. We compute σ\sigma independently of the viewpoint d\mathbf{d} and the appearance code za\mathbf{z}_{a} to encourage multi-view consistency while disentangling shape from appearance. This encourages the network to use the latent codes zs\mathbf{z}_{s} and za\mathbf{z}_{a} to model shape and appearance, respectively, and allows for manipulating them separately during inference. More formally, we have:

All mappings (hθh_{\theta}, cθc_{\theta} and σθ\sigma_{\theta}) are implemented using fully connected networks with ReLU activations. To avoid notation clutter, we use the same symbol θ\theta to denote the parameters of each network.

2.2 Discriminator

The discriminator DϕD_{\phi} is implemented as a convolutional neural network (see supp. material for details) which compares the predicted patch P′\mathbf{P}^{\prime} to a patch P\mathbf{P} extracted from a real image I\mathbf{I} drawn from the data distribution pDp_{\mathcal{D}}. For extracting a K×KK\times K patch from real image I\mathbf{I}, we first draw ν=(u,s)\boldsymbol{\nu}=(\mathbf{u},s) from the same distribution pνp_{\nu} which we use for drawing the generator patch above. We then sample the real patch P\mathbf{P} by querying I\mathbf{I} at the 2D image coordinates P(u,s)\mathcal{P}(\mathbf{u},s) using bilinear interpolation. In the following, we use Γ(I,ν)\Gamma(\mathbf{I},\boldsymbol{\nu}) to denote this bilinear sampling operation. Note that our discriminator is similar to PatchGAN , except that we allow for continuous displacements u\mathbf{u} and scales ss while PatchGAN uses s=1s=1. It is further important to note that we do not downsample the real image I\mathbf{I} based on ss, but instead query I\mathbf{I} at sparse locations to retain high-frequency details, see Fig. 4.

Experimentally, we found that a single discriminator with shared weights is sufficient for all patches, even though these are sampled at random locations with different scales. Note that the scale determines the receptive field of the patch. To facilitate training, we thus start with patches of larger receptive fields to capture the global context. We then progressively sample patches with smaller receptive fields to refine local details.

2.3 Training and Inference

Let I\mathbf{I} denote an image from the data distribution pDp_{\mathcal{D}} and let pνp_{\nu} denote the distribution over random patches (see Section 3.2.1). We train our model using a non-saturating GAN objective with R1-regularization

where f(t)=−log⁡(1+exp⁡(−t))f(t)=-\log(1+\exp(-t)) and λ\lambda controls the strength of the regularizer. We use spectral normalization and instance normalization in our discriminator and train our approach using RMSprop with a batch size of 88 and a learning rate of 0.00050.0005 and 0.00010.0001 for generator and discriminator, respectively. At inference, we randomly sample zs\mathbf{z}_{s}, za\mathbf{z}_{a} and ξ\boldsymbol{\xi}, and predict a color value for all pixels in the image. Details on the network architectures can be found in the supp. material.

Experiments

Datasets: We consider two synthetic and three real-world datasets in our experiments. To analyze our approach in a controlled setting we render 150150k Chairs from Photoshapes following the rendering protocol of . We further use the Carla Driving simulator to create 1010k images of 1818 car models with randomly sampled colors and realistic texture and reflectance properties (Cars). We also validate our approach on three real-world datasets. We use the Faces dataset which comprises celebA and celebA-HQ for image synthesis up to resolution 1282128^{2} and 5122512^{2} pixels, respectively. In addition, we consider the Cats dataset and the Caltech-UCSD Birds-200-2011 dataset. For the latter, we use the available instance masks to composite the birds onto a white background.

Baselines: We compare our approach to two state-of-the-art models for 3D-aware image synthesis using the authors’ implementationsplatonicGAN: https://github.com/henzler/platonicganHoloGAN: https://github.com/thunguyenphuoc/HoloGAN: platonicGAN generates a voxel-grid of the 3D object which is projected to the image plane using differentiable volumetric rendering. HoloGAN instead generates an abstract voxelized feature representation and learns the mapping from 3D to 2D using a combination of 3D and 2D convolutions. To analyze the consequences of a learned projection we further consider a modified version of HoloGAN (HoloGAN w/o 3D Conv) in which we reduce the capacity of the learned mapping by removing the 3D convolutional layers. For reference, we also compare our results to a state-of-the-art 2D GAN model with a ResNet architecture.

Evaluation Metrics: We quantify image fidelity using the Frechet Inception Distance (FID) and additionally report the Kernel Inception Distance (KID) in the supp. material. To assess 3D consistency we perform 3D reconstruction for images of size 2562256^{2} pixels using COLMAP . We adopt Minimum Matching Distance (MMD) to measure the chamfer distance (CD) between 100 reconstructed shapes and their closest shapes in the ground truth for quantitative comparison and show qualitative results for the reconstructions.

We now study several key questions relevant to the proposed model. We first compare our model to several baselines in terms of their ability to generate high-fidelity and high-resolution outputs. We then analyze the implications of learned projections and the importance of our multi-scale discriminator.

We first compare our model against the baselines using an image resolution of 64264^{2} pixels. As shown in Fig. 5, all methods are able to disentangle object identity and camera viewpoint. However, platonicGAN has difficulties in representing thin structures and both platonicGAN and HoloGAN lead to visible artifacts in comparison to the proposed model. This is also reflected by larger FID scores in Table 2.

On Faces and Cats, HoloGAN achieves FID scores similar to our approach as both datasets exhibit only little variation in the azimuth angle of the camera while the other datasets cover larger viewpoint variations. This suggests that it is harder for HoloGAN to accurately capture the appearance of objects from different viewpoints due to its low-dimensional 3D feature representation and the learnable projection. In contrast, our continuous representation does not require a learned projection and renders high-fidelity images from arbitrary views.

Due to the voxel-based representation, platonicGAN becomes very memory intensive when scaled to higher resolutions. Thus, we limit our experiments to HoloGAN and HoloGAN w/o 3D Conv for this analysis. Additionally, we provide results for our model trained at 1282128^{2} pixels, but sampled at higher resolution during inference (Ours sampled). The results in Table 2 show that this significantly improves over naïve bilinear upsampling (Ours upsampled) which indicates that our learned representation generalizes well to higher resolutions. As expected, our method achieves the smallest FID value when trained at full resolution. While our approach outperforms HoloGAN and HoloGAN w/o 3D Conv significantly on the Car dataset, HoloGAN w/o 3D Conv achieves results onpar with us on Faces where viewpoint variations are smaller. Interestingly, we found that HoloGAN w/o 3D Conv achieves lower FID values than the full HoloGAN model originally proposed in despite reduced model capacity. This is due to training instabilities which we observe when training HoloGAN at high resolutions. We even observe mode collapse at a resolution of 5122512^{2} for which we are hence not able to report results.

As illustrated in Fig. 6, HoloGAN (top) fails in disentangling viewpoint from appearance at high resolution, varying different style aspects like facial expression or even completely ignoring the pose input. We identify the learnable projection as the underlying cause for this behavior. In particular, we find that removing the 3D convolutional layers enables HoloGAN to adhere to the input pose more closely, see Fig. 6 (middle). However, images from HoloGAN w/o 3D Conv are still not entirely multi-view consistent.

To better investigate this observation we generate multiple images of the same instance at random viewpoints for both HoloGAN w/o 3D Conv and our approach, and perform dense 3D reconstruction using COLMAP . As reconstruction depends on the consistency across views, reconstruction accuracy can be considered as a proxy for the multi-view consistency of the generated images. As evident from Table 3 and Fig. 7, multi-view stereo works significantly better when using images from our method as input. In contrast, fewer correspondences can be established for HoloGAN w/o 3D Conv which uses learned 2D layers for upsampling. For HoloGAN with 3D convolutions, performance degrades even further as shown in the supp. material. We thus conjecture that learned projections should generally be avoided.

Fig. 8 shows that in addition to disentangling camera and scene properties, our approach learns to disentangle shape and appearance which can be controlled during inference via zs\mathbf{z}_{s} and za\mathbf{z}_{a}. For Cars and Chairs the appearance code controls the color of the object while for Faces it encodes skin and hair color.

To investigate whether we sacrifice image quality by using the proposed multi-scale patch discriminator, we compare our multi-scale discriminator (Patch) to a discriminator that receives the entire image as input (Full). As this is very memory intensive, we only consider images of resolution 64264^{2} and use half the hidden dimensions for hθh_{\theta} and cθc_{\theta} (dim/2). Table 5 shows that our patch discriminator achieves similar performance to the full image discriminator on Cars and performs even better on CelebA. A possible explanation for this phenomenon is that random patch sampling acts as a data augmentation strategy which helps to stabilize GAN training. In contrast, when using only local patches (s=1s=1), we observe that our generator is not able to learn the correct shape, resulting in a high FID value in Table 5. We conclude that sampling patches at random scales is crucial for robust performance.

We ablate the sensitivity of our model to the chosen focal length on Cars under a fixed radius of 1010 in Table 5. Our model performs very similar for changes within 0.7fdata0.7f_{data} to 1.0fdata1.0f_{data} where fdataf_{data} is the focal length we use to render the training images. Even for larger focal lengths up to 1.8fdata1.8f_{data} and with an orthographic projection we observe good performance. Only for very small focal lengths the generated images show distortions at the image borders resulting in higher FID values.

Conclusion

We have introduced Generative Radiance Fields (GRAF) for high-resolution 3D-aware image synthesis. We showed that our framework is able to generate high resolution images with better multi-view consistency compared to voxel-based approaches. However, our results are limited to simple scenes with single objects. We believe that incorporating inductive biases, e.g., depth maps or symmetry, will allow for extending our model to even more challenging real-world scenarios in the future.

Broader Impact

3D-aware image synthesis is a relatively novel research area and does not yet scale to generating complex real-world scenes, preventing immediate applications for society. However, our work takes an important step towards this goal as it enables high-fidelty reconstruction at resolutions beyond 64264^{2} pixels while requiring no 3D supervision as input. In the long run, we hope that our resesarch will facilitate the use of 3D-aware generative models in applications such as virtual reality, data augmentation or robotics. For example, intelligent systems such as autonomous vehicles require large amounts of data for training and validation which will be impossible to collect using static offline datasets. We believe that building generative, photo-realistic and large-scale 3D models of our world will ultimately allow for cost-efficient data collection and simulation. While many use-cases are possible, we believe that these types of models can be particularly beneficial to close the existing domain gap between real-world and synthetic data. However, working with generative models also requires care. While generating photorealistic 3D-scenarios is very intriguing it also bears the risk of manipulation and the creation of misleading content. In particular, models that can create 3D-consistent fake images might increase credibility of fake contents and might potentially fool systems that rely on multi-view consistency, e.g., modern face recognition systems. Therefore, we believe that it is of equal importance for the community to develop methods which are able to clearly distinguish between synthetic and real-world content. We see encouraging progress in this area, e.g., .

Acknowledgments

We acknowledge the financial support by the BMWi in the project KI Delta Learning (project number 19A19013O) and the support from the BMBF through the Tuebingen AI Center (FKZ: 01IS18039A). We thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting Katja Schwarz and Michael Niemeyer. This work was supported by an NVIDIA research gift.

References