GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields

Michael Niemeyer, Andreas Geiger

Introduction

The ability to generate and manipulate photorealistic image content is a long-standing goal of computer vision and graphics. Modern computer graphics techniques achieve impressive results and are industry standard in gaming and movie productions. However, they are very hardware expensive and require substantial human labor for 3D content creation and arrangement.

In recent years, the computer vision community has made great strides towards highly-realistic image generation. In particular, Generative Adversarial Networks (GANs) emerged as a powerful class of generative models. They are able to synthesize photorealistic images at resolutions of 102421024^{2} pixels and beyond .

Despite these successes, synthesizing realistic 2D images is not the only aspect required in applications of generative models. The generation process should also be controllable in a simple and consistent manner. To this end, many works investigate how disentangled representations can be learned from data without explicit supervision. Definitions of disentanglement vary , but commonly refer to being able to control an attribute of interest, e.g. object shape, size, or pose, without changing other attributes. Most approaches, however, do not consider the compositional nature of scenes and operate in the 2D domain, ignoring that our world is three-dimensional. This often leads to entangled representations (\figreffig:controllable-generation) and control mechanisms are not built-in, but need to be discovered in the latent space a posteriori. These properties, however, are crucial for successful applications, e.g. a movie production where complex object trajectories need to be generated in a consistent manner.

Several recent works therefore investigate how to incorporate 3D representations, such as voxels , primitives , or radiance fields , directly into generative models. While these methods allow for impressive results with built-in control, they are mostly restricted to single-object scenes and results are less consistent for higher resolutions and more complex and realistic imagery (e.g. scenes with objects not in the center or cluttered backgrounds).

Contribution: In this work, we introduce GIRAFFE, a novel method for generating scenes in a controllable and photorealistic manner while training from raw unstructured image collections. Our key insight is twofold: First, incorporating a compositional 3D scene representation directly into the generative model leads to more controllable image synthesis. Second, combining this explicit 3D representation with a neural rendering pipeline results in faster inference and more realistic images. To this end, we represent scenes as compositional generative neural feature fields (\figreffig:teaser). We volume render the scene to a feature image of relatively low resolution to save time and computation. A neural renderer processes these feature images and outputs the final renderings. This way, our approach achieves high-quality images and scales to real-world scenes. We find that our method allows for controllable image synthesis of single-object as well as multi-object scenes when trained on raw unstructured image collections. Code and data is available at https://github.com/autonomousvision/giraffe.

Related Work

GAN-based Image Synthesis: Generative Adversarial Networks (GANs) have been shown to allow for photorealistic image synthesis at resolutions of 102421024^{2} pixels and beyond . To gain better control over the synthesis process, many works investigate how factors of variation can be disentangled without explicit supervision. They either modify the training objective or network architecture , or investigate latent spaces of well-engineered and pre-trained generative models . All of these works, however, do not explicitly model the compositional nature of scenes. Recent works therefore investigate how the synthesis process can be controlled at the object-level . While achieving photorealistic results, all aforementioned works model the image formation process in 2D, ignoring the three-dimensional structure of our world. In this work, we advocate to model the formation process directly in 3D for better disentanglement and more controllable synthesis.

Implicit Functions: Using implicit functions to represent 3D geometry has gained popularity in learning-based 3D reconstruction and has been extended to scene-level reconstruction . To overcome the need of 3D supervision, several works propose differentiable rendering techniques. Mildenhall et al. propose Neural Radiance Fields (NeRFs) in which they combine an implicit neural model with volume rendering for novel view synthesis of complex scenes. Due to their expressiveness, we use a generative variant of NeRFs as our object-level representation. In contrast to our method, the discussed works require multi-view images with camera poses as supervision, train a single network per scene, and are not able to generate novel scenes. Instead, we learn a generative model from unstructured image collections which allows for controllable, photorealistic image synthesis of generated scenes.

3D-Aware Image Synthesis: Several works investigate how 3D representations can be incorporated as inductive bias into generative models . While many approaches use additional supervision , we focus on works which are trained on raw image collections like our approach. Henzler et al. learn voxel-based representations using differentiable rendering. The results are 3D controllable, but show artifacts due to the limited voxel resolutions caused by their cubic memory growth. Nguyen-Phuoc et al. propose voxelized feature-grid representations which are rendered to 2D via a reshaping operation. While achieving impressive results, training becomes less stable and results less consistent for higher resolutions. Liao et al. use abstract features in combination with primitives and differentiable rendering. While handling multi-object scenes, they require additional supervision in the form of pure background images which are hard to obtain for real-world scenes. Schwarz et al. propose Generative Neural Radiances Fields (GRAF). While achieving controllable image synthesis at high resolutions, this representation is restricted to single-object scenes and results degrade on more complex, real-world imagery. In contrast, we incorporate compositional 3D scene structure into the generative model such that it naturally handles multi-object scenes. Further, by integrating a neural rendering pipeline , our model scales to more complex, real-world data.

Method

Our goal is a controllable image synthesis pipeline which can be trained from raw image collections without additional supervision. In the following, we discuss the main components of our method. First, we model individual objects as neural feature fields (\secrefsubsec:objects). Next, we exploit the additive property of feature fields to composite scenes from multiple individual objects (\secrefsubsec:scenes). For rendering, we explore an efficient combination of volume and neural rendering techniques (\secrefsubsec:rendering). Finally, we discuss how we train our model from raw image collections (\secrefsubsec:training). \figreffig:method-overview contains an overview of our method.

where tt is a scalar input, e.g. a component of x\mathbf{x} or d\mathbf{d}, and LL the number of frequency octaves. In the context of generative models, we observe an additional benefit of this representation: It introduces an inductive bias to learn 3D shape representations in canonical orientations which otherwise would be arbitrary (see \figreffig:canonical-pose).

Following implicit shape representations , Mildenhall et al. propose to learn Neural Radiance Fields (NeRFs) by parameterizing ff with a multi-layer perceptron (MLP):

where θ\theta indicate the network parameters and Lx,LdL_{\mathbf{x}},L_{\mathbf{d}} the output dimensionalities of the positional encodings.

Generative Neural Feature Fields: While fits θ\theta to multiple posed images of a single scene, Schwarz et al. propose a generative model for Neural Radiance Fields (GRAF) that is trained from unposed image collections. To learn a latent space of NeRFs, they condition the MLP on shape and appearance codes zs,za∼N(0,I)\mathbf{z}_{s},\mathbf{z}_{a}\sim\mathcal{N}(\mathbf{0},I):

where Ms,MaM_{s},M_{a} are the dimensionalities of the latent codes.

In this work we explore a more efficient combination of volume and neural rendering. We replace GRAF’s formulation for the three-dimensional color output c\mathbf{c} with a more generic MfM_{f}-dimensional feature f\mathbf{f} and represent objects as Generative Neural Feature Fields:

Object Representation: A key limitation of NeRF and GRAF is that the entire scene is represented by a single model. As we are interested in disentangling different entities in the scene, we need control over the pose, shape and appearance of individual objects (we consider the background as an object as well). We therefore represent each object using a separate feature field in combination with an affine transformation

In practice, we volume render in scene space and evaluate the feature field in its canonical object space (see \figreffig:teaser):

This allows us to arrange multiple objects in a scene. All object feature fields share their weights and T\mathbf{T} is sampled from a dataset-dependent distribution (see \secrefsubsec:training).

2 Scene Compositions

As discussed above, we describe scenes as compositions of NN entities where the first N−1N-1 are the objects in the scene and the last represents the background. We consider two cases: First, NN is fixed across the dataset such that the images always contain N−1N-1 objects plus the background. Second, NN is varied across the dataset. In practice, we use the same representation for the background as for objects except that we fix the scale and translation parameters sN,tN\mathbf{s}_{N},\mathbf{t}_{N} to span the entire scene, and to be centered at the scene space origin.

While being simple and intuitive, this choice for CC has an additional benefit: We ensure gradient flow to all entities with a density greater than .

3 Scene Rendering

3D Volume Rendering: While previous works volume render an RGB color value, we extend this formulation to rendering an MfM_{f}-dimensional feature vector f\mathbf{f}.

For given camera extrinsics ξ\boldsymbol{\xi}, let {xj}j=1Ns\{\mathbf{x}_{j}\}_{j=1}^{N_{s}} be sample points along the camera ray d\mathbf{d} for a given pixel, and (σj,fj)=C(xj,d)(\sigma_{j},\mathbf{f}_{j})=C(\mathbf{x}_{j},\mathbf{d}) the corresponding densities and feature vectors of the field. The volume rendering operator πvol\pi_{\text{vol}} maps these evaluations to the pixel’s final feature vector f\mathbf{f}:

Using numerical integration as in , f\mathbf{f} is obtained as

where τj\tau_{j} is the transmittance, αj\alpha_{j} the alpha value for xj\mathbf{x}_{j}, and δj=∣∣xj+1−xj∣∣2\delta_{j}={\left|\left|\mathbf{x}_{j+1}-\mathbf{x}_{j}\right|\right|}_{2} the distance between neighboring sample points. The entire feature image is obtained by evaluating πvol\pi_{\text{vol}} at every pixel. For efficiency, we render feature images at resolution 16216^{2} which is lower than the output resolution of 64264^{2} or 2562256^{2} pixels. We then upsample the low-resolution feature maps to higher-resolution RGB images using 2D neural rendering. As evidenced by our experiments, this has two advantages: increased rendering speed and improved image quality.

4 Training

Generator: We denote the full generative process formally as

and NN is the number of entities in the scene, NsN_{s} the number of sample points along each ray, dk\mathbf{d}_{k} is the ray for the kk-th pixel, and xjk\mathbf{x}_{jk} the jj-th sample point for the kk-th pixel / ray.

Discriminator: We parameterize the discriminator DϕD_{\phi} as a CNN with leaky ReLU activation.

Training: During training, we sample the the number of entities in the scene N∼pNN\sim p_{N}, the latent codes zsi,zai∼N(0,I)\mathbf{z}_{s}^{i},\mathbf{z}_{a}^{i}\sim\mathcal{N}(\mathbf{0},I), as well as a camera pose ξ∼pξ\boldsymbol{\xi}\sim p_{\xi} and object-level transformations Ti∼pT\mathbf{T}_{i}\sim p_{T}. In practice, we define pξp_{\xi} and pTp_{T} as uniform distributions over dataset-dependent camera elevation angles and valid object transformations, respectively.Details can be found in the supplementary material. The motivation for this choice is that in most real-world scenes, objects are arbitrarily rotated, but not tilted due to gravity. The observer (the camera in our case), in contrast, can freely change its elevation angle wrt. the scene.

We train our model with the non-saturating GAN objective and R1R_{1} gradient penalty

where f(t)=−log⁡(1+exp⁡(−t))f(t)=-\log(1+\exp(-t)), λ=10\lambda=10, and pDp_{\mathcal{D}} indicates the data distribution.

5 Implementation Details

All object feature fields {hθii}i=1N−1\{h_{\theta_{i}}^{i}\}_{i=1}^{N-1} share their weights and we parametrize them as MLPs with ReLU activations. We use 88 layers with a hidden dimension of 128128 and a density and a feature head of dimensionality 11 and Mf=128M_{f}=128, respectively. For the background feature field hθNNh_{\theta_{N}}^{N}, we use half the layers and hidden dimension. We use Lx=2⋅3⋅10L_{\mathbf{x}}=2\cdot 3\cdot 10 and Ld=2⋅3⋅4L_{\mathbf{d}}=2\cdot 3\cdot 4 for the positional encodings. We sample Ms=64M_{s}=64 points along each ray and render the feature image IV\mathbf{I}_{V} at 16216^{2} pixels. We use an exponential moving average with decay 0.9990.999 for the weights of the generator. We use the RMSprop optimizer with a batch size of 3232 and learning rates of 1\times10−41\text{\times}{10}^{-4} and 5\times10−45\text{\times}{10}^{-4} for the discriminator and generator, respectively. For experiments at 2562256^{2} pixels, we set Mf=256M_{f}=256 and half the generator learning rate to 2.5\times10−42.5\text{\times}{10}^{-4}.

Experiments

Datasets: We report results on commonly-used single-object datasets Chairs , Cats , CelebA , and CelebA-HQ . The first consists of synthetic renderings of Photoshape chairs , and the others are image collections of cat and human faces, respectively. The data complexity is limited as the background is purely white or only takes up a small part of the image. We further report results on the more challenging single-object, real-world datasets CompCars , LSUN Churches , and FFHQ . For CompCars, we randomly crop the images to achieve more variety of the object’s position in the image.We do not apply random cropping for and as we find that they cannot handle scenes with non-centered objects (see supplementary). For these datasets, disentangling objects is more complex as the object is not always in the center and the background is more cluttered and takes up a larger part of the image. To test our model on multi-object scenes, we use the script from to render scenes with 22, 33, 44, or 55 random primitives (Clevr-N). To test our model on scenes with a varying number of objects, we also run our model on the union of them (Clevr-2345).

Baselines: We compare against voxel-based PlatonicGAN , BlockGAN , and HoloGAN , and radiance field-based GRAF (see \secrefsec:rel-work for a discussion of the methods). We further compare against HoloGAN w/o 3D Conv, a variant of proposed in for higher resolutions. We additionally report a ResNet-based 2D GAN for reference.

Metrics: We report the Frechet Inception Distance (FID) score to quantify image quality. We use 2000020000 real and fake samples to calculate the FID score.

Disentangled Scene Generation: We first analyze to which degree our model learns to generate disentangled scene representations. In particular, we are interested if objects are disentangled from the background. Towards this goal, we exploit the fact that our composition operator is a simple addition operation (Eq. 8) and render individual components and object alpha maps (Eq. 10). Note that while we always render the feature image at 16216^{2} during training, we can choose arbitrary resolutions at test time.

fig:disentanglement suggests that our method disentangles objects from the background. Note that this disentanglement emerges without any supervision, and the model learns to generate plausible backgrounds without ever having seen a pure background image, implicitly solving an inpainting task. We further observe that our model correctly disentangles individual objects when trained on multi-object scenes with fixed or varying number of objects. We further find that unsupervised disentanglement is a property of our model which emerges already at the very beginning of training (\figreffig:training-progression). Note how our model synthesizes individual objects before spending capacity on representing the background.

Controllable Scene Generation: As individual components of the scene are correctly disentangled, we analyze how well they can be controlled. More specifically, we are interested if individual objects can be rotated and translated, but also how well shape and appearance can be controlled. In \figreffig:controllable-scene-gen, we show examples in which we control the scene during image synthesis. We rotate individual objects, translate them in 3D space, or change the camera elevation. By modeling shape and appearance for each entity with a different latent code, we are further able to change the objects’ appearances without altering their shape.

Generalization Beyond Training Data: The learned compositional scene representations allow us to generalize outside the training distribution. For example, we can increase the translation ranges of objects or add more objects than there were present in the training data (\figreffig:generalization).

2 Comparison to Baseline Methods

Comparing to baseline methods, our method achieves similar or better FID scores at both 64264^{2} (\tabreftab:tab64) and 2562256^{2} (\tabreftab:tab256) pixel resolutions. Qualitatively, we observe that while all approaches allow for controllable image synthesis on datasets of limited complexity, results are less consistent for the baseline methods on more complex scenes with cluttered backgrounds. Further, our model disentangles the object from the background, such that we are able to control the object independent of the background (\figreffig:qualitative-comparison).

We further note that our model achieves similar or better FID scores than the ResNet-based 2D GAN despite fewer network parameters (0.410.41m compared to 1.691.69m). This confirms our initial hypothesis that using a 3D representation as inductive bias results in better outputs. Note that for fair comparison, we only report methods which are similar wrt. network size and training time (see \tabreftab:netsize).

3 Ablation Studies

Importance of Individual Components: The ablation study in \tabreftab:ablation-study shows that our design choices of RGB skip connections, final activation function, and selected upsampling types improve results and lead to higher FID scores.

Effect of Neural Renderer: A key difference to is that we combine volume with neural rendering. The quantitative (\tabreftab:tab64 and 2) and qualitative comparisons (\figreffig:qualitative-comparison) indicate that our approach leads to better results, in particular for complex, real-world data. Our model is more expressive and can better handle the complexity of real scenes, e.g. note how the neural renderer realistically adapts object appearances to the background (\figreffig:neural-renderer). Further, we observe a rendering speed up: compared to , total rendering time is reduced from 110.1110.1ms to 4.84.8ms, and from 1595.01595.0ms to 5.95.9ms for 64264^{2} and 2562256^{2} pixels, respectively.

Positional Encoding: We use axis-aligned positional encoding for the input point and viewing direction (Eq. 1). Surprisingly, this encourages the model to learn canoncial representations as it introduces a bias to align the object axes with highest symmetry with the canonical axes which allows the model to exploit object symmetry (\figreffig:canonical-pose).

4 Limitations

Dataset Bias: Our method struggles to disentangle factors of variation if there is an inherent bias in the data. We show an example in \figreffig:dataset-bias: In the celebA-HQ dataset, the eye and hair orientation is predominantly pointing towards the camera, regardless of the face rotation. When rotating the object, the eyes and hair in our generated images do not stay fixed but are adjusted to meet the dataset bias.

Object Transformation Distributions: We sometimes observe disentanglement failures, e.g. for Churches where the background contains a church, or for CompCars where the foreground contains background elements (see Sup. Mat.). We attribute these to mismatches between the assumed uniform distributions over camera poses and object-level transformations and their real distributions.

Conclusion

We present GIRAFFE, a novel method for controllable image synthesis. Our key idea is to incorporate a compositional 3D scene representation into the generative model. By representing scenes as compositional generative neural feature fields, we disentangle individual objects from the background as well as their shape and appearance without explicit supervision. Combining this with a neural renderer yields fast and controllable image synthesis. In the future, we plan to investigate how the distributions over object-level transformations and camera poses can be learned from data. Further, incorporating supervision which is easy to obtain, e.g. predicted object masks, is a promising approach to scale to more complex, multi-object scenes.

Acknowledgement

This work was supported by an NVIDIA research gift. We thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting MN. AG was supported by the ERC Starting Grant LEGO-3D (850533) and DFG EXC number 2064/1 - project number 390727645.

References