BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled Images

Thu Nguyen-Phuoc, Christian Richardt, Long Mai, Yong-Liang Yang, Niloy Mitra

Introduction

The computer graphics pipeline has achieved impressive results in generating high-quality images, while offering users a great level of freedom and controllability over the generated images. This has many applications in creating and editing content for the creative industries, such as films, games, scientific visualisation, and more recently, in generating training data for computer vision tasks. However, the current pipeline, ranging from generating 3D geometry and textures, rendering, compositing and image post-processing, can be very expensive in terms of labour, time, and costs.

Recent image generative models, in particular generative adversarial networks [GANs; 16], have greatly improved the visual fidelity and resolution of generated images . Conditional GANs allow users to manipulate images, but require labels during training. Recent work on unsupervised disentangled representations using GANs relaxes this need for labels. The ability to produce high-quality, controllable images has made GANs an increasingly attractive alternative to the traditional graphics pipeline for content generation. However, most work focuses on property disentanglement, such as shape, pose and appearance, without considering the compositionality of the images, i.e., scenes being made up of multiple objects. Therefore, they do not offer control over individual objects in a way that respects the interaction of objects, such as consistent lighting and shadows. This is a major limitation of current image generative models, compared to the graphics pipeline, where 3D objects are modelled individually in terms of geometry and appearance, and combined into 3D scenes with consistent lighting.

Even when considering object compositionality, most approaches treat objects as 2D layers combined using alpha compositing . Moreover, they also assume that each object’s appearance is independent . While this layering approach has led to good results in terms of object separation and visual fidelity, it is fundamentally limited by the choice of 2D representation. Firstly, it is hard to manipulate properties that require 3D understanding, such as pose or perspective. Secondly, object layers tend to bake in appearance and cannot adequately represent view-specific appearance, such as shadows or material highlights changing as objects move around in the scene. Finally, it is non-trivial to model the appearance interactions between objects, such as scene lighting that affects objects’ shadows on a background.

We introduce BlockGAN, a generative adversarial network that learns 3D object-oriented scene representations directly from unlabelled 2D images. Instead of learning 2D layers of objects and combining them with alpha compositing, BlockGAN learns to generate 3D object features and to combine them into deep 3D scene features that are projected and rendered as 2D images. This process closely resembles the computer graphics pipeline where scenes are modelled in 3D, enabling reasoning over occlusion and interaction between object appearance, such as shadows or highlights. During test time, each object’s pose can be manipulated using 3D transforms directly applied to the object’s deep 3D features. We can also add new objects and remove existing objects in the generated image by changing the number of 3D object features in the 3D scene features at inference time. This shows that BlockGAN has learnt a non-trivial representation of objects and their interaction, instead of merely memorizing images.

BlockGAN is trained end-to-end in an unsupervised manner directly from unlabelled 2D images, without any multi-view images, paired images, pose labels, or 3D shapes. We experiment with BlockGAN on a variety of synthetic and natural image datasets. In summary, our main contributions are:

BlockGAN, an unsupervised image generative model that learns an object-aware 3D scene representation directly from unlabelled 2D images, disentangling both between objects and individual object properties (pose and identity);

showing that BlockGAN can learn to separate objects even from cluttered backgrounds; and

demonstrating that BlockGAN’s object features can be added, removed and manipulated to create novel scenes that are not observed during training.

Related work

Unsupervised GANs learn to map samples from a latent distribution to data categorised as real by a discriminator network. Conditional GANs enable control over the generated image content, but require labels during training. Recent work on unsupervised disentangled representation learning using GANs provides controllability over the final images without the need for labels. Loss functions can be designed to maximize mutual information between generated images and latent variables . However, these models do not guarantee which factors can be learnt, and have limited success when applied to natural images. Network architectures can play a vital role in both improving training stability and controllability of generated images . We also focus on designing an appropriate architecture to learn object-level disentangled representations. We show that injecting inductive biases about how the 3D world is composed of 3D objects enables BlockGAN to learn 3D object-aware scene representations directly from 2D images, thus providing control over both 3D pose and appearance of individual objects.

D-aware neural image synthesis.

Introducing 3D structures into neural networks can improve the quality and controllability of the image generation process . This can be achieved with explicit 3D representations, like appearance flow , occupancy voxel grids , meshes, or shape templates , in conjunction with handcrafted differentiable renderers . Renderable deep 3D representations can also be learnt directly from images . HoloGAN further shows that adding inductive biases about the 3D structure of the world enables unsupervised disentangled feature learning between shape, appearance and pose. However, these learnt representations are either object-centric (i.e., no background), or treat the whole scene as one object. Thus, they do not consider scene compositionality, i.e., components that can move independently. BlockGAN, in contrast, is designed to learn object-aware 3D representations that are combined into a unified 3D scene representation.

Object-aware image synthesis.

Recent methods decompose image synthesis into generating components like layers or image patches, and combining them into the final image . This includes conditional GANs that use segmentation masks , scene graphs , object labels, key points or bounding boxes , which have shown impressive results for natural image datasets. Recently, unsupervised methods learned object disentanglement for multi-object scenes on simpler synthetic datasets (single-colour objects, simple lighting, and material). Other approaches successfully separate foreground from background objects in natural images, but make strong assumptions about the size of objects or independent object appearance . These methods treat object components as image patches or 2D layers with corresponding masks, which are combined via alpha compositing at the pixel level to generate the final stage. The work closest to ours learns to generate multiple 3D primitives (cuboids, spheres and point clouds), renders them into separate 2D layers with a handcrafted differentiable renderer, and alpha-composes them based on their depth ordering to create the final image . Despite the explicit 3D geometry, this method does not handle cluttered backgrounds and requires extra supervision in the shape of labelled images with and without foreground objects.

BlockGAN takes a different approach. We treat objects as learnt 3D features with corresponding 3D poses, and learn to combine them into 3D scene features. Not only does this provide control over 3D pose, but also enables learning of realistic lighting and shadows. Our approach allows adding more objects into the 3D scene features to generate images with multiple objects, which are not observed at training time.

Method

Inspired by the computer graphics pipeline, we assume that each image x\mathbf{x} is a rendered 2D image of a 3D scene composed of KK 3D foreground objects {O1,…,OK}\{O_{1},\ldots,O_{K}\} in addition to the background O0O_{0}:

where the function ff combines multiple objects into unified scene features that are projected to the image x\mathbf{x} by pp. We assume each object OiO_{i} is defined in a canonical orientation and generated from a noise vector zi\mathbf{z}_{i} by a function gig_{i} before being individually posed using parameters θi\boldsymbol{\theta}_{i}: Oi=gi(zi,θi)O_{i}=g_{i}(\mathbf{z}_{i},\boldsymbol{\theta}_{i}).

We inject the inductive bias of compositionality of the 3D world into BlockGAN in two ways. (1) The generator is designed to first generate 3D features for each object independently, before transforming and combining them into unified scene features, in which objects interact. (2) Unlike other methods that use 2D image patches or layers to represent objects, BlockGAN directly learns from unlabelled images how to generate objects as 3D features. This allows our model to disentangle the scene into separate 3D objects and allows the generator to reason over 3D space, enabling object pose manipulation and appearance interaction between objects. BlockGAN, therefore, learns to both generate and render the scene features into images that can fool the discriminator.

Figure 1 illustrates the BlockGAN generator architecture. Each noise vector zi\mathbf{z}_{i} is mapped to 3D object features OiO_{i}. Objects are then transformed according to their pose θi\boldsymbol{\theta}_{i} using a 3D similarity transform, before being combined into 3D scene features using the scene composer ff. The scene features are transformed into the camera coordinate system before being projected to 2D features to render the final images using the camera projector function pp. During training, we randomly sample both the noise vectors zi\mathbf{z}_{i} and poses θi\boldsymbol{\theta}_{i}. During test time, objects can be generated with a given identity zi\mathbf{z}_{i} in the desired pose θi\boldsymbol{\theta}_{i}.

BlockGAN is trained end-to-end using only unlabelled 2D images, without the need for any labels, such as poses, 3D shapes, multi-view inputs, masks, or geometry priors like shape templates, symmetry or smoothness terms. We next explain each component of the generator in more detail.

To generate 3D object features, BlockGAN implements the style-based strategy, which helps to disentangle between pose and identity while improving training stability . As illustrated in Figure 2, the noise vector zi\mathbf{z}_{i} is mapped to affine parameters – the “style controller” – for adaptive instance normalization [AdaIN; 21] after each 3D convolution layer. However, unlike HoloGAN , which learns 3D features directly for the whole scene, BlockGAN learns 3D features for each object, which are then transformed to their target poses using similarity transforms, and combined into 3D scene features. We implement these 3D similarity transforms by trilinear resampling of the 3D features according to the translation, rotation and scale parameters θi\boldsymbol{\theta}_{i}; samples falling outside the feature tensor are clamped to zero. This allows BlockGAN to not only separate object pose from identity, but also to disentangle multiple objects in the same scene.

2. Scene composer function

3. Learning to render

The computer graphics pipeline implements perspective projection using a projective transform that converts objects from world coordinates (our scene space) to camera coordinates . We implement this camera transform like the similarity transforms used to manipulate objects in Section 3.1, by resampling the 3D scene features according to the viewing volume (frustum) of the virtual perspective camera (see Figure 3). For correct perspective projection, this transform must be a projective transform, the superset of similarity transforms . Specifically, the viewing frustum, in scene space, can be defined relative to the camera’s pose θcam\boldsymbol{\theta}_{\text{cam}} using the angle of view, and the distance of the near and far planes. The camera-space features are a new 3D tensor of features, of size Hc ⁣× ⁣Wc ⁣× ⁣Dc ⁣× ⁣CsH_{c}\!\times\!W_{c}\!\times\!D_{c}\!\times\!C_{s}, whose corners are mapped to the corners of the camera’s viewing frustum using the unique projective 3D transform computed from the coordinates of corresponding corners using the direct linear transform .

In practice, we combine the object and camera transforms into a single transform by multiplying both transform matrices and resampling the object features in a single step, directly from object to camera space. This is computationally more efficient than resampling twice, and advantageous from a sampling theory point of view, as the features are only interpolated once, not twice, and thus less information is lost by the resampling. The combined transform is a fixed, differentiable function with parameters (θi,θcam)(\boldsymbol{\theta}_{i},\boldsymbol{\theta}_{\text{cam}}). The individual objects are then combined in camera space before the final projection.

4. Loss functions

We train BlockGAN adversarially using the non-saturating GAN loss . For natural images with cluttered backgrounds, we also add a style discriminator loss . In addition to classifying the images as real or fake, this discriminator also looks at images at the feature level. Given image features Φl\mathbf{\Phi}_{l} at layer ll, the style discriminator classifies the mean μ(Φl)\boldsymbol{\mu}(\mathbf{\Phi}_{l}) and standard deviation σ(Φl)\boldsymbol{\sigma}(\mathbf{\Phi}_{l}) over the spatial dimensions, which describe the image “style” . This more powerful discriminator discourages the foreground generator to include parts of the background within the foreground object(s). We provide detailed network and loss definitions in the supplemental material.

Experiments

We train BlockGAN on images at 64×\times64 pixels, with increasing complexity in terms of number of foreground objects (1–4) and texture (synthetic images with simple shapes and simple to natural images with complex texture and cluttered background). These datasets include the synthetic CLEVRnn , Synth-Carnn and Synth-Chairnn, and the real Real-Car , where nn is the number of foreground objects. Additional details and results are included in the supplementary material.

Implementation details.

We assume a fixed and known number of objects of the same type. Fore- and background generators have similar architectures and the same number of output channels, but foreground generators have twice as many channels in the learnt constant tensor. Since foreground objects are smaller than the background, we set scale=1 for the background object, and randomly sample scales <<1 for foreground objects. Please see our supplemental material for more implementation details and an ablation experiment. We make our code publicly available at github.com/thunguyenphuoc/BlockGAN.

1. Qualitative results

Despite being trained with only unlabelled images, Figure 4 shows that BlockGAN learns to disentangle different objects within a scene: foreground from background, and between multiple foreground objects. More importantly, BlockGAN also provides explicit control and enables smooth manipulation of each object’s pose θi\boldsymbol{\theta}_{i} and identity zi\mathbf{z}_{i}. Figure 6 shows results on natural images with a cluttered background, where BlockGAN is still able to separate objects and enables 3D object-centric modifications. Since BlockGAN combines deep object features into scene features, changes in an object’s properties also influence its shadows, and highlights adapt to the object’s movement. These effects can be better observed in the supplementary animations.

2. Quantitative results

We evaluate the visual fidelity of BlockGAN’s results using Kernel Inception Distance [KID; 5], which has an unbiased estimator and works even for a small number of images. Note that KID does not measure the quality of object disentanglement, which is the main contribution of BlockGAN. We first compare with a vanilla GAN [WGAN-GP; 17] using a publicly available implementationhttps://github.com/LynnHo/DCGAN-LSGAN-WGAN-WGAN-GP-Tensorflow​​. Secondly, we compare with LR-GAN , a 2D-based method that learns to generate image background and foregrounds separately and recursively. Finally, we compare with HoloGAN, which learns 3D scene representations that separate camera pose and identity, but does not consider object disentanglement. For LR-GAN and HoloGAN, we use the authors’ code. We tune hyperparameters and then compute the KID for 10,000 images generated by each model (samples by all methods are included in the supplementary material). Table 1 shows that BlockGAN generates images with competitive or better visual fidelity than other methods.

3. Scene manipulation beyond training data

We show that at test time, 3D object features learnt by BlockGAN can be realistically manipulated in ways that have not been observed during training time. First, we show that the learnt 3D object features can also be reused to add more objects to the scene at test time, thanks to the compositionality inductive bias and our choice of scene composer function. Firstly, we use BlockGAN trained on datasets with only one foreground object and one background, and show that more foreground objects of the same category can be added to the same scene at test time. Figure 6 shows that 2–4 new objects are added and manipulated just like the original objects while maintaining

realistic shadows and highlights. In Figure 5, we use BlockGAN trained on CLEVR4 and then remove (top) and add (bottom) more objects to the scene. Note how BlockGAN generates realistic shadows and occlusion for scenes that the model has never seen before.

Secondly, we apply spatial manipulations that were not part of the similarity transform used during training, such as horizontal stretching, or slicing and combining different foreground objects. Figure 6 shows that object features can be geometrically modified intuitively, without needing explicit 3D geometry or multi-view supervision during training.

4. Comparison to 2D-based LR-GAN

LR-GAN first generates a 2D background layer, and then generates and combines foreground layers with the generated background using alpha-compositing. Both BlockGAN and LR-GAN show the importance of combining objects in a contextually relevant manner to generate visually realistic images (see Table 1). However, LR-GAN does not offer explicit control over object location. More importantly, LR-GAN learns an entangled representation of the scene: sampling a different background noise vector also changes the foreground (Figure 7). Finally, unlike BlockGAN, LR-GAN does not allow adding more foreground objects during test time. This demonstrates the benefits of learning disentangled 3D object features compared to a 2D-based approach.

5. Ablation study: Non-uniform pose distribution

For the natural Real-Car dataset, we observe that BlockGAN has difficulties learning the full 360° rotation of the car, even though fore- and background are disentangled well. We hypothesise that this is due to the mismatch between the true (unknown) pose distribution of the car, and the uniform pose distribution we assume during training. To test this, we create a synthetic dataset similar to Synth-Car1 with a limited range of rotation, and train BlockGAN with a uniform pose distribution. To generate the imbalanced rotation dataset, we sample the rotation uniformly from the front/left/back/right viewing directions ±15\pm 15°. In other words, the car is only seen from the front/left/back/right 30°, respectively, and there are four evenly spaced gaps of 60° that are never observed, for example views from the front-right. With the imbalanced dataset, Figure 8 (bottom) shows correct disentangling of foreground and background. However, rotation of the car only produces images with (near-)frontal views (top), while depth translation results in cars that are randomly rotated sideways (middle). We observe similar behaviour for the natural Real-Car dataset. This suggests that learning object disentanglement and full 3D pose rotation might be two independent problems. While assuming a uniform pose distribution already enables good object disentanglement, learning the pose distribution from the training data would likely improve the quality of 3D transforms.

In our supplemental material, we include comparisons to HoloGAN as well as additional ablation studies on comparing different scene composer functions, using a perspective camera versus a weak-perspective camera, adopting the style discriminator for scenes with cluttered backgrounds, and training on images with an incorrect number of objects.

Discussion and Future Work

We introduced BlockGAN, an image generative model that learns 3D object-aware scene representations from unlabelled images. We show that BlockGAN can learn a disentangled scene representation both in terms of objects and their properties, which allows geometric manipulations not observed during training. Most excitingly, even when BlockGAN is trained with fewer or even single objects, additional 3D object features can be added to the scene features at test time to create novel scenes with multiple objects. In addition to computer graphics applications, this opens up exciting possibilities, such as combining BlockGAN with models like BiGAN or ALI to learn powerful object representations for scene understanding and reasoning.

Future work can adopt more powerful relational learning models to learn more complex object interactions such as inter-object shadowing or reflections. Currently, we assume prior knowledge of object category and the number of objects for training. We also assume object poses are uniformly distributed and independent from each other. Therefore, the ability to learn this information directly from training images would allow BlockGAN to be applied to more complex datasets with a varying number of objects and different object categories, such as COCO or LSUN .

Acknowledgments and Disclosure of Funding

We received support from the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement No. 665992, the EPSRC Centre for Doctoral Training in Digital Entertainment (EP/L016540/1), RCUK grant CAMERA (EP/M023281/1), an EPSRC-UKRI Innovation Fellowship (EP/S001050/1), and an NVIDIA Corporation GPU Grant. We received a gift from Adobe.

Broader Impact

BlockGAN is an image generative model that learns an object-oriented 3D scene representation directly from unlabelled 2D images. Our approach is a new machine learning technique that makes it possible to generate unseen images from a noise vector, with unprecedented control over the identity and pose of multiple independent objects as well as the background. In the long term, our approach could enable powerful tools for digital artists that facilitate artistic control over realistic procedurally generated digital content. However, any tool can in principle be abused, for example by adding new, manipulating or removing existing objects or people from images.

At training time, our network performs a task somewhat akin to scene understanding, as our approach learns to disentangle between multiple objects and individual object properties (specifically their pose and identity). At test time, our approach enables sampling new images with control over pose and identity for each object in the scene, but does not directly take any image input. However, it is possible to embed images into the latent space of generative models . A highly realistic generative image model and a good image fit would then make it possible to approximate the input image and, more importantly, to edit the individual objects in a pictured scene. Similar to existing image editing software, this enables the creation of image manipulations that could be used for ill-intended misinformation (fake news), but also for a wide range of creative and other positive applications. We expect the benefits of positive applications to clearly outweigh the potential downsides of malicious applications.

References

Appendix A Additional results

We compare BlockGAN with HoloGAN , which also learns deep 3D scene features but does not consider object disentanglement. In particular, HoloGAN only considers one noise vector z\mathbf{z} for identity and one pose θ\boldsymbol{\theta} for the entire scene, and does not consider translation t\mathbf{t} as part of θ\boldsymbol{\theta}. While HoloGAN works well with object-centred scenes, it struggles with moving foreground objects. Figure 9 shows that HoloGAN tends to associate each pose θ\boldsymbol{\theta} with a fixed object’s identity (i.e., moving objects erroneously changes identity of both foreground and background), while changing z\mathbf{z} only changes a small part of the background. BlockGAN, on the other hand, can separate identity and pose for each object, while being able to learn scene-level effects such as lighting and shadows.

A.2. Increasing the number of foreground objects

To test the capability of our method, we train BlockGAN on CLEVR6 (6 foreground objects). As shown in Figure 10, BlockGAN is still capable of generating and manipulating 3D object features, although the background generator now also produces foreground objects (Figure 10f). Moreover, rotating individual object leads to changes in object’s depth (Figure 10a).

Interestingly, we notice that BlockGAN now generates images with more or less than 6 objects, despite being trained with images that contain exactly 6 objects (Figure 10g). We hypothesise that BlockGAN’s failure in this case is due to our assumption that the poses of all objects are independent from each other (during training, we randomly sample the pose θ\theta for each object). This is not true in the physical world (also in the CLEVR dataset) where objects do not intersect. The more objects there are in the scene, the stronger the interdependence between objects’ poses becomes. Therefore, for future work, we hope to adopt more powerful relation learning structures to learn objects’ pose directly from training images. Another interesting direction is to design object-aware discriminators, which are capable of recognising fake images when the generators produce samples with more objects than the training images.

Appendix B Additional ablation studies

Here we show the advantage of implementing the perspective camera explicitly, compared to using a weak-perspective projection like HoloGAN . Since a perspective camera directly affects foreshortening, it provides strong cues for BlockGAN to solve the scale/depth ambiguity. This is especially important for BlockGAN to learn to project and reason over occlusion by concatenating the depth and channel dimension, followed by an MLP. Since the MLP is flexible, BlockGAN trained without a perspective camera, therefore, tends to learn to associate an object’s identity with scale and depth, while changing depth only changes the object’s appearance (see Figure 11).

B.2. Scene composer function

We consider and compare three scene composer functions: (i) element-wise summation, (ii) element-wise maximum, and (iii) an MLP (multi-layer perceptron). We train BlockGAN with each function and compare their performance in terms of visual quality (KID score) in Table 2. While all three functions can successfully combine objects into a scene, the element-wise maximum performs best and easily generalises to multiple objects. Therefore, we use the element-wise maximum for BlockGAN.

B.3. Learning without the style discriminator

When BlockGAN is trained with a standard discriminator on datasets with a cluttered background, such as the Real-Car dataset, the foreground object features tend to include part of the background object. This creates visual artefacts when objects move in the scene (indicated by red arrows in Figure 12a). We hypothesise that these artefacts should be picked up by the discriminator since generated images should look unrealistic. Therefore, we add more powerful style discriminators to the original discriminator at different layers (see Section D for details). Figure 12b shows that the generator is indeed discouraged from adding background information to the foreground object features, leading to cleaner results.

B.4. Incorrect number of objects

We next investigate the performance of BlockGAN when the training data contains fewer or more objects than expected. In Figure 13, we show BlockGAN configured with 2 foreground object generators when trained with images containing 1 or 3 foreground objects. If only a single object is present (Figure 13, left), changing either of the two foreground generators changes the object’s appearance and pose (top), while changing the background works as expected (bottom). If there are three objects present (Figure 13, right), changing one foreground generator changes one object as expected (top), while changing the background generator simultaneously changes one foreground object and the background (bottom).

Appendix C Comparison to other methods

In Figure 14, 15, 16 and 17, we show generated samples by a vanilla GAN (WGAN-GP ), 2D object-aware LR-GAN , 3D-aware HoloGAN and our BlockGAN. Compared to other models, BlockGAN produces samples with competitive or better quality, and offers explicit control over the poses of objects in the generated images. Notice that although LR-GAN is designed to handle foreground and background objects explicitly, for CLEVR2 with two foreground objects, this method struggles and tends to always place one foreground object at the image centre (see Figure 16).

For WGAN-GP, we use a publicly available implementationhttps://github.com/LynnHo/DCGAN-LSGAN-WGAN-WGAN-GP-Tensorflow​​. For LR-GAN and HoloGAN, we use the code provided by the authors. We conduct hyperparameter search for these models, and report best results for each method. Note that for HoloGAN, we modify the 3D transformation to add translation during training, since this method assumes that foreground objects are at the image centre.

Appendix D Loss function and style discriminator

For datasets with cluttered backgrounds like the natural Real-Car dataset, we adopt style discriminators in addition to the normal image discriminator (see the benefit in Figure 12). Style discriminators perform the same real/fake classification task as the standard image discriminator, but at the feature level across different layers. In particular, style discriminators classify the mean μ\boldsymbol{\mu} and standard deviation σ\boldsymbol{\sigma} of the features Φl\mathbf{\Phi}_{l} at different levels ll (which are believed to describe the image “style”). The mean μ(Φl(x))\boldsymbol{\mu}(\mathbf{\Phi}_{l}(\mathbf{x})) and variance σ(Φl(x))\boldsymbol{\sigma}(\mathbf{\Phi}_{l}(\mathbf{x})) of the features Φl(x)\mathbf{\Phi}_{l}(\mathbf{x}) are computed across batch and spatial dimensions independently using:

The style discriminators are implemented as MLPs with sigmoid activation functions for binary classification. A style discriminator at layer ll is written as

The total loss therefore can be written as

We set λs=1\lambda_{\text{s}}=1 for all natural datasets and λs=0\lambda_{\text{s}}=0 for synthetic datasets.

Appendix E Datasets

We modify the CLEVR dataset to add a larger variety of colours and primitive shapes. Additionally, we use the scene setups provided by CLEVR to render the remaining synthetic datasets (Synth-Carnn and Synth-Chairnn, with nn foreground objects each). These include a fixed, grey background, a virtual camera with fixed parameters but random location jittering, and random lighting. We also use the render script from CLEVR to randomly place foreground objects into the scene and render them. We render all image at resolution 128 ×\times 128, and bi-linearly downsample them to 64 ×\times 64 for training. For the natural Car dataset, each image is first scaled such that the smaller side is 64, then it is cropped to produce a 64×\times64 pixel crop. During training, we randomly move the 64×\times64 cropping window before cropping the image. Figure 18 includes samples from our generated datasets, and Table 3 lists the range of pose parameters used for each dataset during training.

Link for 3D textured chair models: https://keunhong.com/publications/photoshape/

Link for CLEVR: https://github.com/facebookresearch/clevr-dataset-gen

Link for natural Car dataset: http://mmlab.ie.cuhk.edu.hk/datasets/comp_cars/

Appendix F Implementation

We assume a virtual camera with a focal length of 35 mm and a sensor size of 32 mm (Blender’s default values), which corresponds to an angle of view of 2arctan⁡32 mm2×35 mm ⁣= ⁣49.12\arctan\frac{32\,\text{mm}}{2\times 35\,\text{mm}}\!=\!49.1 degrees (we use the same setup for natural images).

Sampling.

We initialise all weights using N(0,0.2)\mathcal{N}(0,0.2) and biases as . For CLEVRnn, we use noise vector dimensions of ∣z0∣=20\left\lvert\mathbf{z}_{0}\right\rvert=20 for the background, and ∣zi∣=60\left\lvert\mathbf{z}_{i}\right\rvert=60 (for i=1,…,ni=1,\ldots,n) for the foreground objects, to account for their relative visual complexity. Similarly, for Synth-Carnn and Synth-Chairnn, we use ∣z0∣=30\left\lvert\mathbf{z}_{0}\right\rvert=30 and ∣zi∣=90\left\lvert\mathbf{z}_{i}\right\rvert=90 (for i=1,…,ni=1,\ldots,n), to account for their relative visual complexity. For the natural Real-Car dataset, we use ∣z0∣=100\left\lvert\mathbf{z}_{0}\right\rvert=100 and ∣z1∣=200\left\lvert\mathbf{z}_{1}\right\rvert=200. Note that we only feed z\mathbf{z} to the 3D features of each object, and not to the 3D scene features and 2D features. Table 3 provides the ranges we use for sampling the pose θi\boldsymbol{\theta}_{i} of foreground objects during training.

Training.

We train BlockGAN using the Adam optimiser , with β1=0.5\beta_{1}=0.5 and β2=0.999\beta_{2}=0.999. We use the same learning rate for both the discriminator and the generator. Empirically, we find that updating the generator twice for every update of the discriminator achieves images with the best visual fidelity. We use a learning rate of 0.0001 for all synthetic datasets. For the natural Cars dataset, we use a learning rate of 0.00005. We train all datasets with a batch size of 64 for 50 epochs. Training takes 1.5 days for the synthetic datasets and 3 days for the natural Real-Cars dataset.

Infrastructure.

All models were trained using a single GeForce RTX 2080 GPU.

F.2. Network architecture

We describe the network architecture for the BlockGAN foreground object generator in Table 5, the BlockGAN background generator in Table 5, and the overall BlockGAN generator in Tables 8 and 8 for synthetic and real datasets, respectively. Note that we use ReLU for the synthetic datasets and LReLU for the natural Car dataset after the AdaIN layer. The discriminator is described in Table 8.

In terms of the notation in Section 3 of the main paper, object features have dimensions Ho×Wo×Do×Co=16×16×16×64H_{o}\times W_{o}\times D_{o}\times C_{o}=16\times 16\times 16\times 64, scene features have the same dimensions Hs×Ws×Ds×Cs=16×16×16×64H_{s}\times W_{s}\times D_{s}\times C_{s}=16\times 16\times 16\times 64, and camera features have dimensions Hc×Wc=16×16H_{c}\times W_{c}=16\times 16 (before up-convolutions to 64×6464\times 64) with Cc=64C_{c}=64 channels for synthetic datasets and Cc=256C_{c}=256 channels for natural image datasets.

As GANs empirically tend to perform better on category-specific datasets, we decided to start with this assumption. A promising future direction is to adopt a shared rendering layer for objects generated by different category-specific generators, similar to Aliev et al. .