RELATE: Physically Plausible Multi-Object Scene Synthesis Using Structured Latent Spaces

Sebastien Ehrhardt, Oliver Groth, Aron Monszpart, Martin Engelcke, Ingmar Posner, Niloy Mitra, Andrea Vedaldi

Introduction

We consider the problem of learning to generate plausible images of scenes starting from parameters that are physically interpretable. Furthermore, we wish to learn such a capability from raw images alone, without any manual or external supervision. Image generation is often approached via Generative Adversarial Networks (GAN) . These models learn to map noise vectors, used as a source of randomness, to image samples. While the resulting images are realistic, the random vectors that parameterize them are not interpretable. To address this issue, authors have recently proposed to structure the latent space of deep generative models, giving it a partial physical interpretability . For example, HoloGAN samples volumes and cameras to generate 2D images of 3D objects, and BlockGAN creates scenes by composing multiple objects. The resulting GANs have been shown to learn concepts such as viewpoint and object disentangling from raw images.

BlockGAN is of particular interest because, via its relatively strong architectural biases, it provides interpretable parameters for the scene, incorporating concepts such as position and orientation. However, BlockGAN comes with a significant limitation in that it assumes that objects are mutually independent. This approximation is acceptable only when objects interact weakly, but it is badly violated for medium to densely packed scenes, or for scenes such as stacking wooden blocks or cars following a path, where the (object) correlation is strong.

Recent work in object-centric generative modeling has attempted to specifically address this by capturing correlations in latent space (e.g., ). However as object state information remains significantly entangled in these models they have, to date, been unable to operate on real-world data.

In this paper, we introduce RELATE, a model which explicitly leverages the strong architectural biases of BlockGAN to effectively model correlations between latent object state variables. This leads to a powerful model class, which is able to capture complex physical interactions, while still being able to learn from raw visual inputs alone. Empirically, we show that only when we model such interactions our GAN model correctly disentangles different objects when they exhibit even a moderate amount of correlation (figures 2 and 3). Without this component, the model may still generate high fidelity images, but it generally fails to establish a physically-plausible association between the parameters and the generated images. Our results also demonstrate that GANs are surprisingly sensitive to the correlation of objects in natural scenes, and can thus be used to directly learn these without resorting to techniques such as variational auto-encoding (VAE ).

We demonstrate the efficacy of RELATE in several scenarios, including balls rolling in bowls of variable shape , cluttered tabletops (CLEVR ), block stacking (ShapeStacks ), and videos of traffic at busy intersection. By ablating the interaction module, we show that modeling the spatial correlation between the objects is key. Furthermore, we compare RELATE to several recent GAN- and VAE-based baselines, including BlockGAN , GENESIS and OCF , in terms of Fréchet Inception Distance (FID) , and outperform even the best state-of-the-art model by up to 29 points.

Qualitatively, we show that modeling spatial relationships strongly affects scene decomposition and the enforcement of spatial constraints in the generated images. We also show that the physically interpretable latent space learned by RELATE can be used to edit scenes as well as to generate scenes outside the distribution of the training data (e.g., containing more or fewer objects). Finally, we show that the parameterization can be used to generate long plausible video sequences (as measured according to FVD score ) by simulating their dynamics while preserving their spatial consistency.

Related Work

Inspired by the analysis-by-synthesis approach for visual perception discussed in cognitive science , recent work propose structured latent space models to explain and synthesize images as sets of constituent components which are individually represented using VAEs or GANs . Other approaches favor explicit symbolic representations over distributed ones when parsing an image or propose probabilistic programming languages to formalize image generation . In both cases, object-centric modeling allows decomposition of images into components and also enables targeted image modification via interpolation in symbol or latent space, e.g., altering position or color of an object parsed by the model. Interpretable and controllable factors of image generation are desirable properties for neural rendering models and have been investigated in recent image generation models, e.g., . However, despite the modeling effort put into the object representations, inter-object interactions are typically only modeled in a less explicit way, e.g., via image layers , depth ordering variables or an autoregressive scene prior . Our work adds to this line of work by proposing a spatial correlation network which facilitates disentanglement of learned object representations and can be trained from raw observations.

Neural Physics Approximation.

Harnessing the power of deep learning to approximate physical processes is an emerging trend in the machine learning community. Especially the approximation of rigid body dynamics with neural networks already boasts a large body of literature, e.g., . Such learned approximations of object interactions have been successfully employed in object manipulation and tracking . However, most entries in this line of work are only applied to visual toy domains or rely on segmentation masks or bounding boxes to initialize their object representations before the networks approximate the object dynamics. While we leverage ideas from neural dynamics modeling, we go beyond the established scope of visual toy domains such as colored point masses or moving MNIST digits and learn directly from rich visual data such as simulated object stacks and real traffic videos without further annotation.

Extraction and Generation of Video Dynamics.

Videos are a natural choice of data source to learn about the physics of rigid bodies. Physical information extracted from videos can either be explicit such as estimates of velocity or friction or implicitly represented in the latent space, e.g., sub-spaces corresponding to pose variation . More recently, object-centric approaches have also been leveraged to acquire better video representations for future frame prediction or model-based reinforcement learning . Several studies have also attempted to learn entire video distributions as spatio-temporal tensors in GAN frameworks yielding impressive first results for full video generation in artificial and real domains. In contrast to prior art, our model departs from a monolithic spatio-temporal tensor representation over an entire video. Instead we cast the video learning and generation process as temporal extension of the object-centric representation of a single frame, lowering the computational burden while still faithfully representing long-range dynamics.

Method

RELATE (figure 1) consists of two main components: An interaction module, which computes physically plausible inter-object and object-background relationships, and a scene composition and rendering module, which features an interpretable parameter space factored into appearance and position vectors. The details are given next.

RELATE considers scenes containing up to KK distinct objects. The model starts by sampling appearance parameters z1,…,zK∼U(Nf)z_{1},\dots,z_{K}\sim\mathcal{U}(^{N_{f}}) for each individual foreground object as well as a parameter z0∼U(Nb)z_{0}\sim\mathcal{U}(^{N_{b}}) for the background. These parameters are small noise vectors, similar to the ones typically used in generative networks. Different from the object poses below, they are sampled independently, thus assuming that the appearance of different objects is independent.

This model is ‘physically interpretable’ in the sense that it captures (1) the identities of KK distinct objects and (2) their pose parameters as translation vectors. This should be contrasted to traditional GAN models, where the code space is given as an uninterpretable, monolithic noise vector zz. Despite the structure given to the code space, there is no guarantee that the model will actually learn to map it to the corresponding structure in the example images. However, we found empirically that this is the case as long as the correlations between the different objects are also captured.

2 Modeling correlations in scene composition

RELATE departs significantly from prior art such as BlockGAN as it does not assume the parameters θi\theta_{i} of the different objects to be independent. In order to model correlation, we propose a two-step procedure, based on a residual sampler. First, we sample a vector of KK i.i.d. poses Θ^∼U([−H′′/2,H′′/2]2K)\hat{\Theta}\sim\mathcal{U}([-H^{\prime\prime}/2,H^{\prime\prime}/2]^{2K}) where H′′<HH^{\prime\prime}<H is smaller than the spatial size HH of the tensor encoding. Then, we pass this vector to a ‘correction’ network Γ\Gamma that remaps the initial configuration to one that accounts for the correlation between object locations and appearances, as well as between objects and the background (coded by the appearance component z0z_{0} in zz): Θ:=Γ(Θ^,Z).\Theta:=\Gamma(\hat{\Theta},Z). In practice, we expect object interactions, as any physical law, to be symmetric with respect to the order of the objects. We obtain this effect by implementing Γ\Gamma as running KK copies of the same corrective function in parallel:

The function ζ\zeta is implemented in a manner similar to the Neural Physics Engine (NPE) :

where ff and gg are Multi Layer Perceptrons (MLPs) (tables A8, A9) operating on stacked vector inputs and hsh^{s} is an embedding capturing the interactions between the KK objects. Besides symmetry, an advantage of this scheme is that it can take an arbitrary number of objects KK due to the sum-pooling operator used to capture the interactions. In this manner, the sampler Γ\Gamma is automatically defined for any value of KK. For each scene, K is sampled uniformly from a fixed interval [Kmin,Kmax][K_{\text{min}},K_{\text{max}}]. Furthermore, sampling independent quantities followed by a correction has the benefit of injecting some variance on the objects positions at the early stage of training, which helps to avoid converging to trivial/bad solutions.

An advantage of RELATE is that it can be easily modified to take advantage of additional structure in the scene. For scenes where objects have natural order, such as stacks of blocks, we experiment with conditioning pose θi\theta_{i} on the preceding pose θi−1\theta_{i-1}, using a Markovian process. This is done by first sampling θ^1∼U([−H′′/2,H′′/2])\hat{\theta}_{1}\sim\mathcal{U}([-H^{\prime\prime}/2,H^{\prime\prime}/2]), and then applying a correction to account for the background z0z_{0} as before, finally sampling the other objects in sequence:

where f0,f1f_{0},f_{1} are implemented as MLPs as before (tables A11, A12). Note that this can be interpreted as a special case of the model above in the sense that we can write Θ:=Γ(Θ^,Z),\Theta:=\Gamma(\hat{\Theta},Z), provided that θ^k=0\hat{\theta}_{k}=0 for k≥2k\geq 2.

Modeling dynamics.

RELATE can also be immediately extended to make dynamic predictions. For this, we sample the initial positions θk(0)\theta_{k}(0) as before and then update them incrementally as θk(t+1)=θk(t)+vk(t+1),\theta_{k}(t+1)=\theta_{k}(t)+v_{k}(t+1), where vk(t)v_{k}(t) is the object velocity. In order to obtain the latter, we let Vk(t)=[vk(t−i)]i=2,1,0V_{k}(t)=[v_{k}(t-i)]_{i=2,1,0} denote the last three velocities of the kk-th object. The initial value Vk(0)=ev(zk,z0,θk(0))V_{k}(0)=e_{v}(z_{k},z_{0},\theta_{k}(0)) is initialized as a function of the appearance parameters and initial positions (table A10); and we use the NPE style update equations , where eve_{v}, fvf_{v} and gvg_{v} are MLPs,

3 Learning objective

Training our model makes use of a training set IiI_{i}, i=1,…,Ni=1,\dots,N of NN images of scenes containing different object configurations. No other supervision is required. Our learning objective is a sum of two high fidelity losses and a structural loss which we describe below.

For high fidelity, images I^\hat{I} generated by the model above are contrasted to real images II from the training set using the standard GAN discriminator LGAN(I^,I)\mathcal{L}_{\text{GAN}}(\hat{I},I) and style Lstyle(I^,I)\mathcal{L}_{\text{style}}(\hat{I},I) losses from (see section A3).

In addition, we introduce a regularizer to encourage the model to learn a non-trivial relationship between object positions and generated images. For this, we train a position regressor network PP that, given a generated image I^\hat{I}, predicts the location of the objects in it. In practice, we simplify this task and generate an image I^′\hat{I}^{\prime} by retaining only object kk of the KK objects at random and minimizing ∥θˇk−P(G(W(z0,zk,θk)))∥22\|\check{\theta}_{k}-P(G(W(z_{0},z_{k},\theta_{k})))\|^{2}_{2}. Here the symbol ⋅ˇ\check{\cdot} means that gradients are not back-propagated through θk\theta_{k}: this is to avoid mode collapse of the position at zero. PP shares most of its weights with the discriminator network (see table A13).

In the case of dynamic prediction, the discriminator takes as input the sequence of images concatenated along the RGB dimension and is tasked to discriminate between fake and real sequences. Similar to a static model we also have a position regressor which is tasked to predict the position of an object rendered at random with zero velocity.

Experiments

We learn mappings Ψb\Psi_{b} and Ψf\Psi_{f} using the same Adaptive Instance Normalization (AdaIN) architecture. The spatial size of their output tensors is set to H=16H=16 and the final output image to 128×128128\times 128 (which is reduced when needed for fair comparison to other methods). We use the Adam optimizer for learning and train for a fixed number of epochs and always select the last model snapshot. We consider two types of baselines: standard generative models such as DCGAN and DRAGAN, and object-centric generative baselines such as GENESIS and OCF , quoting results from the original papers whenever possible. In addition we also add BlockGAN2D as an ablation of our method.

Datasets.

We conduct experiments on four different datasets. First, we consider a relatively simple dataset, BallsInBowl , for assessing the model features and ablations. It consists of videos of two distinctly colored balls rolling in an elliptical bowl of variable orientation and eccentricity. Interactions amount to object collisions and the fact that they must roll within the bowl. To this, we add two popular synthetic datasets CLEVR (cluttered tabletops) and ShapeStacks (block stacking). Finally, we have collected a new dataset RealTraffic containing five hours of footage of a busy street intersection, divided into fragments containing from one to six cars. Especially the last dataset contains many interactions between the individual cars as they adapt their speed to the surrounding traffic which happens frequently when the light changes and cars either slow down because of a queue on red or accelerate when the lights change to green again. Further details about training, evaluation and datasets can be found in the appendix, sections A4, A5 and A6.

1 Generating static scenes

We start experimenting with the comparatively simple BallsInBowl dataset to conduct basic ablations. The first ablation removes the spatial correlation module Γ\Gamma and the position regression loss, therefore reducing RELATE to a 2D version of BlockGAN. We also consider ‘w/o residual’, where the addition θ^k\hat{\theta}_{k} in equation 1 is removed, and ‘w/o pos. loss’, where the position regression loss regularizer is removed. Table 1 shows that each component of RELATE yields an improvement in terms of FID scores on this dataset supporting our spatial modeling decisions. Furthermore, in figure 2 we show qualitatively that only RELATE (a) is able to correctly disentangle the underlying scene factors. We do this by generating the same image while retaining a single factor, which correctly isolates the background, and, in turn, both individual objects. BlockGAN 2D (b) and ‘ours w/o pos. reg.’ (d) fail to disentangle the factors entirely, mapping everything to the background component. ‘Ours w/o residual’ (c) shows that the model partially fails to disentangle, with the background encoding some but not all the objects. Finally, figure 3 visualizes the effect of the interaction module Γ\Gamma. Recall that this is implemented as a ‘correction’ function that accounts for correlation starting from independently-sampled parameters. For BallsInBowl, the correction module moves the balls within the bowl, and for the CLEVR it pushes objects apart if they intersect.

Quantitative evaluation.

In table 2, we compare RELATE to existing scene generators on ShapeStacks, CLEVR and RealTraffic. We report performance in terms of FID score computed between 10,000 images sampled from our model and the respective test sets. For CLEVR, we train RELATE and BlockGAN on a restricted version of the data containing from three to six objects in an imageNote that GENESIS was trained on the full training set featuring three to ten objects., and at test time, we require all models to sample images with three to ten objects. We consistently outperform all prior object-centric methods in all scenes and scenarios according to FID scores. In particular, on CLEVR, RELATE can generate a larger number of objects than seen during training suggesting its improved generalization capabilities which are demonstrated further in figure 6. Our method also consistently out-performs standard GANs on all CLEVR datasets and is on par with DRAGAN on RealTraffic. More qualitative results can be found in section A7.

2 Interpretability of the latent space and scene editing

As shown in the ablation studies in figure 2, RELATE successfully disentangles a scene into independent components - in contrast to BlockGAN2D which struggles to separate individual objects from the background. Figure 4 shows that RELATE can disentangle also far more complex scenes in RealTraffic, ShapeStacks and CLEVR. Note that for RealTraffic and CLEVR we render objects composed with the background. In fact, in these datasets the size and appearance of each object is correlated to their position in the background because of the camera perspective. In addition to qualitative results, we also compute a disentanglement score in table A1 which measures how well our model is disentangling individual components of the scene. We found that our model manages to consistently separate each individual objects of the scene and outperform BlockGAN2D on the most challenging datasets which is in line with the qualitative evidence we observe. Next, in figure 5 we use RELATE to edit a generated scene. For example we can change the position or appearance of individual objects. Finally, we show that RELATE can generate out-of-distribution scenes. This is achieved in particular by sampling a different number of objects. In figure 6, for instance, RELATE is trained on ShapeStacks seeing towers of height two to five. However, it can render taller towers of up to seven objects, or even just a single object. Likewise, in BallsInBowl it can generate bowls with four balls having seen only two during training. Furthermore, in ShapeStacks each tower is composed of blocks of different colors, but RELATE can relax this constraint rendering objects with repeated colors.

3 Simulating dynamics

We train this model on BallsInBowl and RealTraffic to predict 15 and 10 consecutive frames respectively. During generation, we sample videos with a sequence length of 30 frames and measure the faithfulness with respect to the distribution of the test data via the Fréchet Video Distance (FVD) . We achieve FVD scores of 556556 and 22532253 respectively. This is perceptibly better than 920920 and 33703370 for a baseline consisting of time-shuffled sequences from the respective training sets, which feature perfect resolution but poor dynamics. Qualitatively in figure 7 we see that the model does understand the motion and captures interaction with the background. For instance, in BallsInBowl the balls do have a curved motion because of the shape of the bowl and decrease in speed when reaching the edges of the bowl which are in higher position (see first row of figure 7). In RealTraffic the cars do stay in their respective lane. Interestingly our model is able to handle different types of motions correctly (see third row figure 7) and uses the sample vector to decide whether the cars should go straight or turn. Finally, we see that we can also generate videos with much more cars than the upper bound (5) with which the system was trained (see last row figure 7).

While the dynamics in video generation look realistic in most cases, RealTraffic also exposed the limitations of our approach. In this dataset the perspective range of the camera is important. As a result, cars at the bottom of the image, which appear bigger in the training data, often do not get generated properly in the static case (see highlighted frames in figure 8). We hypothesize that this is also the main reason why the cars’ appearance (and hence FVD score) deteriorates in the dynamic scenarios: since the style parameters ziz_{i} are fixed to preserve identity, it is not possible for the model to account accurately for the appearance change introduced by large changes of perspective over the course of a sequence. To verify our hypothesis in section A1.2 we propose a simple modification of our pipeline that accounts for scale changes in an image. This modification simply allows the network to predict H′H^{\prime} for each individual object instead of it being a constant. We found that in general this allows our network to train on datasets with more important range of scales such as CLEVR3 and renders objects at more different scales for CLEVR and RealTraffic (see figure A1). Besides in table 2 we show that this modification almost always results in higher FID scores from our main model as well as a significant boost of 397 points in FVD score for RealTraffic (1855 vs 2253).

Conclusion

We have introduced RELATE, a GAN-based model for object-centric scene synthesis featuring a physically interpretable parameter space and explicit modeling of spatial correlations. Our experimental results suggest that spatial correlation modeling plays a pivotal role in disentangling the constituent components of a visual scene. Once trained, RELATE’s interpretable latent space can be leveraged for targeted scene editing such as altering object positions and appearances, replacing the background or even inserting novel objects. We demonstrate our model’s effectiveness by presenting state-of-the-art scene generation results across a variety of simulated and real datasets. Lastly, we show how our model naturally extends to the generation of dynamic scenes being able to generate entire videos from scratch. A main limitation of our current model is its restriction to planar motions which prevents it from representing arbitrary 3D motions featuring angular rotation more faithfully, most notably highlighted by the experiments for video generation. This effect can however be partially mitigated by a scale-augmented model which we introduce in the appendix. We believe that our work can contribute to future research in object-centric scene representation by providing a scalable, spatio-temporal modeling approach which is conveniently trainable on unlabeled data.

Broader Impact

Our method advances the ability of computers to learn to understand environments in images in an object-centric way. It also enhances the capabilities of generative models to generate realistic images of “invented” environment configurations.

Overall, we believe our research to be at low to no risk of direct misuse. At present, our generation results are insufficient to fool a human observer. However, it has to be noted that the sampling process is, as in many other deep generative models, capable of revealing patterns observed in the training data, e.g., specific textures or object geometries. Such data privacy concerns are not applicable in the street traffic data used in our research, since the resolution of the videos is far too low to identify individual drivers or recognize cars’ license plates. However, ‘training data leakage’ should be taken into consideration when the model is trained on more sensitive datasets.

In a positive prospect, we believe that our model contributes to further the development of less opaque machine learning models. The explicit object-centric modelling of image components and their geometric relationships is in many of its aspects intelligble to a human user. This facilitates debugging and interpreting the model’s behaviour and can help to establish trust towards the model when employed in larger application pipelines.

However, the key value of our paper is in the methodological advances. It is conceivable that, like any advance in machine learning, our contributions could ultimately lead to methods that in turn can and are misused. However, there is nothing to indicate that our contributions facilitate misuse in any direct way; in particular, they seem extremely unlikely to be misused directly.

Acknowledgments

This work is supported by the European Research Council under grants ERC 638009-IDIU, ERC 677195-IDIU, and ERC 335373. The authors acknowledge the use of Hartree Centre resources in this work. The STFC Hartree Centre is a research collaboratory in association with IBM providing High Performance Computing platforms funded by the UK’s investment in e-Infrastructure. The authors also acknowledge the use of the University of Oxford Advanced Research Computing (ARC) facility in carrying out this work (http://dx.doi.org/10.5281/zenodo.22558). Special thanks goes to Olivia Wiles for providing feedback on the paper draft, Thu Nguyen-Phuoc for providing the implementation of BlockGAN and Titas Anciukevičius for providing the generation code for CLEVR variants. We finally would like to thank our reviewers for their diligent and valuable feedback on the initial submission of this manuscript.

References

A1 Additional experiments

We have conducted additional experiments to provide more quantitative insights of the disentanglement capabilities of our model. While measures such as MIG are typically used to quantify disentanglement, computing this score is not applicable in our case since our model does not feature an inference component to compute the posterior q(z∣x)q(z|x). Hence, we have devised a proxy procedure: We toggle each object of an image individually (out of 5 objects generated) and measure how the generated image changes. We report the distance between the pixel location corresponding to the maximum image change and the location (scaled θi\theta_{i}) of the object that was toggled. In table A1 we report the median distance between θi\theta_{i} and the pixel location corresponding to the maximum image change. We note that our model generally outperforms BlockGAN2D, most notably for ShapeStacks where BlockGAN is not able to disentangle different objects at all. In addition, we note that for the model trained on ShapeStacks, the discriminator can predict the position of a stack’s base object with 11.3 mean pixel error on the test set.

A1.2 Scale experiment

In order to tackle the limitation discussed in section 4, we propose to augment our model with scale prediction. Practically, this translates to predicting H′H^{\prime} for each individual object instead of keeping it fixed. Therefore, we now assign Hk′H^{\prime}_{k} instead of H′H^{\prime} to each foreground component of the scene. Hk′H^{\prime}_{k} is computed by a module scsc following the equation: Hk′=H′×(1+sc(z0,θk,zk))H^{\prime}_{k}=H^{\prime}\times(1+sc(z_{0},\theta_{k},z_{k})). More details on scsc can be found in table A2. We summarize all hyper-parameters used for the training of the model in table A3. For evaluation, we sample z0z_{0} from U(Nb)\mathcal{U}(^{N_{b}}), except for the RealTraffic video model where we use U([−0.5,0.5]Nb)\mathcal{U}([-0.5,0.5]^{N_{b}}), a range better suited for optimal background fidelity on this dataset.

A2 Further Discussions

As noted in the conclusions, RELATE is limited to 2D representations. However, our relaionship module is generic enough to be exported to 3D and could be inserted directly into BlockGAN – albeit at the expense of significant additional training time. Finally, we acknowledge that the simplified spatial representation of an object as its centroid position is inferior to other object-centric models which predict full object masks, e.g. . While we believe that an object segmentation mechanism could be easily added based on the already existing rendering of individual latent variables (cf. figure 4), we conjecture that the simplicity of the centroid representation greatly facilitates the learning of spatial correlations within Γ\Gamma. In contrast, using object instance masks as spatial representations might introduce unnecessary complications for the neural physics approximation.

A3 Losses

Our final loss is the sum of three losses mentioned in the main text:

where LGAN(I^,I)\mathcal{L}_{\text{GAN}}(\hat{I},I) is the standard GAN loss:

The style loss follows the implementation of BlockGAN . The input of the style discriminator DlD_{l} are mean μl\mu_{l} and variance σl2\sigma_{l}^{2} across spatial dimensions of Φl(x)∈RWl×Hl×Cl\Phi_{l}(x)\in\mathcal{R}^{W_{l}\times H_{l}\times C_{l}}, the output of the llth layer of DD taken before the normalization step:

The style discriminator DlD_{l} for each layer is then implemented as a linear layer followed by a sigmoid activation function. The resulting style loss is:

A4 Implementation details

For all experiments we use PyTorch 1.4. We train all models on a single NVIDIA Tesla V100 GPU.

Training hyperparameters

We initialize all weights (including instance normalization ones) by drawing from a random normal distribution N(0,0.02)\mathcal{N}(0,0.02). All biases were initialized to 0. For each update of the discriminator we update the generator MM number of times. We use Adam parameters (β1,β2)=(0.,0.999)(\beta_{1},\beta_{2})=(0.,0.999) for all datasets except BallsInBowl where β1=0.5\beta_{1}=0.5. Similarly WW was a max-pooling operator in all datasets except BallsInBowl where we used a sum pooling operator. As in BlockGAN, background and foreground decoders each start from a learned constant tensors TbT_{b} and TfT_{f} respectively with sizes H×H×256H\times H\times 256 and H′×H′×512H^{\prime}\times H^{\prime}\times 512. For BallsInBowl we use a tensor TfT_{f} for each object and use a constant style vector of one.

In the case of dynamic scenarios we reuse same hyperparameters as in the static case except that we use a learning rate of 0.0001 and β1=0.\beta_{1}=0.

Full details of the parameters for each dataset can be found in table A4

Evaluation details

For FID scores computation we draw 10 000 samples from our model which we compare against the same number of images drawn from the test set. To compute FVD score on each dataset, we sample 500 videos of 30 frames from our model and compare them against the videos of the respective test sets (500 for BallsInBowl and 275 videos for RealTraffic). This also applies to the time shuffled baseline.

To be able to compare with other methods we resize our generated images to 96×9696\times 96 on CLEVR5 and CLEVR5-vbg and 64×6464\times 64 for ShapeStacks. For the simple generative baselines, DRAGAN and DCGAN we evaluate FID score on the generated 64×6464\times 64 images from these models. We evaluate on the generated 128×128128\times 128 images otherwise.

We empirically found that background was rendered with better quality for lower values of z0z_{0}. Hence at test time we sampled z0z_{0} from U([−0.5,0.5]Nb)\mathcal{U}([-0.5,0.5]^{N_{b}}) for optimal results.

A4.1 Architecture details

In this work we maintain the core of our architecture fixed as much as possible. Since the dimension of the sample ziz_{i} does not necessarily match the channel dimension where it is injected before applying Adaptive Instance Normalisation (AdaIN) to a layer ll we map ziz_{i} to a vector z^i\hat{z}_{i} transformed such that

Where (Wl,bl)(W_{l},b_{l}) are learnable parameters. AdaIN is applied at the end of the layers (after the activation). All LeakyReLU layers are using a parameter of 0.2.

Discriminator

We describe the architecture of the discriminator network in more details in table A13. We use spectral normalization at almost every layer. Positions are directly regressed from the last feature output of the discriminator (see last line PendP_{end}). Therefore in practice PP and DD share the same backbone DbD_{b} (see table table A13 until flatten) for every image I:

Input for style discriminator are taken after the convolution of (Convd_2, Convd_3, Convd_4, Convd_5) in table A13 before the normalization. Spectral Normalization was not applied to any DlD_{l}.

A5 Baselines

We used an online pytorch implementationhttps://github.com/LynnHo/DCGAN-LSGAN-WGAN-GP-DRAGAN-Pytorch with default hyper-parameters. We trained these models to generate 64×6464\times 64 images and therefore only evaluated FID score at the same resolution (see section A4).

OCF.

OCF results were copied from original paper of .

BlockGAN2D.

We use the same hyperparameters and network architecture as RELATE except for learning rate and MM. In all cases we report the best results over models trained with variations of learning rate in (0.001, 0.0001) and MM in (2,3).

GENESIS.

We use the official implementationhttps://github.com/applied-ai-lab/genesis of GENESIS for all experiments. For the ShapeStacks dataset, we use the official model snapshot released with the original paperhttps://drive.google.com/drive/folders/1uLSV5eV6Iv4BYIyh0R9DUGJT2W6QPDkb?usp=sharing. For all other datasets, we train GENESIS for 500,000 iterations with the default learning parameters and select the last model checkpoint for evaluation. When training GENESIS we use constrained ELBO optimization controlled via g_goal in the training script which influences the decomposition capability of GENESIS. We perform a grid search over g_goal in the range of 0.5635 to 0.5655 and select the model with the lowest ELBO after 500,000 iterations.

A6 Datasets

This dataset is a replica of the two balls synthetic dataset of . It consists of 2500 training sequences and 500 test sequences of two balls of different fixed colour rolling in bowls of various shapes. We count an epoch as 10,000 iterations over the data. In figure A2 we show some sample data from this dataset.

CLEVR.

We used the official CLEVR from . We train on data from train and validation set and evaluate on the test set. Both ours and BlockGAN2D were trained on the subset containing 3 to 6 objects and evaluated on the entire test set.

CLEVR5/CLEVR5-vbg.

We use online code provided by the authorshttps://github.com/TitasAnciukevicius/clevr-dataset-gen. to generate CLEVR5 and CLEVR5-vbg. As done in we generate 100,000 images keep 90,000 for training and 10,000 for testing.

ShapeStacks

We use the official release of the ShapeStacks datasethttps://shapestacks.robots.ox.ac.uk/#data. We use the default partitioning provided with the dataset and merge the training and validation splits for a total of 264,384 training images. All FID comparisons are made against 10,000 images randomly sampled from the test set which contains 46,560 images in total. Since the original resolution of the images is 224×224224\times 224 pixels, we re-scale them to 128×128128\times 128 before feeding them to our network.

RealTraffic.

We recorded 5 hours from Youtubehttps://www.youtube.com/watch?v=5_XSYlAfJZM of a live traffic camera at a crossing. The video was then unrolled at 10 fps and manually processed to keep only sequences with a number of cars in . We kept 560 videos for the training set and 123 in test (80/20 ratio). This dataset will be publicly released.

A7 Qualitative results

We provide additional qualitative generation results. Figure A3 shows a failure case of BlockGAN2D mentionned in the paper. In fact, when the scene is more structured BlockGAN2D fails to be object centric and let the background render the entire scene. In addition figures A4, A8, A6, A7, LABEL: and A9 provide more samples on every dataset for all the models we trained. In particular we can see that when inter-objects relations are weak in CLEVR5 or CLEVR5-vbg, BlockGAN2D performs qualitatively similar to ours (see figures A6, LABEL: and A7). However when the scene is more crowded and the objects have higher correlation BlockGAN2D quality decreseases significantly (see figures A4, A8, LABEL: and A9).