DriveGAN: Towards a Controllable High-Quality Neural Simulation

Seung Wook Kim, Jonah Philion, Antonio Torralba, Sanja Fidler

Introduction

The ability to simulate is a key component of intelligence. Consider how animals make thousands of decisions each day. Some of the decisions are critical for survival, such as deciding to step away from an approaching car. Mentally simulating the future given the current situation is key in planning successfully. In robotic applications such as autonomous driving, simulation is also a scaleable, robust and safe way of testing self-driving vehicles in safety-critical scenarios before deploying them in the real world. Simulation further allows for a fair comparison of different autonomous driving systems since one has control over the repeatability of the scenarios.

Desired properties of a good robotic simulator include accepting an action from an agent and generating a plausible next world state, allowing for user control over the scene elements, and the ability to re-simulate an observed scenario with plausible variations. This is no easy feat as the world is incredibly rich in situations one can encounter. Most of the existing simulators are hand-designed in a game engine, which involves significant effort in content creation, and designing complex behavior models to control non-player objects. Grand Theft Auto, one of the most realistic driving games to date, set in a virtual replica of Los Angeles, took several years to create and involved hundreds of artists and engineers. In this paper, we advocate for data-driven simulation as a way to achieve scaleability.

Data-driven simulation has recently gained attention. LidarSim used a catalog of annotated 3D scenes to sample layouts into which reconstructed objects obtained from a large number of recorded drives are placed, in the quest to achieve diversity for training and testing a LIDAR-based perception system. , on the other hand, learn to synthesize road-scene 3D layouts directly from images without supervision. These works do not model the dynamics of the environment and object behaviors.

As a more daring alternative, recent works attempted to create neural simulators that learn to simulate the environment in response to the agent’s actions directly in pixel-space by digesting large amounts of video data along with actions. This line of work provides a scaleable way to simulation, as we do not rely on any human-provided annotations, except for the agent’s actions which are cheap to obtain from odometry sensors. It is also a more challenging way, since the complexity of the world and the dynamic agents acting inside it, needs to be learned in a high-resolution camera view. In this paper, we follow this route.

We introduce DriveGAN, a neural simulator that learns from sequences of video footage and associated actions taken by an ego-agent in an environment. DriveGAN leverages Variational-Auto Encoder and Generative Adversarial Networks to learn a latent space for images on which a dynamics engine learns the transitions within the latent space. The key aspects of DriveGAN are its disentangled latent space and high-resolution and high-fidelity frame synthesis conditioned on the agent’s actions. The disentanglement property of DriveGAN gives users additional control over the environment, such as changing the weather and locations of non-player objects. Furthermore, since DriveGAN is an end-to-end differentiable simulator, we are able to re-create the scenarios observed from real video footage allowing the agent to drive again through the recorded scene but taking different actions. This property makes DriveGAN the first neural driving simulator of its kind. By learning on 160 hours of real driving data, we showcase DriveGAN to learn high-fidelity simulation, surpassing all existing neural simulators by a significant margin, and allowing for the control over the environment not possible previously.

Related Work

As in image generation, the standard architectures for video generation are VAEs , auto-regressive models , flow-based models , and GANs . For a generator to sample videos, it must be able to generate realistic looking frames as well as realistic transitions between frames. Video prediction models learn to produce future frames given a reference frame, and they share many similarities to video generation models. Similar architectures can be applied to the task of conditional video generation in which information such as semantic segmentation is given as input to the model . In this work, we use a VAE-GAN based on StyleGAN to learn a latent space of natural images, then train a dynamics model within the space.

2 Data-driven Simulation and Model-based RL

The goal of data-driven simulation is to learn simulators given observations from the environment to be simulated. Meta-Sim learns to produce scene parameters in a synthetic scene. LiDARSim leverages deep learning and physics engine to produce LiDAR point clouds. In this work, we focus on data-driven simulators that produce future frames given controls. World Model use a VAE and LSTM to model transition dynamics and rendering functionality. In GameGAN , a GAN and a memory module are used to mimic the engine behind games such as Pacman and VizDoom. Model-based RL also aims at learning a dynamics model of some environment which agents can utilize to plan their actions. While prior work has applied neural simulation to simple environments in which a ground-truth simulator is already known, we also apply our model to real-world driving data and focus on improving the quality of simulations. Furthermore, we show how users can interatively edit scenes to create diverse simulation environments.

Methodology

Our objective is to learn a high-quality controllable neural simulator by watching sequences of video frames and their associated actions. We aim to achieve controllability in two aspects: 1) We assume there is an egocentric agent that can be controlled by a given action. 2) We want to control different aspects of the current scene, for example, by modifying an object or changing the background color.

Let us denote the video frame at time tt as xtx_{t} and the continuous action as ata_{t}. We learn to produce the next frame xt+1x_{t+1} given the previous frames x1:tx_{1:t} and actions a1:ta_{1:t}. Fig 2 provides an overview of our model. Image encoder ξ\xi produces the disentangled latent codes zthemez^{\text{theme}} and zcontentz^{\text{content}} for xx in an unsupervised manner. We define theme as information that does not depend on pixel locations such as the background color or weather of the scene, and content as spatial content (Fig 4). Dynamics Engine, a recurrent neural network, learns to produce the next latent codes zt+1themez_{t+1}^{\text{theme}}, zt+1contentz_{t+1}^{\text{content}} given ztthemez_{t}^{\text{theme}}, ztcontentz_{t}^{\text{content}}, and ata_{t}. zt+1themez_{t+1}^{\text{theme}} and zt+1contentz_{t+1}^{\text{content}} go through an image decoder that generates the output image.

Generating high-quality temporally-consistent image sequences is a challenging problem . Rather than generating a sequence of frames directly, we split the learning process into two steps, motivated by World Model . Sec 3.1 introduces our encoder-decoder architecture that is pre-trained to produce the latent space for images. We propose a novel architecture that disentangles themes and content while achieving high-quality generation by leveraging a Variational Auto-Encoder (VAE) and Generative Adversarial Networks (GAN). Sec A.2 describes the Dynamics Engine that learns the latent space dynamics. We also show how the Dynamics Engine further disentangles action-dependent and action-independent content.

We build our image decoder on top of the popular StyleGAN , but make several modifications that allow for theme-content disentanglement. Since extracting the GAN’s latent code that corresponds to an input image is not trivial, we introduce an encoder ξ\xi that maps an image xx into its latent code zz. We utilize the VAE formulation, particularly the β\beta-VAE to control the KL term better. Therefore, on top of the adversarial losses from StyleGAN, we add the following loss at each step of generator training:

where p(z)p(z) is the standard normal prior distribution, q(z∣x)q(z|x) is the approximate posterior from the encoder ξ\xi, and KLKL is the Kullback-Leibler divergence. For the reconstruction term, we reduce the perceptual distance between the input and output images rather than the pixel-wise distance.

We observe that balancing the KL loss with suitable β\beta in LVAEL_{VAE} is essential. Smaller β\beta gives better reconstruction quality, but the learned latent space could be far away from the prior, in which case the dynamics model (Sec.A.2) had a harder time learning the dynamics. This causes zz to be overfit to xx, and it becomes more challenging to learn the transitions between frames in the overfitted latent space.

2 Dynamics Engine

With the pre-trained encoder and decoder, the Dynamics Engine learns the transition between latent codes from one time step to the next given an action ata_{t}. We fix the parameters of the encoder and decoder, and only learn the parameters of the engine. This allows us to pre-extract latent codes for a dataset before training. The training process becomes faster and significantly easier than directly working with images, as latent codes typically have dimensionality much smaller than the input. In addition, we further disentangle content information from zcontentz^{\text{content}} into action-dependent and action-independent features without supervision.

In a 3D environment, the view-point shifts as the ego agent moves. This shifting naturally happens spatially, so we employ a convolutional LSTM module (Figure 18) to learn the spatial transition between each time step:

where we denote convolution and AdaINAdaIN layers as C\mathcal{C} and A\mathcal{A}, respectively. An MLPMLP is used to produce α\bm{\alpha} and β\bm{\beta}. We reparameterize zadepz^{a_{\text{dep}}},zaindepz^{a_{\text{indep}}}, zthemez^{\text{theme}} into the standard normal distribution N(0,I)N(0,I) which allows sampling at test time:

where μ\mu and σ\sigma are the intermediate variables for the mean and standard deviation for each reparameterization step.

Intuitively, zaindepz^{a_{\text{indep}}} is used as stylestyle for the spatial tensor zadepz^{a_{\text{dep}}} through AdaINAdaIN layers. zaindepz^{a_{\text{indep}}} does not get action information, so it alone cannot learn to generate plausible next frames. This architecture thus allows disentangling action-dependent features such as the layout of a scene from action-independent features such as object types. Note that the engine could ignore zaindepz^{a_{\text{indep}}} and only use zadepz^{a_{\text{dep}}} to learn dynamics. If we keep the model size small and use a high KLKL penalty on the reparameterized variables, it will utilize full model capacity and make use of zaindepz^{a_{\text{indep}}}. We can also enforce disentanglement between zaindepz^{a_{\text{indep}}} and zadepz^{a_{\text{dep}}} using an adversarial loss . In practice, we found that our model was able to disentangle information well without such a loss.

3 Differentiable Simulation

One compelling aspect of DriveGAN is that it can create an editable simulation environment from a real video. As DriveGAN is fully differentiable, it allows for recovering the scene and scenario by discovering the underlying factors of variations that comprise a video, while also recovering the actions that the agent took, if these are not provided. We refer to this as differentiable simulation. Once these parameters are discovered, the agent can use DriveGAN to re-simulate the scene and take different actions. DriveGAN further allows sampling and modification of various components of a scene, thus testing the agent in the same scenario under different weather conditions or objects.

First, note that reparametrization steps (Eq. 10) involve a stochastic variable ϵ\epsilon which gives stochasticity in a simulation to produce diverse future scenarios. Given a sequence of frames from a real video x0,...,xTx_{0},...,x_{T}, our model can be used to find the underlying a0,...,aT−1,ϵ0,...ϵT−1a_{0},...,a_{T-1},\epsilon_{0},...\epsilon_{T-1}:

where ztz_{t} is the output of our model, z^t\hat{z}_{t} is the encoding of xtx_{t} with the encoder, and λ1,λ2\lambda_{1},\lambda_{2} are hyperparameters for regularizers. We add action regularization assuming the action space is continuous and ata_{t} does not differ significantly from at−1a_{t-1}. To prevent the model from utilizing ϵt\epsilon_{t} to explain all differences between frames, we also add the ϵ\epsilon regularizer.

Experiments

We perform thorough quantitative (Sec 4.1) and qualitative (Sec 4.2) experiments on the following datasets.

Carla simulator is an open-source simulator for autonomous driving research. We use five towns in Carla to generate the dataset. The ego-agent and other vehicles are randomly placed and use random policy to drive in the environment. Each sequence has a randomly sampled weather condition and consists of 80 frames sampled at 4Hz. 48K sequences are extracted, and 43K are used for training.

Gibson environment virtualizes real-world indoor buildings and has an integrated physics engine with which virtual agents can be controlled. We first train a reinforcement learning agent that can navigate towards a given destination coordinate. In each sequence, we randomly place the agent in a building and sample a destination. 85K sequences each with 30 frames are extracted from 100 indoor environments, and 76K sequences are used for training.

Real World Driving (RWD) data consists of real-world recordings of human driving on multiple different highways and cities. It was collected in a variety of different weather and times. RWD is composed of 128K sequences each with 36 frames extracted at 8Hz. It corresponds to ∼\sim 160 hours of driving, and we use 125K sequences for training.

Figure 7 illustrates scenes from the datasets. Each sequence consists of the extracted frames (256×\times256) and the actions the ego agent takes at each time step. The 2-dim actions consist of the agent’s speed and angular velocity.

The quality of simulators needs to be evaluated in two aspects. The generated videos from simulators have to look realistic, and their distribution should match the distribution of the real videos. They also need to be faithful to the action sequences used to produce them. This is essential to be useful for downstream tasks, such as training a robot. Therefore, we use two automatic metrics to measure the performance of models. The experiments are carried out by using the first frames and action sequences of the test set. The remaining frames are generated autoregressively.

We compare with four baseline models: Action-RNN is a simple action-conditioned RNN model trained with reconstruction loss on the pixel space, Stochastic Adversarial Video Prediction (SAVP) and GameGAN are trained with adversarial loss along with reconstruction loss on the pixel space, World Model trains a vision model based on VAE and an RNN based on mixture density networks (MDN-RNN). World Model is similar to our model as they first extract latent codes and learn MDN-RNN on top of the learned latent space. However, their VAE is not powerful enough to model the complexities of the datasets studied in this work. Fig 9 shows how a simple VAE cannot reconstruct the inputs; thus, the plain World Model cannot produce realistic video sequences by default. Therefore, we include a variant, denoted as World Model*, that uses our proposed latent space to train the MDN-RNN component.

We also conduct human evaluations with Amazon Mechanical Turk. For 300 generated sequences from each dataset, we show one video from our model and one video from a baseline model for the same test data. The workers are asked to mark their preferences on ours versus the baseline model on visual qulity and action consistency (Fig 10).

Video Quality: Tab 1 shows the result on Fréchet Video Distance (FVD) . FVD measures the distance between the distributions of the ground truth and generated video sequences. FVD is an extension of FID for videos and is suitable for measuring the quality of generated videos. Our model achieves lower FVD than all baseline models except for GameGAN on Gibson. The primary reason we suspect is that our model on Gibson sometimes slightly changes the brightness. In contrast, GameGAN, being a model directly learned on pixel space, produced more consistent brightness. Human evaluation of visual quality (Fig 10) shows that subjects strongly prefer our model, even for Gibson.

Action Consistency: We measure if generated sequences conform to the input action sequences. We train a CNN model that takes two images from real videos as input and predicts the action that caused the transition between them. The model is trained by reducing the mean-squared error loss between the predicted and input actions. The trained model can be applied to the generated sequences from simulator models to evaluate action consistency. Table 2 and human evaluation (Fig 10) show that our model achieves the best performance on all datasets.

2 Controllability and Differentiable Simulation

DriveGAN learns to disentangle factors comprising a scene without supervision, and it naturally allows controllability on all zzs as zadepz^{a_{\text{dep}}}, zaindepz^{a_{\text{indep}}}, zcontentz^{\text{content}} and zthemez^{\text{theme}} can be sampled from the prior distribution. Fig 4 demonstrates how we can change the background color or weather condition by sampling and swapping zthemez^{\text{theme}}. Fig 12 shows how sampling different zaindepz^{a_{\text{indep}}} modifies the interior parts, such as object shapes, while keeping the layout and theme consistent. This allows users to sample various scenarios for specific layout shapes. As zcontentz^{\text{content}} is a spatial tensor, we can sample each grid cell to change the content of the cell. In the bottom row of Fig 11, a user clicks specific locations to erase a tree, add a tree, and add a building.

We also record the sampled zzs corresponding to specific content and build an editable neural simulator, as in Fig 1. This editing procedure lets users create unique simulation scenarios and selectively focus on the ones they want. Note that we can even sample the first screen, unlike some previous works such as GameGAN .

Differentiable Simulation: Sec 3.3 introduces how we can create an editable simulation environment from a real video by recovering the underlying actions aa and stochastic variables ϵ\epsilon with Eq.(9). Fig 13 illustrates the result of differentiable simulation. The third row exhibits how we can recover the original video A by running DriveGAN with optimized aa and ϵ\epsilon. To verify we have recovered aa successfully and not just overfitted using ϵ\epsilon, we evaluate the quality of optimized aa from test data using the Action Prediction loss from Tab 2. Optimized aa results in a loss of 1.91 and 0.57 for Carla and RWD, respectively. These numbers are comparable to Tab 2 and much lower than the baseline performances of 3.64 and 1.01, calculated with the mean of actions from the training data, demonstrating that DriveGAN can recover unobserved actions successfully. We can even recover aa and ϵ\epsilon for non-existing intermediate frames. That is, we can do frame interpolation to discover in-between frames given a reference and a future frame. If the time between the two frames is small, even a naive linear interpolation could work. However, for a large gap (≥\geq 1 second), it is necessary to reason about the environment’s dynamics to properly interpolate objects in a scene. We modify Eq.(9) to minimize the reconstruction term for the last frame zTz_{T} only, and add a regularization ∣∣zt−zt−1∣∣||z_{t}-z_{t-1}|| on the intermediate zzs. Fig 14 shows the result. Top row, which shows interpolation in the latent space, produces reasonable in-between frames, but if inspected closely, we can see the transition is unnatural (\ega tree appears out of nowhere). On the contrary, with differentiable simulation, we can see how it learns to utilize the dynamics of DriveGAN to produce plausible transitions between frames. In Fig 15, we calculate the action prediction loss with optimized actions from frame interpolation. We discover optimized actions that follow the ground-truth actions closely when we interpolate frames one second apart. As the interpolation interval becomes larger, the loss increases since many possible action sequences lead to the same resulting frame. This shows the possibility of using differentiable simulation for video compression as it can decode missing intermediate frames.

Differentiable simulation also allows replaying the same scenario with different inputs. In Fig 13, we get optimized aA,ϵAa^{A},\epsilon^{A} and aB,ϵBa^{B},\epsilon^{B} for two driving videos, A and B. We replay starting with the encoded first frame z0Az_{0}^{A} of A. On the fourth row, ran with (aBa^{B},ϵA\epsilon^{A}), we see that a vehicle is placed at the same location as A, but since we use the slightly-left action sequence aBa^{B}, the ego agent changes the lane and slides toward the vehicle. The fifth row, replayed with (aAa^{A},ϵB\epsilon^{B}), shows the same ego-agent’s trajectory as A, but it puts a vehicle at the same location as B due to ϵB\epsilon^{B}. This effectively shows that we can blend-in two different scenarios together. Furthermore, we can modify the content and run a simulation with the environment inferred from a video. In Fig 6, we create a simulation environment from a RWD test data, and replay with modified objects and weather.

3 Additional Experiments

LiftSplat proposed a model for producing the Bird’s-Eye-View (BEV) representation of a scene from camera images. We use LiftSplat to get BEV lane predictions from a simulated sequence from DriveGAN (Fig 16). Simulated scenes are realistic enough for LiftSplat to produce accurate predictions. This shows the potential of DriveGAN being used with other perception models to be useful for downstream tasks such as training an autonomous driving agent. Furthermore, in real-time driving, LiftSplat can potentially employ DriveGAN’s simulated frames as a safety measure to be robust to sudden camera drop-outs.

Plain StyleGAN latent space: StyleGAN proposes an optimization scheme to project images into their latent codes without an encoder. The projection process optimizes each image and requires significant time (∼\sim19200 GPU hours for Gibson). Therefore, we use 25% of Gibson data to compare with the projection approach. We train the same dynamics model on top of the projected and proposed latent spaces. The projection approach resulted in FVD of 636.8 with the action prediction loss of 0.225, whereas ours achieved 411.9 (FVD) and 0.050 (action prediction loss).

Conclusion

We proposed DriveGAN for a controllable high-quality simulation. DriveGAN leverages a novel encoder and an image GAN to produce a latent space on which the proposed dynamics engine learns the transitions between frames. DriveGAN allows sampling and disentangling of different components of a scene without supervision. This lets users interactively edit scenes during a simulation and produce unique scenarios. We showcased differentiable simulation which opens up promising ways for utilizing real-world videos to discover the underlying factors of variations and train robots in the re-created environments.

References

Appendix A Model Architecture and Training

We provide detailed descriptions of the model architecture and training process for the pre-trained image encoder-decoder (Sec. A.1) and dynamics engine (Sec. A.2). Unless noted otherwise, we denote tensor dimensions by H×W×DH\times W\times D where HH and WW are the spatial height and width of a feature map, and DD is the number of channels.

The latent space is pretrained with an encoder, generator and discrminator. Figure 17 shows the overview of the pretraining model.

The above tables show the architecture for each component. ξfeat\xi^{\text{feat}} takes xx as input and consists of several convolution layers whose output is passed to the two heads. Conv2d 3×\times3 denotes a 2D convolution layer with 3×\times3 filters and padding of 1 to produce the same spatial dimension as input. ResBlock denotes a residual block with downsampling by 2×\times which is composed of two 3×\times3 convolution layers and a skip connection layer. After each layer, we put the leaky ReLU activation function, except for the last layer of ξcontent\xi^{\text{content}} and ξtheme\xi^{\text{theme}}. The outputs of ξcontent\xi^{\text{content}} and ξtheme\xi^{\text{theme}} are equally split into two chunks by the channel dimension, and used as μ\mu and σ\sigma for the reparameterization steps:

A.1.2 Generator

The generator architecture closely follows the generator of StyleGAN . Here, we discuss a few differences. zcontentz^{\text{content}} goes through a 3×\times3 convolution layer to make it a 4×\times4×\times512 tensor. StyleGAN takes a constant tensor as an input to the first layer. We concatenate the constant tensor with zcontentz^{\text{content}} channel-wise and pass it to the first layer. zthemez^{\text{theme}} goes through 8 linear layers, each outputting a 1024-dimensional vector, and the output is used for the adaptive instance normalization layers in the same way style vectors are used in StyleGAN. The generator outputs a 256×\times256×\times3 image.

A.1.3 Discriminator

Dicriminator takes the real and generated images (256×\times256×\times3) as input. We use multi-scale multi-patch discriminators , which results in higher quality images for complex scenes.

We use three discrminators D1,D2,D_{1},D_{2}, and D3D_{3}. D1D_{1} takes a 256×\times256×\times3 image as input and produces a single number. D2D_{2} takes a 256×\times256×\times3 image as input and produces 16×\times16 patches each with a single number. D3D_{3} takes a 128×\times128×\times3 image as input and produces 8×\times8 patches each with a single number. The adversarial losses for D2D_{2} and D3D_{3} are averaged across the patches. The inputs to D1,D2,D3D_{1},D_{2},D_{3} are the real and generated images, except that the input to D3D_{3} is downsampled by 2×\times. The model architectures are described in the above tables, and we use the same convolution layer and residual blocks from the previous sections. Each layer is followed by a leaky ReLU activation function except for the last layer.

A.1.4 Training

We combine the loss functions of VAE and GAN , and let Lpretrain=LVAE+LGANL_{pretrain}=L_{VAE}+L_{GAN}. We use the same loss function for the adversarial loss LGANL_{GAN} from StyleGAN , except that we have three terms for each discriminator. LVAEL_{VAE} is defined as:

where p(z)p(z) is the standard normal prior distribution, q(z∣x)q(z|x) is the approximate posterior from the encoder ξ\xi, and KLKL is the Kullback-Leibler divergence. For the reconstruction term, we reduce the perceptual distance between the input and output images rather than the pixel-wise distance, and this term is weighted by 25.0. We use separate β\beta values βtheme\beta^{\text{theme}} and βcontent\beta^{\text{content}} for zcontentz^{\text{content}} and zthemez^{\text{theme}}. We also found different β\beta values work better for different environments. We use βtheme=1.0,βcontent=2.0\beta^{\text{theme}}=1.0,\beta^{\text{content}}=2.0 for Carla, βtheme=1.0,βcontent=4.0\beta^{\text{theme}}=1.0,\beta^{\text{content}}=4.0 for Gibson, and βtheme=1.0,βcontent=1.0\beta^{\text{theme}}=1.0,\beta^{\text{content}}=1.0 for RWD. Adam optimizer is employed with learning rate of 0.002 for 310,000 optimization steps. We use a batch size of 16.

A.2 Dynamics Engine

With the pre-trained encoder and decoder, the Dynamics Engine learns the transition between latent codes from one time step to the next given an action ata_{t}. We first pre-extract the latent codes for each image in the training data, and only learn the transition between the latent codes. All neural network layers described below are followed by a leaky ReLU activation function, except for the outputs of discriminators, the outputs for μ,σ\mu,\sigma variables used for reparameterization steps, and the outputs for the AdaIN parameters.

The major components of the Dynamics Engine are its two LSTM modules. The first one learns the spatial transition between the latent codes and is implemented as a convolutional LSTM module (Figure 18).

Finally, zt+1adepz_{t+1}^{a_{\text{dep}}} and zt+1aindepz_{t+1}^{a_{\text{indep}}} are used as inputs to two AdaINAdaIN + Conv blocks.

We use disciminators on the flattened 1152 dimensional latent codes zz (concatenation of zthemez^{\text{theme}} and flattened zcontentz^{\text{content}}). There are two discriminators 1) single latent discriminator DsingleD_{single}, and 2) temporal action-conditioned discriminator DtemporalD_{temporal}.

We denote SNLinear and SNConv as linear and convolution layers with Spectral Normalization applied, and BN as 1D Batch Normalization layers . DsingleD_{single} is a 6-layer MLPMLP that tries to discriminate generated zz from the real latent codes. It takes a single zz as input and produces a single number. For the temporal action-conditioned discriminator DtemporalD_{temporal}, we first reuse the 1024-dimensional feature representation from the fourth layer of DsingleD_{single} for each ztz_{t}. The represenations for ztz_{t} and zt−1z_{t-1} are concatenated and go through a SNLinear layer to produce the 1024-dimensional temporal discriminator feature. Let us denote the temporal discriminator feature as zt,t−1z_{t,t-1}. The action ata_{t} also goes through a SNLinear layer to produce the 1024-dimensional action embedding. zt,t−1z_{t,t-1} and the action embedding are concatenated and used as the input to DtemporalD_{temporal}. We use 32 time-steps to train DriveGAN, so the input to DtemporalD_{temporal} has size 2048×\times31 where 31 is the temporal dimension. Table 8 shows the architecture of DtemporalD_{temporal}. After each layer of DtemporalD_{temporal}, we put a 3-timestep wide convolution layer that produces a single number for each resulting time dimension. Therefore, there are three outputs of DtemporalD_{temporal} with sizes 14, 11, and 4 which can be thought of as patches in the temporal dimension. We also sample negative actions aˉt\bar{a}_{t}, and the job of DtemporalD_{temporal} is to figure out if the given sequence of latent codes is realistic and faithful to the given action sequences. aˉt\bar{a}_{t} is sampled randomly from the training dataset.

A.2.2 Training

We use Adam optimizer with learning rate of 0.0001 for 400,000 optimization steps. We use batch size of 128 each with 32 time-steps and train with a warm-up phase. In the warm-up phase, we feed in the ground-truth latent codes as input for the first 18 time-steps and linearly decay the number to 1 at 100-th epoch, which corresponds to completely autoregressive training at that point. We use the loss LDE=Ladv+Llatent+Laction+LKLL_{DE}=L_{adv}+L_{latent}+L_{action}+L_{KL}. LadvL_{adv} is the adversarial losses, and we use the hinge loss . We also add a R1R_{1} gradient regularizer to LadvL_{adv} that penalizes the gradients of discriminators on true data . LactionL_{action} is the action reconstruction loss (implemented as a mean squared error loss) which we obtain by running the temporal discriminator features zt,t−1z_{t,t-1} through a linear layer to reconstruct the input action at−1a_{t-1}. Finally, we add the latent code reconstruction loss LlatentL_{latent} (implemented as a mean squared error loss) so that the generated ztz_{t} matches the input latent codes, and reduce the KLKL penalty LKLL_{KL} for ztadepz_{t}^{a_{\text{dep}}},ztaindepz_{t}^{a_{\text{indep}}}, ztthemez_{t}^{\text{theme}}. LlatentL_{latent} is weighted by 10.0 and we use different β\beta for the KLKL penalty terms. We use βadep=0.1,βaindep=0.1,βtheme=1.0\beta^{a_{\text{dep}}}=0.1,\beta^{a_{\text{indep}}}=0.1,\beta^{\text{theme}}=1.0 for Carla, and βadep=0.5,βaindep=0.25,βtheme=1.0\beta^{a_{\text{dep}}}=0.5,\beta^{a_{\text{indep}}}=0.25,\beta^{\text{theme}}=1.0 for Gibson and RWD.

Appendix B Additional Analysis on Experiments

Multi-patch Multi-scale discriminator We experimented with Carla dataset to choose the image discriminator architecture. In contrast to the plain StyleGAN, the datasets studied in this work contain much more diverse objects in multiple locations. Using a multi-patch multi-scale discriminator improved our FID score on Carla images from 72.3 to 67.1 over the StyleGAN discriminator.

LiftSplat proposed a model for producing the Bird’s-Eye-View (BEV) representation of a scene from camera images. Section 4.3 in the main text shows how we can leverage LiftSplat to get BEV lane predictions from a simulated sequence from DriveGAN. We can further analyze the qualitative result by comparing how the perception model (LiftSplat) perceives the ground truth and generated sequences differently. We fit a quadratic function to the LiftSplat BEV lane prediction for each image in the ground-truth sequence, and compare the distance between the fitted quadratic and the predicted lanes.

We show results on different look-ahead distances, which denote how far from the ego-car we are making the BEV predictions for. The above table lists the mean distance from the BEV lane predictions and the fitted quadratic function. Random compares the distance between the fitted quadratic and the BEV prediction for a randomly sampled RWD sequence. DriveGAN compares the distance for the BEV prediction for the optimized sequence with differentiable simulation of DriveGAN. Ground-Truth compares the distance for the BEV prediction for the ground-truth image. Note that Ground-Truth is not 0 since the fitted quadratic does not necessarily follow the lane prediction from the ground-truth image exactly. We can see that DriveGAN-optimized sequences produce lanes that follow the ground-truth lanes, which demonstrates how we could find the underlying actions and stochastic variables from a real video through differentiable simulation.

Appendix C DriveGAN Simulator User Interface

We build an interactive user interface for users to play with DriveGAN. Figure 19 shows the application screen. It has controls for the steering wheel and speed, which can be controlled by the keyboard. We can randomize different components by sampling zaindep,zcontentz^{a_{\text{indep}}},z^{\text{content}} or zthemez^{\text{theme}}. We also provide a pre-defined list of themes and objects that users can selectively use for specific changes. The supplementary video demonstrates how this UI can enable interactive simulation.