Guided Conditional Diffusion for Controllable Traffic Simulation
Ziyuan Zhong, Davis Rempe, Danfei Xu, Yuxiao Chen, Sushant Veer, Tong Che, Baishakhi Ray, Marco Pavone
I Introduction
Simulation is crucial to comprehensively evaluate modern autonomous vehicles (AVs). Due to the difficulty and danger of running large-scale real-world tests , AV developers rely on extensive testing in simulation to produce reliable systems . To be most useful, simulators must embody both realism and controllability, especially for the models of traffic. Realistic traffic allows developments made in simulation to faithfully transfer to the real world, while controllability enables constructing fine-grained traffic scenarios to analyze specific AV behavior. Yet, developing realistic traffic models is still an open challenge, and little attention has been devoted to making such models easily controllable.
Existing driving simulators commonly synthesize agent behaviors by either replaying recorded driving logs or using heuristic-based controllers . Such behaviors can often be controlled through a user-friendly high-level programming API or scenario editor UI . However, these methods lack realism as they either cannot react to other traffic agents or are not expressive enough to appear human-like. To address these issues, recent works propose to learn generative models of traffic behavior from large-scale driving datasets . While generated behaviors from these models appear reasonable, they merely reflect the distribution of the training data: they lack any mechanism to control the generated traffic flow, making these methods less useful in practice. For example, to test an AV against cut-in from a neighboring lane, a user needs an interface to steer the vehicle to perform the maneuver. Unfortunately, the opaque nature of neural network models makes this difficult for current approaches.
We look to bridge the gap between realism and controllability by studying the problem of controllable traffic behavior generation. We seek a model that can generate trajectories to meet specific desirable objectives from a user at inference time, which may differ from objectives used in training. This is different from traffic models that aim to learn maximum-likelihood behaviors, and therefore are not flexible to new objectives. Moreover, controllable generation is an especially challenging for learned approaches, which typically struggle to produce samples outside of the training distribution.
To enable such flexible generation, we leverage recent advances in diffusion modeling, which have achieved state-of-the-art performance in several domains including images, audio, and pedestrian trajectories. Importantly, diffusion models allow a notion of control at generation time through so-called guidance, which has benefited several tasks including conditional image generation , language generation, and offline reinforcement learning. Inspired by these works, we propose a conditional agent-centric diffusion model (Fig. 1) that allows flexible traffic behavior generation. Unlike prior diffusion models for trajectories, our model is conditioned on a holistic context surrounding the vehicle including the local roadmap and neighboring agents. Since traffic has an inherent ground-truth transition function, we also enforce realistic vehicle dynamics into the design of the diffusion model state space to guarantee that generated trajectories are physically feasible.
To achieve controllable generation, Diffuser guides each step of the denoising process in a diffusion model by perturbing network outputs with the gradient of some differentiable objective to encourage desired properties. For traffic simulation, however, deriving and implementing objectives like collision avoidance, goal reaching, and road rules is complex due to its spatio-temporal and multi-agent nature. We propose to leverage the established syntax of Signal Temporal Logic (STL) . As a formal language designed for specifying spatio-temporal constraints, STL allows one to easily and scalably define driving rules; it also incorporates a notion of robustness that measures how well rules are satisfied. Concretely, we use this measure of rule satisfaction as the objective function for guiding the diffusion model (Fig. 1, right) by leveraging differentiable frameworks to make STL compatible with guidance. Since our diffusion model generates trajectories independently for each agent in a scene, we further propose a joint guidance procedure for rules involving multi-agent interactions (e.g., no collisions), which simultaneously denoises all agents in the scene to mitigate interaction rule violations.
We evaluate Controllable Traffic Generation (CTG), on the nuScenes driving dataset, demonstrating the ability to meet user constraints while maintaining realistic trajectory generation. In summary, we contribute (1) the problem formulation of controllable traffic generation, (2) a conditional diffusion-based method to generate realistic traffic satisfying physical feasibility and user-specified STL rules, and (3) extensive evaluation comparing CTG to several strong baselines, demonstrating its superiority in terms of the trade-off between controllability and realism. Demos can be found at https://aiasd.github.io/ctg.github.io.
II Related Work and Background
Traffic simulation approaches can be categorized into rule-based and learning-based. Rule-based approaches rely on analytical models such as cellular automata and intelligent driver model. These approaches typically have fixed routes for vehicles to follow and separate longitudinal and lateral motions of agents, thus having limited expressiveness for behavior simulation. Learning-based approaches mimic real-world driving behavior based on trajectory datasets . For example, TrafficSim uses a trajectory prediction model to perform scene-level traffic simulation. BITS decouples the problem into a high-level latent inference and a low-level driving behavior imitation. However, none of these methods allow a user to specify customized properties of the generated traffic behaviors at inference time.
Recent work in adversarial or safety-critical scenario generation can be seen as one instance of controllable traffic simulation, namely generating trajectories that cause an AV to misbehave in some way . STRIVE generates near-collision scenarios by searching in the latent space of a trained trajectory prediction model via a test-time optimization. Abeysirigoonawardena et al. and Chen et al. use Bayesian optimization and reinforcement learning, respectively, to generate adversarial trajectories for vehicles in a specific scenario like intersection crossing or lane changing. While these works focus specifically on adversarial objectives, our approach is general and can generate trajectories to meet several different objectives.
II-B Diffusion Modeling
Controllable diffusion models have been explored with classifier , classifier-free , and reconstruction guidance for image and video generation. Li et al. use a diffusion model with different pre-trained classifiers to guide language generation on different natural language tasks. Diffuser uses a diffusion model to plan robot behavior (state-action trajectories). Gu et al. model pedestrian trajectories for forecasting. Different from these works that use no or limited context, we adapt conditional diffusion to condition on decision-relevant information such as map and state of nearby agents. Additionally, we leverage known vehicle dynamics models to ensure physical feasibility of the generated trajectories. Diffuser also introduces test-time guidance to generate trajectories that optimize a given reward function. We build on this formulation to generate controllable traffic trajectories, but rather than learning reward functions we use analytical loss functions based on STL rules that are easy and scalable for driving applications.
II-C Signal Temporal Logic (STL)
Importantly, STL includes a formal notion of robustness, which measures how much a signal satisfies or violates a formula . We use robustness formulas as guidance functions for the proposed conditional diffusion model. In contrast to Leung and Pavone that focus on a single-controller setting and rules fixed at training time, we consider a more flexible inference-time guidance with multiple agents.
III Controllable Traffic Generation
Next, we detail our approach CTG. Sec. III-A formulates the problem of controllable traffic generation. We then describe the two stages of CTG. The offline stage trains a dynamics-enforced conditional diffusion model to capture diverse behaviors from real-world driving data (Sec. III-B). Then during online inference, CTG generates rule-compliant behaviors by sampling the model using a novel iterative joint STL guidance process (Sec. III-C). Taken together, the conditional diffusion model and STL-based guidance enable realistic and controllable generation of traffic trajectories.
For some target vehicle (tgt) that we would like to simulate, let the state at a timestep be , including 2D location, speed, and yaw. Similarly, let the action (i.e., control) be with acceleration and yaw rate. We denote to be decision-relevant context for the target agent. This consists of a local (agent-centric) semantic map and the previous states of both the target agent and its neighbors . To obtain state for vehicle at time , we assume a transition function that computes given the previous state and control . We use a unicycle dynamics model for .
III-B Conditional Diffusion for Traffic Modeling
Diffusion models pose data generation as an iterative denoising process by learning to reverse a forward diffusion process. As shown in Fig. 2, our diffusion model operates primarily on a (future) trajectory of states and actions , but is conditional as it receives the context as input at each step of denoising. Starting from Gaussian noise, the diffusion model is applied iteratively to predict a clean, denoised trajectory of states and actions.
Trajectory representation. In this section, we denote the (future) trajectory that the model operates on as:
Unlike which directly predicts states and actions jointly, our model only predicts actions and we leverage the known dynamics to infer states via rollout starting at the initial state (included in the past context). In other words, in the following formulation, always refers to a state trajectory resulting from actions, or, more formally: . This ensures physical feasibility of the state trajectory throughout the denoising process.
Formulation. Let be the action trajectory at the th diffusion step where is at the original clean trajectory. The forward diffusion process acting on is defined as:
where are a fixed variance schedule that controls the scale of the injected noise at each diffusion step. As noise is gradually added, the signal is corrupted into an isotropic Gaussian distribution.
For trajectory generation, we seek to reverse this diffusion process using a learned conditional denoising model (Fig. 2) that is iteratively applied starting from sampled noise. Given the context information , the reverse diffusion process is:
where and denotes the parameters of the diffusion model. Note that at each step, the model receives both actions and the resulting states as input. Following , the variance term of the Gaussian transition is fixed as .
Training. At each training iteration, the context and ground truth clean trajectory are sampled from a real-world driving dataset and the denoising step is uniformly sampled from . We compute the noisy input from by first corrupting the action trajectory , with , and then computing the corresponding state . The diffusion model indirectly parameterizes in Eq. 3 by instead predicting the uncorrupted trajectory where is the direct network output (see ). Finally, we use a simplified loss function to train the model:
Note that both the action and state trajectories are supervised by this loss since we found that the additional state trajectory information improves generation quality.
Implementation details. The input context containing agent-centric map information and past trajectories is represented in a rasterized format. This context is processed by a ResNet encoder before being passed to the diffusion model. Similar to Diffuser , we use a diffusion model architecture like U-Net containing several blocks of temporal 1D convolutions over the input trajectory. We incorporate conditioning information by first concatenating with the diffusion step input and then adding this conditioning feature to the convolutional features at each block of the U-Net. The diffusion process uses a cosine variance schedule and diffusion steps for all experiments.
III-C Guided Generation with Signal Temporal Logic (STL)
To enforce desired rules on realistic samples from the trained diffusion model, we introduce an iterative guidance algorithm with rules specified as STL formulas.
Conditional guidance formulation. Diffuser introduces the notion of guidance to sample trajectories from an unconditional diffusion model to meet some pre-defined objective. We can do the same for our conditional diffusion model by defining a binary random variable that indicates if a trajectory is optimal, with based on the rule satisfaction “reward” . To approximately sample from the distribution of optimal trajectories, the steps of the denoising process can be modified to :
where as in Eq. 3, and the added gradient is computed from a guide based on satisfaction:
The process of perturbing the predicted means from the diffusion model using gradients of a specified objective is summarized in Alg. 1. Different from Diffuser , an iterative inner gradient descent with clipping using the Adam optimizer is incorporated rather than using a single-step gradient update. This gives flexibility to trade off rule compliance and realism by adjusting learning rate and the number of optimization steps. When generating a future trajectory in practice, we guide several samples from the diffusion model and choose the one with the best rule satisfaction according to at the end of denoising. We refer to this as filtration.
STL as guidance. Instead of training a classifier or reward function for as in prior works , our guidance functions are implemented analytically based on STL. For each rule we wish to apply through guidance, the grammar in Eq. 1 is used to generate a corresponding STL formula describing how the trajectory should be constrained. Examples of various STL rules used are shown in Tab. I; note the formulas are relatively simple despite the complex behavior they describe. By construction, every STL formula admits a robustness formula measuring the degree of rule satisfaction, which is used as the guide . Since it is necessary to compute a gradient through this function, STL formulas are implemented using differentiable frameworks.
Multi-agent guidance. A particular challenge is applying scene-level rules that involve multiple agents (e.g., no collisions). For this purpose, guided sampling is performed in a batched fashion over all agents in the same scene simultaneously. This way guidance can be computed across all trajectories, and the corresponding gradients are collected and propagated back to agents as needed.
Simulating traffic. To perform closed-loop traffic simulation of a scene with many agents, the same model is used for each agent in a standard control loop: for each agent at each step of simulation, a guided sample is generated from the model and the first few actions are taken before re-planning at a prescribed frequency. In all experiments (Sec. IV), each scene is rolled out for 20 seconds starting from a ground truth driving log, and the re-plan rate is 2 Hz as in .
IV Experiments
We conduct experiments to validate that: (1) CTG can generate controllable traffic behaviors that satisfy user-specified rules, and (2) compared to strong baselines, CTG achieves better rule satisfaction while maintaining realism. As discussed in Sec. III-A, there is often a trade-off between realism and rule compliance; ideally a method will strike a reasonable balance and achieve good performance for both. After describing the experimental design (Sec. IV-A), we compare to baselines in single-rule (Sec. IV-B) and multi-rule (Sec. IV-C) settings both quantitatively and qualitatively, and finally conduct an ablation study (Sec. IV-D).
Datasets. nuScenes is a large-scale real-world driving dataset, which consists of 5.5 hours of accurate trajectories across two cities with diverse scenarios and dense traffic. We train all models on scenes from the train split and evaluate on 100 scenes randomly sampled from the validation split. In the current work, we focus only on vehicle simulation and defer other types (e.g., pedestrians, cyclists) to future works.
Metrics. Our evaluation focuses on controllability, realism, and stability (i.e., avoiding collisions and off-road driving). We consider rule-specific violation metrics (rule), detailed in Tab. I, to evaluate controllability. Metrics are computed for each scene ( denotes each vehicle in the scene), then averaged across all testing scenes. To evaluate realism, we follow and compare data statistics between generated traffic simulations and ground truth trajectories in the dataset. This comparison is computed via the Wasserstein distance between the normalized histograms of the driving profiles for the simulated and recorded trajectories. We define an aggregated metric – realism deviation (real) – as the mean of the realism for three properties from : longitudinal acceleration magnitude, latitudinal acceleration magnitude, and jerk. We further provide failure rate (fail) to evaluate the stability of generated trajectories. This is measured as the average fraction of agents experiencing a critical failure, i.e. collision or road departure, in a scene.
Baselines. Since there are no comparable works on rule-compliant traffic generation, we augment state-of-the-art traffic simulation models by adding a test-time optimization to meet specified rules. For a fair comparison, this optimization uses the same loss function as used for guidance in CTG. SimNet is a deterministic behavior-cloning model. We apply an optimization on its output action trajectory (SimNet+opt). TrafficSim is a CVAE-based trajectory generation method. We consider a variant with filtration (TrafficSim) using our loss function, and another with both filtration and latent space optimization (TrafficSim+opt). BITS is a bi-level imitation learning model and we adapt its sampling ranking function to use our loss function (BITS). We also use a variant that employs optimization on the output action trajectory (BITS+opt). Finally, we compare to CTG without filtration and guidance (CTG w/o f+g), i.e. random samples from the diffusion model.
IV-B Single Rule Evaluation
We first evaluate how well methods satisfy a single specified rule, which is crucial for applications such as traffic scene editing. We apply five STL rules formulated in Tab. I: speed limit, target speed, goal waypoint, no collision, and no off-road. For rules that require specific parameters to be set (e.g., the goal waypoint location), we select reasonable values based on the ground truth log in the dataset to avoid setting out-of-distribution values (e.g. off-road waypoints).
Speed Limit. Vehicles should not exceed a speed limit threshold. Since the speed limit of each road is not available in the dataset, we set the limit per scene to be the speed at the quantile of all moving vehicles in that scene.
Target Speed. Vehicles should follow a specified speed at each time step. For each vehicle, the speed is set to of its speed in the ground truth scene, similar to a traffic jam.
Goal Waypoint. Vehicles should reach a specified waypoint at any time in the future. We set waypoints to be the position at s along the ground truth data trajectory for each vehicle. Hitting a waypoint from ground truth data is not trivial since sampled trajectories often greatly deviate from the dataset.
No Collision. Vehicles should not collide with each other.
No Off-road. Vehicles should not leave the drivable area.
Quanitative results are shown in Table II. In general, CTG achieves lower values for rule violation, realism deviation, and failure rate than the baselines. Among all five settings, CTG has the lowest rule violation in three and is competitive in the others. For realism deviation and failure rate, CTG is usually top two. Fig. 3 shows qualitative results for target speed and waypoint rules. Compared to TrafficSim+opt and BC+opt, the strongest baselines for target speed, CTG has the lowest rule violation using more realistic trajectories. For the waypoint example, although BC+opt provides better rule satisfaction than CTG, both BC+opt and BITS+opt predict curvy, unrealistic trajectories resulting in multiple collisions.
IV-C Multiple Rules Evaluation
Stop Sign and No Off-road. Vehicles should stop if they enter a stop sign region and not go off-road. The expression for this rule (Tab. I) is relatively involved using “Implies” and “Eventually” operators, making it a good test for our STL-based approach. Stop regions are m boxes with centers set to be s along the ground truth data trajectories.
Goal Waypoint and Target Speed. Vehicles should reach their goal following the specified target speeds. Waypoints are set s along the ground truth trajectories, and the target speeds at each time step are the same as the ground truth scene. This setting is similar to a “reactive replay” use case in AV testing, where a user reconstructs a driving log using the traffic model, which allows agents to realistically react to any subsequent changes to AV behavior or environment.
Results are shown in Table III. For both settings, CTG variants achieve the top two lowest rule violation and realism deviation, with only slightly higher failure rates.
IV-D Ablation Study.
To analyze design choices, we conduct an ablation study under the speed limit rule setting. Results are shown in Table IV where the top row is our proposed version of CTG. The first section compares different combinations of guidance (g) and fitration (f). Without guidance, rule violation increases greatly. Filtration is more effective paired with guidance than by itself, as it can choose the best from several already-guided samples that may satisfy rules to differing degrees. In the next part of the table, variants using an additional output action optimization (a) are evaluated. Replacing guidance with optimization is worse on all metrics, while combining guidance with optimization reduces rule violation at the cost of higher realism deviation and failure rate. Next, more inner optimization steps (op) are used in guidance, improving rule violation but giving worse realism and failure rate. Finally, in the bottom section, we evaluate a variant that only supervises the action trajectory, and variants where unicycle dynamics (dyn) are not enforced. Supervising only actions (instead of states and actions) gives more faithful accelerations resulting in lower realism deviation, but failure is more frequent without state supervision to regularize. Additionally, we find that enforcing dynamics using the unicycle model is key: all metrics degrade without this.
V Conclusion
We proposed CTG, a conditional diffusion model for the task of controllable traffic simulation, which opens several exciting future research directions. Currently, we have only used CTG to model vehicles, but cyclists and pedestrians are also important agents to simulate for AV interactions. Additionally, using collision and off-road guidance to enable very long-term, robust traffic simulation is an important application. Outside of AV, the proposed guidance framework may benefit many tasks where learned models in-the-loop must be reactive and follow novel objectives online.
Acknowledgments. The authors thank Or Litany, Sanja Fidler, and Karen Leung for valuable discussions and feedback.