Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following

Brian Yang, Huangyuan Su, Nikolaos Gkanatsios, Tsung-Wei Ke, Ayush Jain, Jeff Schneider, Katerina Fragkiadaki

Introduction

Diffusion models have shown to excel at modeling highly complex and multimodal trajectory distributions for decision-making and control . Reward-gradient guidance has been used to test-time optimize differentiable reward functions by alternating between denoising diffusion steps and backpropagating reward gradients to the noised trajectory. In this way, sampled trajectories are pushed towards the trajectory data manifold while also maximizing the reward function at hand . This decoupling of the reward function from trajectory diffusion permits a single trajectory diffusion model to be used for maximizing a variety of reward functions at test time. Reward-gradient guidance requires the reward function to be differentiable and fitted in both noisy and clean trajectories, which usually requires re-training. This limits its applicability as a general solver for trajectory optimization.

We propose Diffusion-ES, a reward-guided denoising method for optimization of non-differentiable, black-box objectives that samples and mutates trajectories using a diffusion model, guided by a reward function that operates only on the clean, final, denoised samples. Naively combining diffusion with sampling-based optimization does not work: sampling-based optimizers, like CEM or MPPI , typically require a large population of samples across multiple iterations of selection and mutation to converge to good solutions, which, when combined with the computational cost of denoising inference, results in a prohibitively slow search process. In Diffusion-ES, high-scoring trajectories are mutated using a truncated diffusion-denoising process, by adding a small amount of noise and denoising them back, as shown in Figure 2 right. The amount of added noise is progressively decreased across search iterations making Diffusion-ES computationally viable.

The trajectory diffusion model used for test-time optimization in Diffusion-ES can in principle condition on any scene-relevant information to narrow the sampling to a distribution of scene-relevant trajectories. In fact, the amount of conditioning information controls a continuum between train-then-test learning and test-time planning, and a corresponding trade-off between inference speed and out-of-distribution (OOD) generalization: 1. The more conditioning information, the narrower the distribution to draw trajectory samples from, the faster the search. In the extreme, no test-time reward optimization is used and our diffusion model operates as a reactive policy at test time. Indeed, many methods train diffusion policies as conditional diffusion trajectory prediction models using rewards as conditioning information to the trajectory diffusion model or finetune a diffusion policy with reinforcement learning or imitation learning . While these methods can handle black-box reward functions or good behaviours to imitate, we show that using them as is or with reward-guidance at test time often under-performs reward-guided denoising of an unconditional diffusion model, which completely decouples trajectory and reward modelling. 2. The less conditioning information, the wider the distribution to draw trajectory samples from, the slower the search, but the better the generalization to OOD tasks and scenarios that require novel pairings of trajectories and scene contexts, not present in the training data. Indeed, this is the premise of test-time planning over train-then-test learning: test-time optimization of a composition of energy functions , here, the energy of the trajectory data distribution and the energy of arbitrary reward functions, should be able to synthesize novel behaviours not seen at training time.

We show how Diffusion-ES, when combined with an unconditional diffusion model over trajectories, can achieve state-of-the-art planning performance purely through diffusion-guided black-box reward maximization. Our approach is evaluated on nuPlan , an established driving benchmark built on real-driving logs and estimated ground-truth perception. We achieve state-of-the-art performance for closed-loop driving, matching the performance of the previous SOTA, PDM-Closed, a sampling-based planner tailored to the nuPlan benchmark , as well as reactive driving policies , deterministic or diffusion-based. Moreover, we illustrate the flexibility of Diffusion-ES by test-time optimizing language-shaped reward functions generated using few-shot LLM prompting. Using language instructions, we can solve the most challenging nuPlan scenarios, as well as synthesize entirely novel driving behaviors. We then test our model and baselines in their ability to optimize the generated reward functions to elicit the desired behaviours. Qualitative examples of behaviors generated by instruction following using our method can be found in Figure 1. We show Diffusion-ES dramatically outperforms PDM-Closed , other sampling-based planners, as well as ablative versions of Diffusion-ES that either condition the diffusion model on the surrounding scene, or do not use any guidance at all.

In summary, our contributions are as follows:

We introduce Diffusion-ES, a trajectory optimization method for optimizing black-box objectives that uses a trajectory diffusion model for sampling and mutating trajectory proposals during sampling-based search. We show Diffusion-ES matches the SOTA performance of engineered planners in closed-loop driving in nuPlan, and much outperforms them when optimizing more complex reward functions that require flexible driving behaviour, beyond lane following. To the best of our knowledge this is the first work to combine evolutionary search with diffusion models.

We show that Diffusion-ES can be used to follow language instructions and steer the closed-loop driving behaviour of an autonomous vehicle by optimizing the LLM-shaped reward functions, without any training data of language and actions. We showed that such instruction following can solve the most challenging driving scenarios in nuPlan.

We show extensive ablations of our model with varying amount of conditioning information which clearly reveals the trade off between inference speed and out-of-distribution (OOD) generalization in driving.

We believe Diffusion-ES will be useful to the community as a general trajectory optimizer with applicability beyond driving. Our code and models will be publicly available upon publication to aid reproducibility in the project webpage: diffusion-es.github.io.

Related work

Diffusion models for decision-making and trajectory optimization Diffusion models learn to approximate the data distribution through an iterative denoising process and have shown impressive results on image generation . They have been used for imitation learning for manipulation tasks , for controllable vehicle motion generation and for video generation of manipulation tasks . Works of use diffusion models to forecast offline vehicle trajectories. To the best of our knowledge, this is the first work to use diffusion models in closed-loop driving.

Learning versus planning for autonomous driving Learning to drive from imitating driving demonstrations is prevalent in the research and development of autonomous vehicles . Many preeminent imitation methods assume the underlying action distribution is unimodal, which is problematic when training from multimodal expert demonstrations. Objectives and architectures that can better handle multimodal trajectory prediction have been proposed . We show that diffusion models are well-suited for driving and can be used to synthesize rich complex behaviors from multimodal demonstrations.

On the other hand, conventional autonomy stacks do not rely on learning at all for decision making, and rather rely on optimizing manually engineered cost functions online . Recently, PDM-Closed achieved state-of-the-art performance on the nuPlan driving benchmark by purely relying on test-time planning and heuristics for selecting trajectory proposals. Other prior works aim to incorporate the benefits of offline learning for test-time planning by performing sampling-based planning over learned cost maps or doing gradient-based optimization over learned dynamics models . We extend this line of work by showing how diffusion-based generative models can be combined with sampling-based planning.

Language-conditioned policies for autonomous driving Recently there has been significant progress made towards language-conditioned policies for driving. GAIA-1 is a generative world model capable of multimodal video generation that leverages video, language and actions to synthesize driving scenarios which can comply with given language instructions. However, GAIA-1 does not execute any actual control inputs.

LLMs trained from Internet-scale text have shown impressive zero-shot reasoning capabilities for a variety of downstream language tasks when prompted appropriately, without any weight fine-tuning . Recent works have shown that LLMs can be prompted to map language instructions to language subgoals action programs or cost maps with appropriate plan-like or program-like prompts. Our work follows few-shot prompting of LLMs to shape driving reward functions. We extend previous methods by using Python generators to produce reward functions which maintain internal state across calls. Works of use LLMs to predict low-level control signals given high-level scene descriptions and language instructions. However, none of these approaches evaluate their performance on closed-loop driving, which is significantly more difficult than open-loop trajectory forecasting . is similar to us, but only considers a highly simplified driving setup and does not report results on a standardized benchmark with strong baselines.

Method

Diffusion models A diffusion model captures the probability distribution p(x)p(x) through the inversion of a forward diffusion process, that gradually adds Gaussian noise to the intermediate distribution of a initial sample xx. The amounts of added noise depend on a predefined variance schedule βt∈(0,1)t=1T{\beta_{t}\in(0,1)}_{t=1}^{T}, where TT denotes the total number of diffusion timesteps. At diffusion timestep tt, the forward diffusion process adds noise into xx using the formula xt=αˉtx+1−αˉtϵx_{t}=\sqrt{\bar{\alpha}_{t}}x+\sqrt{1-\bar{\alpha}_{t}}\epsilon, where ϵ∼N(0,1)\epsilon\sim\mathcal{N}(\mathbf{0},\mathbf{1}) is a sample from a Gaussian distribution with the same dimensionality as xx. Here, αt=1−βt\alpha_{t}=1-\beta_{t}, and αˉt=∏i=1tαi\bar{\alpha}t=\prod{i=1}^{t}\alpha_{i}. For denoising, a neural network ϵ^=ϵθ(xt;t)\hat{\epsilon}=\epsilon_{\theta}(x_{t};t) takes input as the noisy sample xtx_{t} and the diffusion timestep tt, and learns to predict the added noise ϵ\epsilon. To generate a sample from the learned distribution pθ(x)p_{\theta}(x), we start by drawing a sample from the prior distribution xT∼N(0,1)x_{T}\sim\mathcal{N}(\mathbf{0},\mathbf{1}) and iteratively denoise this sample TT times with ϵθ\epsilon_{\theta}. The application of ϵθ\epsilon_{\theta} depends on a specified sampling schedule , which terminates with x0x_{0} sampled from pθ(x)p_{\theta}(x). Diffusion models can be easily extended to model p(x∣c)p(x|\textbf{c}), where c is some conditioning signal, such as the expected future rewards , by adding an additional input to the denoising neural network ϵθ\epsilon_{\theta}.

2 Diffusion-ES

Diffusion-ES is a trajectory optimization method that leverages gradient-free evolutionary search to perform reward-guided sampling from trained diffusion models, for any black-box reward function R(x)R(x). Specifically, we use a trained diffusion model ϵθ\epsilon_{\theta} to initialize the sample population and we use a truncated diffuse-denoise process to mutate samples while staying in the data manifold. The control flow of Diffusion-ES is shown in Figure 2 (left) and in Algorithm 1.

Initializing the population with diffusion sampling We begin by sampling an initial population X0X^{0} of MM trajectory samples using our diffusion model:

where XkX^{k} is the population at iteration kk. This involves a complete pass through the reverse diffusion process. Given we use an unconditional diffusion model, these samples are scene agnostic and can always be used without re-sampling them at each timestep. We can also modify the initial population by including samples generated by other approaches or mixing in solutions from the previous timestep to warm-start our optimization.

Sample scoring At each iteration kk, we score the samples in our population {R(xi)∣xi∈Xk}i=1M\{R(x_{i})|x_{i}\in X^{k}\}_{i=1}^{M}. Note that our population consists of “clean” samples so we do not need a reward function which can handle ”noisy” samples, which gives us significant flexibility compared to guidance methods that perform classifier-based guidance.

Selection We use rewards to decide which samples we should select to propagate to the next iteration. Similar to MPPI , we resample Xk+1X^{k+1} as follows:

where Ek+1E^{k+1} represents our elite set which is kept from iteration kk, and τ\tau is a tunable temperature parameter controlling the sharpness of qq.

Mutation using truncated diffusion-denoising We apply randomized mutations to Ek+1E^{k+1} for exploration. Prior evolutionary search methods resort to naive Gaussian perturbations which do not exploit any prior knowledge about the data manifold. Our key insight is to leverage a truncated diffusion-denoising process to mutate trajectories in a way the resulting mutations are part of the data manifold. We can run the first tt steps of the forward diffusion process to get noised elite samples Eˉk+1\bar{E}^{k+1}:

where ϵ∼N(0,1)\epsilon\sim\mathcal{N}(\mathbf{0},\mathbf{1}). Then we can run the last tt steps of the reverse diffusion process to denoise the samples again, giving us clean samples Xk+1X^{k+1}:

In practice, the number of timesteps tt of the truncated diffusion process is a tunable time-dependent hyperparameter tkt_{k} which controls the mutation strength at each iteration kk. This is visualized in Figure 2 (right). We find that linearly decaying the number of mutation diffusion steps tkt_{k} from 5 to 1 over 20 search steps works best in our experiments.

3 Mapping language instructions to reward functions with LLM prompting

To follow driving instructions given in natural language, we map them to black-box reward functions which we optimize with Diffusion-ES. We adopt a similar approach to which uses LLMs to synthesize reward functions from language instructions. Reward functions are composable and allow us to seamlessly combine language guidance with other constraints. This is crucial in driving where we constantly optimize many different objectives at once (e.g., safety, driver comfort, route adherence).

Similar to prior work, we expose a Python API which can be used to extract information about entities in the road scene. Since many of the basic reward signals in driving do not change from scenario to scenario (e.g., collision avoidance or drivable area compliance), we allow the LLM to write reward shaping code which modifies the behavior of the base reward function, as opposed to generating everything from scratch. Reward shaping can add auxiliary reward terms (e.g., a dense lane-reaching reward) or re-weight existing reward terms. We show generated code examples in Figure 3.

Our goal is to handle general and complex language instructions with temporal dependencies, such as “Change lanes to the left, then pass the car on the right, then take the exit”. Previous works produce a stationary reward function, i.e., one which is fixed during planning. This can make it challenging to express sequential plans solely through rewards. We find that a much more natural and succinct way of capturing these plans in code is through the use of generator functions which retain internal state between calls. All our prompts and code examples can be found in the supplementary file. In Section 4, we show how optimizing these language-shaped reward functions can synthesize rich and complex driving behaviors that comply with the language instructions.

Experiments

We first evaluate Diffusion-ES on closed-loop driving in nuPlan , an established benchmark that uses estimated perception for vehicles, pedestrians, lanes and traffic signs. Our model and baselines are evaluated on their ability to drive safely and efficiently while having access to close-to-ground-truth perception output provided by the dataset. We also consider a suite of driving instruction following tasks. We map instructions to shaped reward functions with LLM prompting, and evaluate Diffusion-ES and baselines in their ability to optimize the generated reward functions and accurately follow the instructions. Our experiments aim to answer the following questions:

How does Diffusion-ES compare to existing sampling-based planners and reward-gradient guidance for trajectory optimization?

How does Diffusion-ES compare to SOTA reactive driving policies that directly map environments’ state to vehicle trajectories?

Can the hardest nuPlan driving scenarios be solved by assuming access to a human teacher giving instructions in natural language, without any additional training data?

Does scene conditioning for the diffusion model benefit Diffusion-ES?

We evaluate our model on the nuPlan Val14 planning benchmark . We specifically consider the reactive agent track of the nuPlan benchmark since it is the most difficult and realistic of the evaluation settings in nuPlan.

For Diffusion-ES, we train a diffusion model over ego-vehicle trajectories consisting of 2D poses (x,y,θ)(x,y,\theta) predicted 8 seconds into the future at 2Hz, leading to an overall action dimension of 48. Unlike prior work , we model the distribution over actions only rather than modeling states and actions. We use a population size of M=128M=128 in our experiments. Our diffusion model is trained with T=100T=100 denoising steps.

Reward function

We adopt a modified version of the scoring function used in PDM-Closed as our reward function. To compute rewards, we convert our predicted trajectory to low-level control inputs using an LQR tracker. These control inputs are fed into a kinematic bicycle model which propagates the dynamics of the ego-vehicle. We follow and forecast the motion of other agents by assuming constant velocity. These simulated rollouts are then scored following the nuPlan benchmark evaluation metrics. We also add auxiliary reward terms to penalize proximity to the leading agent and enforce speed limits.

Note that this reward function is not differentiable due to the tracker and the use of non-differentiable heuristics for assessing traffic violations. Additionally, training a model to regress rewards is challenging since the nuPlan dataset contains no instances of serious traffic infractions.

Evaluation metric

We report our results using driving score, which aggregates multiple planning metrics related to traffic rule compliance, safety, route progress, and rider comfort. This is the standard evaluation metric used in nuPlan.

Baselines

UrbanDriverOL , a deterministic transformer policy trained with behaviour cloning and augmentations. Unlike in , the nuPlan implementation does not perform closed-loop training.

PlanCNN a deterministic imitation policy which encodes a rasterized BEV map using a CNN backbone.

IDM : a heuristic rule-based planner which adjusts its speed to maintain a safe distance to the leading vehicle. It is also used to control the behaviour of agents in nuPlan.

PDM-Closed : an MPC-based planner which generates path proposals using lane centerlines, and rolls out trajectories similarly to us. Instead of iteratively optimizing rewards, PDM-Closed simply executes the highest-performing proposal after one round of scoring. It is the current state-of-the-art on the nuPlan Val14 benchmark.

Diffusion Policy: a diffusion model we consider that conditions on scene features to predict a vehicle trajectory directly. We encode scene features using the transformer feature backbone from Urban Driver . We train it with imitation learning and augmentations. This is similar to the unconditional trajectory model in Diffusion-ES with additional conditioning on scene features.

We show quantitative results in Table 1. We draw the following conclusions:

1. Diffusion-ES matches the prior state-of-the-art, PDM-Closed and substantially outperforms all other baselines. PDM-Closed is a sampling-based planner that relies on domain-specific heuristics to generate trajectory proposals whereas Diffusion-ES learns these proposals from data. Both use a similar reward function and dynamics model for the agents in the scene.

2. There is a large gap in performance between reactive neural policies and test-time planners, also pointed out in recent work . We hypothesize that this is because compared to other control benchmarks, nuPlan has a much richer observation space as scenes are densely populated by dynamic actors, many of which are irrelevant to the ego-agent. This can make it challenging for learning-based methods to generalize, which motivates the need of test-time optimization.

3. Diffusion-ES substantially outperforms diffusion policy. Qualitatively, our diffusion policy has a tendency to randomly change lanes, which causes the ego-vehicle to reach out-of-distribution scenarios faster. Diffusion-ES leverages the expressiveness of generative modeling while using test-time optimization to improve generalization.

2 Language instruction following

One drawback of the nuPlan driving benchmark is that encourages highly conservative driving behaviors. For example, PDM-Closed holds the state-of-the-art in the nuPlan Val14 benchmark while being unable to change lanes, since its path proposals only consider the lane the ego-vehicle is currently on. However, lane changing is not necessary for good driving performance in the current nuPlan benchmark.

To evaluate Diffusion-ES and baselines in their ability to optimize arbitrary reward functions, we consider eight language instruction following tasks, each taken from an existing driving log in the nuPlan benchmark. In each task, the language instruction requires the ego-vehicle to perform a specific driving maneuver that solves a challenging driving scenario. In most scenarios there will be no examples of the instructed behavior anywhere in nuPlan. For instance, the lane weaving task requires the ego-vehicle to aggressively change multiple lanes in dense urban traffic. Task descriptions, language instructions, and prompts are provided in the appendix. We use the method described in 3 to generate executable Python code given a language instruction that adapts the initial reward function of Section 4.1, giving us a language-shaped reward function for each scenario. Our model and baselines will optimize the same language-shaped reward function.

We evaluate our model and baselines using task success rate, which measures how frequently the agent was able to successfully complete the designated task. To increase the difficulty of the tasks, we randomize the behavior of other vehicle agents by adding noise to their IDM parameters at sporadic intervals during each episode. All scores reported are averaged across ten random seeds.

Baselines

PDM-Closed: this is adapted to this setting by using our language-shaped reward function in place of the original reward function.

PDM-Closed-Multilane: a modified variant of PDM-Closed which considers a wider range of laterally offset paths, allowing for lane changes.

Conditional Diffusion-ES: Diffusion-ES that uses a conditional diffusion model instead of an unconditional one.

Figure 4 shows the success rates on the controllability tasks. We draw the following conclusions:

1. Diffusion-ES outperforms all baselines. Although PDM-Closed-Multilane has substantially improved performance compared to PDM-Closed due to more diverse trajectory proposals, it is still weaker than Diffusion-ES on 4 out of 8 tasks. This highlights the weakness of relying on handcrafted rules for proposal generation. 2. Diffusion-ES performs significantly worse with a conditional diffusion model. The conditional diffusion model is much harder to guide since the scene context causes fewer samples to be in-distribution.

3 Lane following

To compare Diffusion-ES against reward-gradient guidance, we consider a simplified lane following task with a differentiable reward function, which consists of two terms: one penalizing lateral deviation from the lane (lane error), and one penalizing deviation from a target speed (speed error). We sample 14 scenarios, one of each scenario type in nuPlan, and report average planning costs across all scenarios.

CEM : a widely used ES method which parameterizes the search distribution qq as a Gaussian. CEM iterates between sampling from qq and fitting qq to the best samples.

MPPI : similar to CEM but rather than keep a fixed number of elites, MPPI samples proportional to rewards.

Reward-gradient guidance : we directly optimize the ground-truth planning objective with gradient descent during the denoising process.

We show quantitative results in Table 2. We draw the following conclusions: 1.Diffusion-ES outperforms the differentiable reward-gradient guidance baseline even though the objective is differentiable. We hypothesize that this is because although the ground truth reward function is available, it may not provide suitable guidance for intermediate noisy trajectories. This highlights a key advantage of our method over prior work, which is that we can optimize novel objectives without needing to train a reward regressor on noisy samples. 2. Both diffusion-based methods significantly outperform sampling-based planners that do not leverage diffusion. This is consistent with our hypothesis that diffusion guidance can optimize trajectories much more efficiently than conventional ES methods. Videos of our method driving in all experimental settings can be found in our project page diffusion-es.github.io.

4 Runtime analysis

Diffusion-ES can be used in real-time with some minor optimizations. By using fewer diffusion steps T=10T=10, smaller population size M=32M=32 and less iterations K=2K=2, our method can be run at the same frequency as the simulator (2 Hz) at a small cost to performance (nuPlan driving score drops from 92 to 91). We report the average wallclock time for inference at every timestep over 100 trials in Table 3.

5 Discussion - Limitations - Future work

As seen in Section 4.4, our approach does introduce computational overhead. We believe that these issues can be mitigated by incorporating recent advances in diffusion modeling such as faster samplers. Our reward function assumes other agents will travel at constant velocity, which could clearly be improved. This also assumes that other agents cannot react to the ego-vehicle, which has been shown to be a major limitation for planners in self-driving . However, our instruction following experiments suggest that even if we cannot ever forecast perfectly, we can use language-shaped rewards to solve the hardest driving scenarios. We aim to explore memory-prompted analogical reward shaping for handling long-tail scenarios without a human teacher in our future work.

Conclusion

We presented Diffusion-ES, a method for black-box reward guided diffusion sampling. We showed that Diffusion-ES can effectively optimize reward functions in nuPlan for driving and instruction following, and outperforms engineered sampling-based planners, reactive deterministic or diffusion policies, as well as differentiable reward-gradient guidance. We showed how our method can be used to follow language instructions without any language-action trajectory data, simply using LLM prompting to generate shaped reward maps for test-time optimization. Our future work will explore retrieving the right reward shaping to optimize to handle long-tailed driving scenarios in the absence of human teachers. Our experiments show the trade-off between inference speed and OOD generalization during scene conditioning in diffusion policies: more scene conditioning throws the model out-of-distribution, less conditioning enables a more effective test-time search which requires more time to complete. Our future work will explore ways to amortize the result of such searches to fast reactive policies, and to device a continuum among the two extremes, so that a variable amount of compute would be spent depending on the difficulty of the scenario at hand.

References

Appendix

We train our diffusion models on 1% of the released nuPlan training data, which is subsampled from the original 20Hz to 0.5Hz. Our diffusion models are DDIMs trained with T=100T=100 diffusion steps. We use the scaled linear beta schedule and predict ϵ\epsilon. Our models are implemented in PyTorch and we use the HuggingFace Diffusers library to implement our diffusion models.

Our base diffusion architecture is as follows. Each trajectory waypoint is linearly projected to a latent feature with hidden size 256. The noise level is encoded with sinusoidal positional embeddings followed by a 2-layer MLP. Noise features are fused with the trajectory tokens by concatenating the noise feature to all trajectory features along the feature dimension and projecting back to the hidden size of 256. We also apply rotary positional embeddings to the trajectory tokens as temporal embeddings. We then pass all the trajectory tokens through 8 transformer encoder layers, and each trajectory token is decoded to a corresponding waypoint. The final trajectory consists of the stacked waypoint predictions. Our conditional diffusion policy baseline uses a similar architecture, except we featurize the scene using the backbone from the nuPlan re-implementation of Urban Driver and pass those tokens into the self-attention layers.

We train our models with batch size 256 and use the AdamW optimizer with learning rate 1e-4, weight decay 5e-4, and (β1,β2)=(0.9,0.999)(\beta_{1},\beta_{2})=(0.9,0.999).

Our trajectories consist of 16 2D pose waypoints each with 3 features (x,y,θx,y,\theta). We preprocess these trajectory features by applying Verlet wrapping as described in MotionLM , which we found to improve performance by encouraging smooth trajectories.

2 Language instruction following tasks

Here, we describe in detail each of the controllability tasks. For each, we list the task goal as well as the specific language instruction used.

Lane change: the ego-vehicle must execute a lane change. The language instruction is ”Change lanes to the left”. The episode is considered a success if the ego-vehicle reaches the left lane.

Unprotected left turn: the ego-vehicle must perform an unprotected left turn. The language instruction is ”If car 18 is within 20 meters yield to it. Otherwise it will slow for you”, where car 18 is the incoming car. The episode is considered a success if the ego-vehicle either completes the turn before the incoming car, or the incoming car passes the ego freely indicating a successful yield. Due to the randomized agent behaviors, it is not always possible to safely execute the turn in this task.

Unprotected right turn: the ego-vehicle must perform an unprotected right turn. The language instruction is ”Change to lane 33.”, where lane 33 is the target lane. The episode is considered a success if the ego-vehicle completes the right turn.

Overtaking: the ego-vehicle must overtake the target car. The language instruction is ”Car 21 will slow for you. Change to the right lane. Once ahead of car 4 change to the left lane.”, where car 21 is the incoming car in the right lane and car 4 is the target car. The episode is considered a success if the ego-vehicle is ahead of the target car while in the same lane.

Extended overtaking: the ego-vehicle must overtake the target car across several lanes of dense traffic. The language instruction is ”Change two lanes to the left. Then if you are ever ahead of car 3 change lanes to the right”. The episode is considered a success if the ego-vehicle is ahead of the target car while in the same lane.

Yielding: the ego-vehicle must allow a car approaching quickly from behind to pass by changing lanes. The language instruction is ”Slow down and change lanes to the left. Then once car 8 is ahead of you change lanes to the right.”, where car 8 is the incoming car. The episode is considered a success if the ego-vehicle is behind the target car.

Cut in: the ego-vehicle must cut in to a column of cars. The language instruction is ”Slow down. Car 2 will slow for you. Change two lanes to the left”, where car 2 is the car we are cutting in front of. The episode is considered a success if the ego-vehicle is in front of the target car and in the same lane.

Lane weaving: the ego-vehicle must reach a specific gap between two cars across several lanes of dense traffic. The language instruction is ”Once ahead of car 12 change lanes to the left. Then slow down a lot and change lanes to the left. Once car 9 is ahead of you by a few meters change lanes to the left”. The episode is considered a success if the ego-vehicle successfully reaches the target gap.

3 Language instruction following prompts

We collect 24 instruction-program pairs in total, and they are shown in full below. To organize and manage our prompts, we use DSPy .