End-to-End Driving with Online Trajectory Evaluation via BEV World Model
Yingyan Li, Yuqi Wang, Yang Liu, Jiawei He, Lue Fan, Zhaoxiang Zhang
Introduction
End-to-end autonomous driving has achieved remarkable success in recent years and garnered widespread attention. While most approaches primarily focus on predicting a high-quality trajectory, recent works suggest that predicting multiple trajectories simultaneously better reflects the multi-modal nature of driving and leads to improved performance. Given these multiple possible trajectories, it becomes crucial to effectively evaluate them in order to ensure the safety and reliability of end-to-end autonomous driving systems.
Traditional trajectory evaluation methods are rule-based and rely on perception results, such as bounding boxes and maps, to assess trajectories. These methods are sensitive to perception inaccuracies and may not be directly optimized in an end-to-end manner. Recently, some end-to-end driving approaches have incorporated trajectory evaluation. Their evaluations rely on the current state. However, evaluating a trajectory candidate is particularly challenging without knowledge of the future led by this trajectory. Similar to human drivers, who anticipate futures before making decisions, a trajectory evaluation network would benefit from predicting the corresponding future states of trajectory candidates.
To effectively predict future states, it is intuitive to adopt methodologies from reinforcement learning, where world models have successfully facilitated future state predictions. In reinforcement learning, model-based methods have consistently shown superior performance compared to model-free approaches , largely due to their ability of modeling the environmental dynamics and reasoning about future scenarios through world models. Inspired by this success, we propose adopting a world model within end-to-end autonomous driving frameworks to facilitate accurate prediction of future states, thus significantly improving trajectory evaluation.
However, leveraging the world model for trajectory evaluation presents some technical issues. First, a good representation of future scenarios is needed. Many existing world models for driving predict the images of future scenarios with diffusion models , which can be time-consuming and not suitable for real-time driving. Although another line of work for driving simulation can render free-viewpoint future videos in real-time, they need off-board reconstruction and cannot be adopted in our onboard driving setting. Second, we lack supervision for future states. Supervising the predicted future states of the world model is difficult because world models need to imagine multiple future states based on the multiple trajectory candidates, while only one future state is available in the real-world datasets.
To address these technical issues, we propose WoTE as shown in Fig.1. We propose to predict future states in the BEV space. The information in BEV space is much more compact than in raw images, which enables a more efficient single-step feed-forward future prediction than the time-consuming multi-step denoising in previous driving world models . This characteristic of BEV space makes WoTE highly suitable for real-time driving applications. Additionally, the predicted future BEV states can be easily supervised using off-the-shelf BEV-space traffic simulators such as nuPlan . These simulators model the scenario in BEV space by representing the map and objects using BEV semantic maps, creating a dynamic simulation of the driving environment. Consequently, we can use the BEV semantic maps from these simulated scenarios as supervision for the prediction of the BEV world model. Moreover, these simulators are equipped with well-established evaluation rules to assess the simulated future BEV scenarios, which provide target rewards.
We validate our framework on both NAVSIM benchmark and closed-loop Bench2Drive benchmark based on the CARLA simulator and achieve state-of-the-art performance. Our experiments show that incorporating future states into trajectory evaluation is essential for improved performance. Additionally, the trajectory evaluation module, built on the world model shows great generalization ability across diverse trajectories. Our contributions are as follows:
We highlight the importance of utilizing imagined future states for trajectory evaluation in end-to-end autonomous driving.
We introduce WoTE, which consists of an end-to-end trajectory evaluation module based on a BEV world model. The world model in BEV space enables efficient real-time future prediction and overcomes the scarcity of multi-future supervisions.
WoTE is validated on both the NAVSIM benchmark and the closed-loop CARLA-based Bench2Drive benchmark, achieving state-of-the-art performance.
Related Works
End-to-end autonomous driving is to train a model that directly maps sensor inputs to trajectories or control signals instead of decomposing the driving task into traditional sub-tasks. These works can be divided into two categories according to the training paradigm: imitation-learning-based and reinforcement-learning-based. Imitation learning is the most common practice in end-to-end autonomous driving. Their predicted trajectories are supervised by the expert trajectories. For instance, P3 , takes a differentiable semantic occupancy as the intermediate representation and learns actions from human trajectories. Following this, ST-P3 learns spatial-temporal features for perception, prediction and planning tasks simultaneously. UniAD proposed goal-oriented planning, combining and leveraging advantages of perception and motion prediction modules. VAD introduces vectorized representations for end-to-end planning. There are also some reinforcement-learning-based methods. For example, MaRLn utilizes implicit affordances to develop a reinforcement learning algorithm. LBC trained a vision-based sensorimotor agent from the privileged agent. Combining imitation learning and reinforcement learning, Hydra-MDP utilizes knowledge distillation to obtain supervisions from both rule-based planners and human drivers. It is a promising research direction as more supervisions are utilized.
2 Trajectory Evaluation
Trajectory evaluation is essential for ensuring the reliability and safety of autonomous driving systems. We divide the main trajectory evaluation methods into two categories: model-free and model-based. Model-free methods directly assess the learned policy using observed data without explicitly modeling environmental dynamics. For instance, Monte Carlo policy evaluation uses the empirical mean return as the metric. The value function approximates the expectation based on sampled returns. Temporal-difference learning uses bootstrapping, where the current estimate of the value function is used to generate target values for updating the value function. Subsequently, model-based trajectory evaluation methods have gained increasing popularity in recent years. By leveraging learned world models that capture environmental dynamics, these methods enable more accurate and predictive policy evaluation. For example, PPUU measures policy performance by quantifying uncertainty in predicted future states. Rails evaluates trajectories through a tabular dynamic programming mechanism, and MARL-CCE employs learned environmental dynamics to assess multi-agent policies in driving scenarios. However, these methods often rely on explicit trajectory representations and non-differentiable metrics, which limits their end-to-end optimization capabilities. In contrast, our approach evaluates trajectories using BEV features and encoded trajectory embeddings, enabling fully differentiable and thus end-to-end trajectory evaluation.
Method
In this section, we introduce four key components of WoTE: Trajecoty Predictor (Sec.˜3.1), BEV World Model (Sec.˜3.2), Reward Model (Sec.˜3.3) and BEV Space Supervision (Sec.˜3.4). As illustrated in Figure 2, our method integrates future state prediction via the BEV World Model and reward prediction via the Reward Model within a unified framework. Specifically, the Trajecoty Predictor encodes the multi-modal observations into BEV state, represented as a BEV feature map, and provides multiple trajectory proposals. The BEV World Model predicts corresponding future BEV states based on these trajectory proposals. With the help of future states, the Reward Model predicts the reward of each trajectory. Finally, we choose the trajectory with the highest reward as the final trajectory.
As shown in Figure 2 (a), the Trajectory Predictor aims to predict multiple trajectory candidates. First, the BEV encoder encodes the multi-modal sensor inputs into a unified BEV feature map. Then, given trajectory anchors, the trajectory refinement module refines trajectories based on the BEV feature map. We will now describe each part in detail.
Trajectory Anchors.
Following methods in , we generate trajectory anchors by K-Means clustering. This clustering is performed on all the expert trajectories collected from the training dataset. The number of trajectory anchors is denoted as .
Trajectory Refinement.
As illustrated in Figure 2 (a), the trajectory refinement module takes the trajectory anchors and the current BEV states as inputs to predict the refined trajectories. Specifically, we first use a trajectory encoder , which is implemented by an MLP, to encode the trajectory anchors into feature vectors. Then, these feature vectors are used as queries to perform cross attention with the current BEV state , which serves as both the key and value. The output of the cross-attention is then fed into an MLP head to predict the offsets. The offset is added to the trajectory anchors to produce the refined trajectories. The entire process can be expressed as
where is the refined trajectories.
2 BEV World Model
In this section, we use the BEV World Model to predict the future BEV states, as illustrated in Figure 2 (b). The process is divided into two steps. First, the current BEV state and refined trajectories are combined to form the inputs for the world model. Next, the world model recurrently predicts future states over multiple time steps.
Given refined trajectories and the current BEV state at time , we construct state-action pairs as the inputs of world model. We use the trajectory encoder in Eq. 1, with shared parameters, to encode into action embeddings . Given , each state-action pair is represented as , where .
Recurrent Future Prediction.
Furthermore, this world model is able to predict the states of the future steps in a recurrent manner, as shown by the following equation
The process above is conducted for every state-action pair in a parallel manner. It is important to note that the size of BEV states is relatively small (e.g., ), so the world model does not introduce significant computational overhead. Next, we introduce the Reward Model, which predicts rewards for a trajectory based on its future states.
3 Reward Model
In this section, we first present the reward types and then describe how they are learned.
Following , we have two types of rewards: the imitation reward and the simulation reward. The imitation reward measures how well the predicted trajectory mimics an expert trajectory. The detailed implementation will be introduced in Sec. 3.4. The simulation rewards measures trajectory quality based on simulator-defined criteria. Following , we assess trajectories using the following five criteria: no collisions (NC), drivable area compliance (DAC), time-to-collision (TTC), comfort (Comf), and ego progress (EP). As a result, we have Given and , the final reward is defined as
Reward Prediction.
Our method needs these rewards for trajectory evaluation in the inference time, so we propose reward prediction. Take the -th refined trajectory for illustration, the previous approaches typically predict a reward based solely on the current state. With the help of the world model, our reward model predicts the reward of -th trajectory based on both the current BEV state and future steps BEV states. This process is illustrated in the following equation
The refined trajectory with the highest final reward is selected as the final trajectory. Next, we introduce the supervision of these predicted rewards and BEV states.
4 Supervision in BEV Space
The world model predicts multiple future feature states corresponding to different trajectories. However, supervising these predictions in the image space is challenging, as driving logs provide only a single future trajectory. To address this, we propose supervising in the BEV space. Supervising in BEV space alleviates the difficulty of accurately modeling future states in image space while also improving computational efficiency. Many real-world traffic simulators support traffic simulations in the BEV space. They generate realistic and reliable future BEV scenarios along with rewards. We adopt nuPlan as our traffic simulator for its broad adoption and realistic scenario simulations.
Supervision of Simulation Rewards.
Regarding the simulation rewards, we leverage the simulator to produce five rewards for evaluating a trajectory. Specifically, given the -th trajectory , the simulator simulates the future location of the other agents and ego vehicle and produces a simulated future BEV semantic map with agents. Next, it uses a rule-based evaluator to evaluate this simulated future BEV scenarios for producing the corresponding target simulation rewards of . Given the rewards provided by the simulator, we use the Binary Cross Entropy (BCE) loss to supervise the predicted simulation reward . As a result, the loss is formulated as
Supervision of Imitation Reward.
In addition to the supervision from the simulator, our framework is also guided by the expert driver.
Supervision of Trajectory.
For supervising the refined trajectories, we adopt the commonly used winner-take-all strategy . In this strategy, we select the trajectory anchor that is closest to the expert trajectory. Only the corresponding refined trajectory of this anchor will be supervised. Let this trajectory anchor be denoted as . We use the L1 loss to measure the difference between the expert trajectory and the refined trajectory as
Instead of supervising all refined trajectories, we focus solely on the most optimal one. This strategy helps the model capture a diverse range of trajectory patterns, with each anchor specializing in a particular trajectory modality.
The overall training loss is as follow
Experiments
NAVSIM Dataset The NAVSIM dataset is constructed based on nuPlan . Specifically, OpenScene first down-sampled the nuPlan data from 10Hz to 2Hz to condense it into 120 hours of driving logs. NAVSIM resampled the data from OpenScene to emphasize challenging scenarios, reducing simple situations like straight-line driving. The dataset is divided into two parts: Navtrain and Navtest, comprising 1192 scenarios for training and validation and 136 scenarios for testing. Unlike nuScenes , which collects data in simpler, slower driving scenarios, NAVSIM ensures that challenging scenarios are prioritized, making simply fitting ego status insufficient for planning.
NAVSIM Metrics The original end-to-end driving metrics were primarily designed to evaluate the deviation between the predicted trajectories and those driven by human experts. However, this evaluation approach has proven inadequate, resulting in poor performance in real-world driving contexts. NAVSIM introduces more practical and reliable metrics, focusing on aligning open-loop and closed-loop metrics. Specifically, NAVSIM evaluates model performance using the Predictive Driver Model Score (PDMS), which is calculated based on five factors: No At-Fault Collision (NC), Drivable Area Compliance (DAC), Time-to-Collision (TTC), Comfort (Comf.), and Ego Progress (EP). The PDMS is calculated as
Bench2Drive Dataset and Metrics To further evaluate our framework, we conduct closed-loop experiments in the CARLA simulator . Specifically, we adopt the recently proposed Bench2Drive benchmark as our closed-loop evaluation benchmark. Bench2Drive provides a standardized training dataset on CARLA, ensuring fair comparisons across different methods. The benchmark consists of 220 short evaluation routes (around 150 meters each) distributed across all CARLA towns, with each route containing one safety-critical scenario. This design effectively reduces the evaluation variance. For metric, it utilizes the standard CARLA metric, Driving Score (DS), as the primary metric. Additionally, Bench2Drive reports the success rate, which represents the proportion of successfully completed routes. A route is considered successful if and only if DS reaches 100% on this route.
2 Implementation Details
NAVSIM For input data, it is aligned with TransFuser . We concatenate the front-view image with center-cropped front-left and front-right images, resulting in a combined resolution of pixels. For LiDAR, the point cloud surrounding the ego vehicles is used. For network architecture, we employ ResNet34 as the backbone for BEV feature extraction following TransFuser . The world model consists of two transformer decoder layers. We use trajectory anchors. The reward weights are set as follows: . The training is conducted on the Navtrain split using 8 NVIDIA L20 GPUs with a total batch size of 128, distributed across 30 epochs. In the ablation study section, we train our model for 20 epochs instead of 30 to accelerate the training process. For fast training, we pre-compute the simulation results (BEV semantic maps and scores) based on the trajectory anchors. We then feed trajectory anchors into the trajectory evaluation module during training. During testing, the trajectory evaluation module evaluates the refined trajectory instead of the anchor trajectories. We utilize the Adam optimizer with a learning rate of 1e-4.
Bench2Drive For Bench2Drive, we adopt TCP as our baseline and integrate the trajectory evaluation module. We choose TCP as it is the best open-source framework officially provided by Bench2Drive. Similar to NAVSIM, we utilize imitation rewards along with simulation-based rewards, including No Collision (NC) and Drivable Area Compliance (DAC). The number of trajectory anchors is set to 256. Notably, rather than employing the winner-take-all strategy (Eq. 9) to supervise only the best trajectory, we find that supervising all trajectories leads to improved performance. The model is trained for 27 epochs with a batch size of 300. We use the Adam optimizer with a learning rate of 1e-4.
3 Comparison with SOTA
NASIM Table 1 compares our method with state-of-the-art approaches on the NAVSIM test set. By leveraging the world model to simulate future scenarios and guide trajectory evaluation, our method achieves a PDMS of 87.1. WoTE outperforms the previous model-free approach, Hydra-MDP, highlighting the advantages of model-based trajectory evaluation.
Bench2Drive To further validate our approach, we conduct a closed-loop evaluation on the Bench2Drive benchmark within the CARLA simulator. As shown in Table 2, our method improves the Driving Score by 1.81 points, demonstrating its effectiveness in various evaluation settings.
4 Ablation Studies
Trajectory evaluation and future state prediction matter. In this experiment, we investigate the importance of the trajectory evaluation and future state prediction as Table 3 shows. The first row in Table 3 represents our baseline TransFuser , which performs trajectory prediction only. In the second row, we change to our framework, which consists of both trajectory prediction and trajectory evaluation. However, we do not used the predicted future states as the input of Reward Model. That is, and in Eq. 5. In the third row, we integrate the BEV world model to predict future states, leading to notable performance improvements across multiple metrics.
Imitation and simulation rewards are complementary. To investigate how each type of reward affects driving performance, we conduct an ablation study, as shown in Table 4. During inference, we adjust the reward weights in Eq. 4 to isolate their effects: setting enables evaluation with only simulation rewards, while setting tests performance with only an imitation reward. The results indicate that imitation and simulation rewards emphasize different aspects of planning. Trajectory evaluation using imitation rewards performs better in NC and TTC, while using simulation rewards excels in DAC and EP. By integrating both, the model leverages their respective strengths, achieving improved overall performance across all key metrics.
Recurrent future state prediction helps trajectory evaluation. Our framework supports predicting future states recurrently, as detailed in Eq. 3. In Table 5, we present the results obtained by varying the number of BEV states predicted by the world model. Our findings indicate that using a finer time step for future state prediction significantly improves performance. This is because a finer time step provides richer temporal information, which aids trajectory evaluation.
This ablation study explores the effect of varying the number of trajectories on overall performance. Specifically, the model in each row is trained and tested under the same number of trajectories, and the world model is fed with refined trajectories. As shown in Table 6, using a small number of trajectories, such as 64, significantly degrades performance. As the number of trajectories increases, performance improves notably, with a marked improvement observed when increasing from 64 to 128 trajectories. However, as the number of trajectories increases from 128 to 256, the performance gains diminish, indicating that the performance of the model approaches its optimal level. Based on these results, we choose 256 as the default number of trajectories for our experiments.
Latency analysis. A common concern is the overall latency of our framework. To address this, we measured its latency on an NVIDIA L20 GPU, as shown in Table 6. Since the GPU processes state-action pairs in parallel, our framework remains efficient even as the number of trajectories increases. The total latency is only 18.7 ms, well within the real-time requirements for end-to-end autonomous driving.
Generalization ability across different trajectories. As shown in Table 7, we investigate the generalization ability of our trajectory evaluation module. To be specific, we sample instead of trajectory anchors from the dataset. Since our model is trained under the setting of trajectory anchors, the model has never seen these trajectories. Surprisingly, we find a PDMS increase when using these unseen trajectory anchors. Besides, the model performs even better when using refined trajectories as inputs. These refined trajectories continuously change based on varying sensor inputs. This shows that the proposed world model has great generalization ability when dealing with unseen trajectories.
5 Visual Analysis
End-to-end planning trajectories. To highlight the improvements in end-to-end driving achieved through trajectory evaluation, we present visualizations of end-to-end planning trajectories in Figure 3.
Trajectory rewards. Figure 4 illustrates the predicted rewards for different trajectories.
Conclusion
In this paper, we introduced a novel framework that leverages a BEV world model for end-to-end trajectory evaluation in autonomous driving. By integrating a BEV world model, we enable a more informed evaluation process that considers the dynamic evolution of driving scenarios, leading to more effective trajectory evaluation. our approach benefits from dense supervision provided by BEV-space traffic simulators, which supply both semantic future states and rule-based reward targets. Our method achieves state-of-the-art performance on the NAVSIM and Bench2Drive benchmarks while maintaining real-time efficiency. Overall, our work establishes trajectory evaluation as an important research direction for end-to-end autonomous driving. We hope our framework serves as a strong baseline and inspires further research in end-to-end online trajectory evaluation.