NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking

Daniel Dauner, Marcel Hallgarten, Tianyu Li, Xinshuo Weng, Zhiyu Huang, Zetong Yang, Hongyang Li, Igor Gilitschenski, Boris Ivanovic, Marco Pavone, Andreas Geiger, Kashyap Chitta

Introduction

Autonomous vehicles (AVs) have gained immense research interest due to their potential to change transportation and improve traffic safety . This has created a large community working on the development of AV algorithms, which map high-dimensional sensor data to desired vehicle control outputs. Therefore, measuring and comparing the performance of AV algorithms is a crucial task.

Unfortunately, it is extremely challenging to evaluate driving performance, and the most widely-used benchmarks today fall short in several respects: (1) the datasets used, such as nuScenes , were created for perception tasks such as object detection. As such, they focus on visual diversity and label quality instead of the relevance of the data for research on planning. Often, most frames have a trivial solution of extrapolating the historical driving behavior, leading to “blind” driving policies that observe only the vehicle’s past trajectory obtaining state-of-the-art performance . (2) Due to the fact that driving is an inherently multifaceted task where the algorithm must balance several desired properties such as safety, comfort, and progress, the evaluation must also involve multiple complementary metrics. However, as shown in Fig. 1, existing metrics such as the average displacement error (ADE) between a predicted and recorded human trajectory often misrepresent the relative accuracy of trajectories. (3) Since driving involves interactions among multiple agents, evaluation must ideally be interactive, e.g., in simulation. Unfortunately, existing simulators with synthetic sensor data exhibit a significant domain gap to real-world driving. (4) Besides, the lack of a standardized evaluation setup has led to subtle inconsistencies between metrics in existing work, leading to unfair comparisons and inaccurate conclusions . Collectively, these problems hinder progress in the development of AVs, emphasizing the need for more principled benchmarks.

In this work, we take steps towards alleviating these issues. First, we propose a strategy for sampling interesting driving scenarios and apply it to the largest publicly-available driving dataset . We obtain, for the first time, over 100k challenging real-world driving scenarios for training and evaluating sensor-based driving policies. We show that in these scenarios, “blind” driving policies fail to compete with more principled sensor-based policies. Second, we draw inspiration from the literature of rule-based planning for AVs to identify a set of diverse, efficient, and principled metrics that cover multiple facets of the autonomous driving task. Third, we circumvent the need for inaccurate sensor simulation with domain gaps by simplifying our simulation to a non-reactive one. Given an observed real-world sensor input, the agent under test commits to a set of actions for a specific time horizon. Further, these actions are assumed to not affect the future behavior of other agents in the scene. Under this setting, it is possible to simulate the expected motion of all agents over this time horizon in a simplified bird’s-eye-view (BEV) abstraction of the scene, and incorporate metrics that involve interactions, as we observe in Fig. 1. Empirically, we demonstrate that our selected metrics are well-correlated to the outcomes of closed-loop simulations. Finally, we establish an official evaluation server on the open-source HuggingFace platform, which is free, has a low maintenance overhead, and enables future scaling to more challenging datasets and metrics.

We combine these ideas to propose NAVSIM, a comprehensive tool for AV data curation, simulation, and benchmarking. We instantiate standardized training and evaluation splits for NAVSIM with the OpenScene dataset , though our framework can be extended to other datasets. With these splits, we present a detailed analysis of popular end-to-end driving models previously benchmarked either exclusively on CARLA or nuScenes , providing the first direct comparison between these families of approaches in an independent evaluation setting. Interestingly, we find that the performances of the best methods developed in both settings are similar, despite a vast difference in computational requirements for their training. Finally, we review the insights gained through the 2024 NAVSIM challengehttps://opendrivelab.com/challenge2024/#end_to_end_driving_at_scale, hosted in conjunction with the CVPR 2024 Workshop on Foundation Models for Autonomous Systems. For the challenge, 143 teams from 13 countries developed diverse methods that competed on the proposed benchmark. The top methods ranged from multi-billion parameter vision language models to more efficient and recently overlooked approaches based on trajectory sampling and scoring , demonstrating the remarkable ability of the broader community to advance AV research when provided with the right tools.

Contributions. (1) We build NAVSIM, a framework for non-reactive AV simulation, with standardized protocols for training and testing, data curation tools ensuring broad accessibility, and an official public evaluation server used for the inaugural NAVSIM challenge. (2) We develop configurable simulation-based metrics that are well-suited for evaluating sensor-based motion planning. (3) We reimplement a collection of end-to-end approaches for NAVSIM including TransFuser, UniAD, and PARA-Drive, showcasing the surprising potential of simple models in our challenging scenarios.

Related Work

End-to-End Driving. End-to-end driving streamlines the entire stack from perception to planning into a single optimizable network. This eliminates the need for manually designing intermediate representations. Following pioneering work , a diverse landscape of end-to-end models has emerged. For instance, an extensive body of end-to-end approaches focuses on closed-loop simulators, utilizing single-frame cameras, LiDAR point clouds, or a combination of both for expert imitation . More recently, developing end-to-end models on open-loop benchmarks has gained traction . Our work introduces a new evaluation scheme with which we compare end-to-end models from both communities.

Closed-Loop Benchmarking with Simulation. Driving simulators allow us to evaluate autonomous systems in a closed-loop manner and collect downstream driving statistics, including collision rates, traffic-rule compliance, or comfort. A broad body of research conducts evaluations in simulators, such as CARLA or Metadrive with sensor simulation, or nuPlan and Waymax for data-driven simulation. Unfortunately, ensuring realism when simulating traffic behavior or sensor data remains a challenging task. To simulate camera or LiDAR sensors, most established simulators rely on graphics-based rendering methods, leading to an inherent domain gap in terms of visual fidelity and sensor characteristics. Data-driven simulators for motion planning incorporate traffic recordings but do not support image or LiDAR-based methods . Data-driven sensor simulation leverages and adapts real-world sensor data to create new simulations where the vehicle may move differently, but the rendering quality of existing tools is subpar . Further, while promising image or LiDAR synthesis approaches exist, efficiently simulating sensors entirely from data remains an open problem. In this work, we provide an approach for the evaluation of real sensor data with simulation-based metrics by making a simplifying assumption that the agent and environment do not influence each other over a short simulation horizon. Despite this strong assumption, when benchmarking on real data, NAVSIM better reflects planning performance than established evaluation protocols, as demonstrated through our systematic experimental analysis.

Open-Loop Benchmarking with Displacement Errors. Open-loop evaluation protocols commonly measure displacement errors between trajectories of a recorded expert (i.e., of a human driver) and a motion planner. However, several issues concerning evaluation with displacement errors have surfaced recently, particularly on the nuScenes dataset . Given that nuScenes does not provide standardized planning metrics, prior work relied on independent implementations, which led to inconsistencies when reporting or comparing results . Next, most planning models in nuScenes receive the human trajectory endpoint as a discrete direction command , thereby leaking ground-truth information into inputs. Moreover, about 75% of the scenarios in nuScenes involve straight driving , leading to simple solutions when extrapolating the ego-motion. For instance, AD-MLP demonstrates that an MLP on the kinematic ego status (ignoring perception completely) can achieve state-of-the-art displacement errors . Such blind agents are undeniably dangerous, which highlights a broader concern: displacement metrics are not correlated to closed-loop driving . In this work, we address prevalent issues of nuScenes and propose a standardized driving benchmark with challenging scenarios and an official evaluation server. We derive a navigation goal from the lane graph instead of the human trajectory to prevent label leakage, and propose principled simulation-based metrics as an alternative to displacement errors.

NAVSIM: Non-Reactive Autonomous Vehicle Simulation

NAVSIM combines the ease of use of open-loop benchmarks such as nuScenes with metrics based on closed-loop simulators such as nuPlan . In the following, we give a detailed introduction to the task and metrics that driving agents are challenged with in NAVSIM. Subsequently, we propose a filtering method to obtain standardized train and test splits covering challenging scenes.

Task description. Driving agents in NAVSIM must plan a trajectory, defined as a sequence of future poses, over a horizon of four seconds. Their input contains streams of past frames from onboard sensors, such as cameras, LiDAR, as well as the vehicle’s current speed, acceleration, and navigation goal, jointly termed the ego status. For compatibility with prior work , we provide the navigation goal as a one-hot vector with three categories: left, straight, or right.

Non-Reactive Simulation. Traditional closed-loop benchmarks normally infer planners at high frequencies (e.g., 10Hz) . However, this requires efficient simulation of all input modalities of the driving agent, including high-dimensional sensor streams in the case of sensor-based approaches. To sidestep this, the core idea of NAVSIM is to evaluate driving agents using a non-reactive simulation. This means driving agents are only employed in the initial frame of each scene. Afterwards, the planned trajectory is kept fixed for the entire trajectory duration. Over this short horizon, no environmental feedback is provided to the driving agent, and the NAVSIM evaluation is purely based on the initial real-world sensor sample. This makes the agent’s task more challenging, as it requires a safe plan up to the entire four-second horizon. Therefore, it is not possible to scale this kind of simulation to arbitrarily long time horizons. Despite this limitation, non-reactive simulation offers a key advantage: unlike traditional open-loop benchmarks, which mainly compare the planned trajectory to the human driver’s trajectory in a similar setting, it enables the use of simulation outcomes to compute metrics reflecting safety, comfort, and progress. An LQR controller is applied at each simulation iteration to calculate steering and acceleration values, and a kinematic bicycle model propagates the ego vehicle. We execute this pipeline at 1010Hz over the 44s trajectory horizon. In Sec. 4.1, we show that despite our simplifying assumption, our evaluation results in a much better alignment with closed-loop metrics than traditional open-loop metrics achieve.

PDM Score. NAVSIM scores driving agents in two steps. First, subscores in range $arecomputedaftersimulation.Second,thesesubscoresareaggregatedintothePDMScore(PDMS)are computed after simulation. Second, these subscores are aggregated into the PDM Score (PDMS)\in$. It is named after the Predictive Driver Model (PDM) , a state-of-the-art rule-based planner which uses this scoring function to evaluate trajectory proposals during closed-loop simulation in nuPlan. The metric is also an efficient reimplementation of the nuPlan closed-loop score metric . In NAVSIM, the PDMS can be adapted by adding or removing subscores, changing aggregation parameters, or making subscores more challenging, e.g., by adapting their internal thresholds. It is calculated per frame and averaged across frames. In this work, we use the following aggregation of subscores:

Subscores are categorized by their importance as penalties or terms in a weighted average. A penalty punishes inadmissible behavior such as collisions with a factor <1<1. The weighted average aggregates subscores for other objectives such as progress and comfort. In the following, we briefly describe each subscore. More details can be found in the supplementary material.

Penalties. Avoiding collisions and staying on the road is imperative for motion planning as it ensures traffic rule compliance and the safety of pedestrians and road users. Thus, failing to drive with no collisions (NC) with road users (vehicles, pedestrians, and bicycles) or infractions with regard to drivable area compliance (DAC) result in hard penalties of scoreNC=0\texttt{score}_{\texttt{NC}}=0 or scoreDAC=0\texttt{score}_{\texttt{DAC}}=0 respectively. This results in a PDMS of 0 for the current scene. We ignore certain collisions that are not considered "at-fault" in the non-reactive environment, e.g. when the ego vehicle is static. For collisions with static objects, we apply a softer penalty of scoreNC=0.5\texttt{score}_{\texttt{NC}}=0.5.

Weighted Average. The weighted average accounts for ego progress (EP), time-to-collision (TTC), and comfort (C). The ego progress subscore scoreEP\texttt{score}_{\texttt{EP}} represents the agent progress along the route center as a ratio to an approximated safe upper bound from the PDM-Closed planner . PDM-Closed obtains a possible progress value without collisions or off-road driving with a search-based strategy based on trajectory proposals. The final ratio is clipped to $whilediscardinglowornegativeprogressscoresiftheupperboundisbelow5meters.Next,theTTCsubscoreensuresthatdrivingagentsrespectthesafetymarginstoothervehicles.Defaultingtoavalueofwhile discarding low or negative progress scores if the upper bound is below 5 meters. Next, the TTC subscore ensures that driving agents respect the safety margins to other vehicles. Defaulting to a value of1,thissubscoreissettoifforanysimulationstepwithinthe, this subscore is set to if for any simulation step within the4\text{s}horizon,theego−vehicle’stime−to−collison,whenprojectedforwardwithaconstantvelocityandheading,islessthanacertainthreshold.Finally,thecomfortsubscoreisobtainedbycomparingtheaccelerationandjerkofthetrajectorytopredeterminedthresholds.FollowingthecostweightsusedbythePDM−Closedplanner,wesetthecoefficientsoftheweightedaverageashorizon, the ego-vehicle’s time-to-collison, when projected forward with a constant velocity and heading, is less than a certain threshold. Finally, the comfort subscore is obtained by comparing the acceleration and jerk of the trajectory to predetermined thresholds. Following the cost weights used by the PDM-Closed planner, we set the coefficients of the weighted average as\texttt{weight}_{\texttt{EP}}=5,,\texttt{weight}_{\texttt{TTC}}=5,and, and\texttt{weight}_{\texttt{C}}=2$.

Dataset. The NAVSIM framework is agnostic to the choice of driving dataset. We choose OpenScene , a redistribution of nuPlan , the largest annotated public driving dataset. OpenScene includes 120 hours of driving at a reduced frequency of 22Hz typically considered by end-to-end planning algorithms, resulting in a 10×10\times reduction of data storage requirements compared to nuPlan from over 20 TB to 2 TB. Our agent input, based on OpenScene, comprises eight cameras, each with a resolution of 1920×10801920\times 1080 pixels, and a merged LiDAR point cloud from five sensors. The input includes the current time-step and optionally 3 past frames, totaling 1.51.5s at 22Hz. In principle, any driving dataset that provides annotated HD maps, object bounding boxes, and sensor data can be converted into this format and thus be used with NAVSIM.

Filtering for challenging scenes. A majority of human driving data involves trivial situations such as being stationary or straight driving at a near constant speed. These can be solved efficiently by simple heuristics, e.g., as depicted in Fig. 2 (a), the baseline of maintaining a constant velocity and heading achieves a PDMS of 7979% on the OpenScene dataset, where human-level performance corresponds to 9191%. In NAVSIM, we propose the use of a filtered dataset to remove frames with (1) near-trivial solutions and (2) significant annotation errors. We remove highly simplistic scenes by detecting if the previously mentioned constant velocity agent exceeds a PDMS of 0.80.8. Similarly, we remove scenes in which the human trajectory results in a PDMS of less than 0.80.8. This ensures that an acceptable solution exists to these difficult scenarios and filters out noisy annotations such as inaccurate bounding boxes. These thresholds can be adjusted based on the desired filtered dataset size. The resulting scenarios are challenging, which is underlined by the score of the constant velocity agent dropping to 2222%, whereas the human expert achieves a score of 9595%. The higher ratio of non-trivial scenarios, such as turning, also results in endpoints being less distant longitudinally when nonzero, and more evenly distributed laterally, as seen in Fig. 2 (b-c). We employ this filtering strategy to provide standardized splits for training and testing, called navtrain and navtest, with 103k and 12k samples respectively. This curated data serves as a benchmark accessible as a standalone download option with a moderate storage demand given its large scale and diversity (450 GB).

Experiments

In this section, we present the results of our experiments aimed at answering the following questions: (1) Can non-reactive open-loop simulation provide sufficient correlation to closed-loop metrics? (2) What new conclusions do experiments on NAVSIM provide compared to prior benchmarks?

Open-loop metrics should ideally be aligned with closed-loop metrics in their evaluation of different driving algorithms. In this section, we benchmark a large set of planners to analyze the alignment of closed-loop metrics with traditional distance-based open-loop metrics and the proposed PDMS.

Benchmark. Studying the relation of closed-loop and open-loop metrics necessitates access to a fully reactive simulator. To stay compatible with the dataset, we use the nuPlan simulator , which enables simulation for privileged planners with access to ground-truth perception and HD map inputs. Similar to PDMS, nuPlan combines weighted averages and multiplied penalties in two official scores: the open-loop score (OLS) aggregates displacement and heading errors with a multiplied miss-rate, and the closed-loop score (CLS) implements similar metrics from Section 3. Including PDMS, all metrics are in $$ with higher scores indicating better performance.

Due to the heavy computational requirements of closed-loop simulation, we evaluate on the navmini split. This is a new split we create for rapid testing, with 396 scenarios in total that are independent of both navtrain and navtest but filtered using the same strategy (Section 3.1) and hence similarly distributed. We note that nuPlan offers two kinds of background agents: reactive agents along lane centers based on the Intelligent Driver Model (IDM) , and non-reactive agents replayed from the dataset, which we employ unless otherwise stated. We also default to a closed-loop simulation duration of d=15sd=15\textrm{s}, and a planning frequency of f=10Hzf=10\textrm{Hz}, which are standard for nuPlan.

Motion Planners. Open-loop metrics favor learned planners while rule-based approaches perform well in closed-loop evaluation in nuPlan . We use a combination of both planner types in this experiment to cover different performance levels. In total, we include 37 rule-based planners with 2 constant velocity and 8 constant acceleration models, 15 IDM planners , and 12 PDM-Closed variants which differ in hyperparameters for trajectory generation. For learned planning, we evaluate Urban Driver models of 2 model sizes and 2 training lengths, and PlanCNN models with 15 input combinations of the BEV raster, ego status, centerline, and navigation goal. We train all models on {25%,50%,100%}\{25\%,50\%,100\%\} of navtrain and an equally sized uniformly sampled subset of OpenScene, giving 114 learned planners. See the supplementary material for additional details.

Results. The alignment between metrics is presented in Fig. 3 (a-e). Compared to OLS, we consistently observe better closed-loop correlation for PDMS, in terms of Spearman’s (rank) and Pearson’s (linear) correlation coefficients. As shown in (a), PDMS can capture the closed-loop properties of both learned and rule-based planners, whereas distance-based open-loop metrics show a clear misalignment. Decreasing the CLS duration in (b) from d=15d=15s to d=4d=4s further raises the correlation of PDMS and OLS, as the simulation horizon more closely matches the open-loop counterparts. Interestingly, we observe a higher correlation of open-loop metrics in (c) when reducing the planning frequency to 22Hz. We expect a lower planning frequency to mitigate cumulative errors and enhance the controller’s stability in simulation, leading to more precise trajectory execution. Moreover, we observe an increase in correlation for longer PDMS horizons in (d), ranging from h=2h=2s to h=8h=8s. While predicting the future motion over 8s is challenging in uncertain scenarios, our results indicate the value of long horizons when evaluating motion planners. Lastly, replacing the non-reactive background agents with reactive IDM vehicles in (e) has little effect on the correlation, possibly due to the similar difficulty of both tasks .

2 Analysis of the State of the Art in End-to-End Autonomous Driving

In this section, we benchmark a collection of end-to-end architectures, which previously achieved state-of-the-art performance on existing open- or closed-loop benchmarks.

Methods. As a lower bound, we consider the (1) Constant Velocity baseline detailed in Section 3.1. We include an (2) Ego Status MLP as a second "blind" agent, which leverages an MLP for trajectory prediction given only the ego velocity, acceleration and navigation goal. As an established architecture on CARLA, we evaluate our reimplementation of (3) TransFuser , which uses three cropped and downscaled forward-facing cameras, concatenated into a 1024×2561024\times 256 image, and a rasterized BEV LiDAR input for predicting waypoints. It performs 3D object detection and BEV semantic segmentation as auxiliary tasks. We then consider (4) Latent TransFuser (LTF) , which shares the same architecture as TransFuser but replaces the LiDAR input with a learned embedding, hence requiring only camera inputs. Moreover, we provide two state-of-the-art end-to-end architectures for open-loop trajectory prediction on nuScenes. (5) UniAD incorporates a wide range of tasks, such as mapping, tracking, motion, and occupancy prediction in a semi-sequential architecture, which processes feature representations through several transformer decoders culminating in a trajectory planning module. (6) PARA-Drive uses the same auxiliary tasks, but parallelizes the network architecture, and the auxiliary task heads are trained in parallel with a shared encoder. Both UniAD and PARA-Drive use a BEVFormer backbone , which encodes the eight surround-view 1920×10801920\times 1080 camera images over four temporal frames into a BEV feature representation. Implementation details for all methods are provided in the supplementary material.

Results. We show our results on navtest in Table 1. The Constant Velocity model is a lower bound, as the agent is used to identify trivial driving scenes excluded from the benchmark. The Ego Status MLP achieves a PDMS of 65.665.6, showing the value of the acceleration and navigation goal for avoiding collisions and driving off-road. However, we observe a clear gap between agents relying solely on the ego status and those considering sensor data, in contrast to results on nuScenes . All sensor agents achieve a PDMS of over 8383, where TransFuser and PARA-Drive marginally perform best, with a PDMS of 84.084.0. Surprisingly, the camera-only LTF achieves similar results (83.883.8). UniAD reaches a PDMS of 83.483.4, which, together with PARA-Drive, do not surpass the performance of TransFuser and LTF, despite the need for more demanding training, e.g., 80 GPUs for 3 days to train PARA-Drive versus 1 GPU for 1 day for TransFuser on the navtrain split. Due to the definition of at-fault collisions, which discard certain rear-collisions into the ego vehicle, we suspect that surround-view cameras used by UniAD and PARA-Drive, and LiDAR input of TransFuser, are less important than the wide-angle front camera which is the only input of LTF. The 1010 PDMS discrepancy to the human operator demonstrates that navtest poses challenges even to well-studied end-to-end architectures. Specifically, the drivable area compliance (DAC) and ego progress (EP) subscores remain the most challenging. Notably, EP cannot be solved purely by human imitation, given that the maximum progress estimate used for normalization is based on a privileged rule-based motion planner. Interestingly, all agents achieve near-perfect comfort scores, indicating that smooth acceleration and jerk profiles are learned naturally from human imitation.

Analyzing TransFuser. In Table 2, we compare several training settings for TransFuser. For the three training seeds in configs A1-A3, we observe a standard deviation of ±\pm 0.560.56 in PDMS, which is relatively small compared to variance among training seeds for closed-loop simulations in CARLA . Further, unlike CARLA, NAVSIM is deterministic, and we obtain identical scores when repeating evaluations of a deterministic driving agent. Discarding velocity and acceleration (B1) lowers PDMS by 1.5−2.61.5-2.6, whereas only removing the acceleration (B2) lowers the score by 1.0−2.11.0-2.1. We conclude that while TransFuser benefits from the ego status, it is not purely relying on the kinematic state for planning. Next, only considering the front camera (C1) with a 60∘60^{\circ} FOV leads to a small drop in almost all subscores, compared to our default setting of three cropped and concatenated images with a FOV of 140∘140^{\circ}. However, expanding the FOV with additional cameras does not result in substantially improved scores. Interestingly, restricting the LiDAR range to 1616m in all directions (D1), results in a score of 7979, which is lower than dropping LiDAR altogether (see LTF in Table 1). Expanding the LiDAR range to 6464m in the forward direction (D2) or all directions (D3) does not provide significant improvements. We suspect that changes in the LiDAR range overly simplify or complicate the auxiliary 3D object detection and BEV semantic segmentation tasks, which operate in the LiDAR coordinate frame, hindering effective imitation learning. We check the impact of the auxiliary tasks by excluding them, where performance drops without BEV Segmentation (E1).

CVPR 2024 NAVSIM Challenge. We organized the inaugural NAVSIM challenge which ran from March - May 2024. To ensure integrity, we used a private dataset and only gave participants access to sensor inputs, withholding all annotations. Competitors could submit their agent’s trajectories to our leaderboard, where they were simulated and scored to obtain the PDMS. We received 463 submissions from 143 teams, of which 78 submissions were made publicly visible. We summarize their scores in Fig. 4, relative to the constant velocity and TransFuser baselines from Table 1. The winning entry extended TransFuser and learned to predict proxy subscores for trajectory samples , with a sampling strategy inspired by VADv2 . These predicted subscores were weighted alongside a human imitation score to select the output plan. While the idea of sampling and scoring trajectories is well-known , it has recently been overlooked in favor of approaches which predict a single trajectory. This result prompts a reassessment of such methods. The team that placed second employed a vision language model (VLM) for driving, which is rapidly emerging as a sub-field in the AV literature . Several submissions attempted to reimplement or extend prior work on nuScenes such as UniAD and VAD , but were unable to outperform the TransFuser baseline by the challenge submission deadline, given the significant engineering challenge and compute requirements. The diversity of the solutions on the leaderboard shows the potential of NAVSIM as a framework for pushing the frontiers of autonomous driving research. We aim to hold future competitions with more challenging data and metrics. Detailed competition results and statistics are provided in the supplementary material.

Discussion

We present NAVSIM, a framework for non-reactive AV simulation. We address several shortcomings of existing driving benchmarks and propose standardized but configurable simulation-based metrics for benchmarking driving policies. We improve accessibility for conducting AV research by providing downloadable challenging scenario splits and simple data curation methods. We demonstrate that our evaluation protocol is better aligned to closed-loop driving, benchmark an established set of end-to-end planning baselines, and present the results of our inaugural competition. We hope that NAVSIM can serve as an accessible toolkit for AV researchers that bridges the gap between simulated and real-world driving. We aim to support more datasets in the future, and advocate for more open dataset releases by the community for accelerating progress in autonomous driving.

Limitations. While we show improvements over displacement error-based benchmarking, several aspects of driving remain unaddressed by evaluation in NAVSIM. A high PDMS does not always imply a high CLS, since our framework does not consider reactiveness or compounding accumulation of errors in closed-loop simulation. Moreover, closed-loop metrics also face problems, i.e., PDMS inherits several weaknesses of nuPlan’s CLS. Both scores do not regard certain traffic rules (e.g., stop-sign or traffic light compliance) or concepts such as transit and fuel efficiency. Moreover, as in CLS, rear-end collisions into the ego vehicle are currently not classified as "at-fault", resulting in little importance given to the scene behind the vehicle in NAVSIM. Further, certain limitations of the nuPlan dataset persist in NAVSIM, such as minor errors in camera parameters or noise in poses and 3D annotations. Our analysis might favor methods that are robust to such inconsistencies. In the future, we aim to improve the subscore definitions (e.g. the at-fault collision logic) and add more subscores during aggregation. Additionally, we would like to explore pose or route augmentations to account for control drifts in closed-loop driving and improve our evaluation’s diversity. Given these limitations, we strongly encourage the use of graphics-based closed-loop simulators, such as CARLA , as complementary benchmarks to NAVSIM when developing planning algorithms.

Acknowledgments

This work was supported by the ERC Starting Grant LEGO-3D (850533), the DFG EXC number 2064/1 - project number 390727645, the German Federal Ministry of Education and Research: Tübingen AI Center, FKZ: 01IS18039A and the German Federal Ministry for Economic Affairs and Climate Action within the project NXT GEN AI METHODS. We thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting Daniel Dauner and Kashyap Chitta. We also thank HuggingFace for hosting our evaluation servers, the team members of OpenDriveLab for their organizational support, as well as Napat Karnchanachari and his team from Motional for open-sourcing their dataset and providing us the private test split used in the 2024 NAVSIM Challenge.

References