Towards learning-based planning:The nuPlan benchmark for real-world autonomous driving
Napat Karnchanachari, Dimitris Geromichalos, Kok Seang Tan, Nanxiang Li, Christopher Eriksen, Shakiba Yaghoubi, Noushin Mehdipour, Gianmarco Bernasconi, Whye Kit Fong, Yiluan Guo, Holger Caesar
I Introduction
In the last decade, autonomous vehicle perception and prediction have been revolutionized by deep learning-based methods trained on large-scale datasets . While similar attempts have been made in the field of learning-based or neural planning, these are not yet able to surpass their rule-based counterparts. One possible reason is the difficulty of generalizing driving scenarios when learned from a limited number of examples. Furthermore, driving scenarios typically follow a long-tail distribution, which further exacerbates the generalization issue. Finally, learning-based planning lacks formal safety guarantees, thus making it potentially unsafe and challenging to certify.
We introduce the nuPlan dataset and simulation framework for autonomous vehicle planning. Our goal is to create a testbed for open-loop and closed-loop planning starting in real-world scenarios. This test bed is then used to compare traditional, learning-based, and hybrid planners. nuPlan enables numerous novel types of research, such as learning-based planning, the interplay between prediction and planning, and end-to-end planning using a large amount of published sensor data.
We release the largest dataset for autonomous driving to date, with a total of 1282h from 4 cities. We also publish an unprecedented 128h of sensor data.
We develop techniques to auto-label the dataset with accurate object tracks, traffic lights, and scenario labels.
We publish our closed-loop simulation and evaluation framework (Fig. 1) and compare the performance of traditional and learning-based planners to identify gaps.
II Related Work
In this section we discuss the most important datasets and simulators for autonomous vehicle planning. Based on these, we introduce the main categories of planners: classical, learning-based, and hybrid planners.
Tab. I provides an overview of large-scale prediction and planning datasets with more than 20h of data. We omit smaller datasets like Interaction , highD , inD , OpenDD and CommonRoad and subsets of datasets not focused on prediction or planning . With the exception of nuPlan and CommonRoad , all datasets focus on prediction (motion forecasting). Offline perception is crucial to train and evaluate planners using high-quality object tracks, but only present in Waymo , MONA and nuPlan. Likewise, the availability of traffic light statuses is crucial for realistic traffic simulation, but only Waymo and Lyft contain these and only from an online perception system, rather than developing offline traffic light status inference as in nuPlan. To assess planning performance we need to focus on specific scenarios and evaluate them in closed-loop. nuPlan is the first dataset to feature both scenario tags and a closed-loop simulation framework. Lyft provides interactive tutorials for closed-loop simulation, but they lack the modular framework, evaluation server, and hold-out test set of nuPlan. Finally, we need a large-scale dataset to generalize well. Only the Lyft , Shifts and nuPlan datasets provide more than 1000h of driving data and nuPlan was the first such dataset that provides lidar and camera sensor data (128h), although Waymo later also released compressed lidar data.
II-B Simulation
Many works in the literature use proprietary simulators . While sharing their approaches to planning, they do not provide enough information to reproduce their experimental results. Graphical simulators like CARLA and AirSim focus on photorealistic rendering, but lack the realism of real-world maps and agent behavior. CommonRoad was the first open-source simulator focused on planning. However, CommonRoad does not provide a real world dataset as the basis for the simulation, instead resorting to a small number of manually crafted scenarios and tools to import other datasets, albeit without any sensor data. With nuPlan we aim to overcome the above limitations by releasing a large-scale real-world dataset and an open-source closed-loop simulator. Following the release of nuPlan, ScenarioNet focused specifically on Reinforcement Learning and integrated nuPlan and other datasets into their pipeline. They also interfaced with MetaDrive that enables graphical simulation of nuPlan.
II-C Planning
Classical planning. The planning problem has long been treated as an optimization problem in the traditional approaches . By carefully designing a cost function, the optimization aims to generate the optimal trajectory that minimizes the cost function in the corresponding search space (e.g., A* search , sampling-based methods , dynamic programming ). While these approaches enjoy the theoretical guarantees on the convergence to an optimal solution, hand-crafting the cost function that represents the human-like driving behavior is challenging. In practice, many studies rely on tremendous engineering efforts to fine-tune the solution.
Learning based planning. Pioneering by the study of , the idea of using a neural network to imitate expert driver and directly output driving control command provides an alternative planning solution. With the recent success of deep learning, the learning based planning received considerable attention .
Imitation learning (IL) and inverse reinforcement learning (IRL): IL trains a model to either map the sensor data directly (end-to-end system), or indirectly through the perception and prediction models (modularized system), to the expert driver actions (e.g. steering and speed profile) . With the advancements in deep learning literature, IL studies adopt the state-of-the-art supervised learning models architectures to learn better scene representations . IL often suffers from a poor generalization where the compounding error leads to driving scenarios that are outside of the training data, known as “covariate shift” . Carefully designed data augmentation is often used to address this issue . As an alternative to directly imitating the driver behavior, IRL aims to learn an unknown reward function that explains expert demonstrations . Once learned, such reward function is used to infer the optimal trajectory from a set of pre-defined or generated trajectories . The maximum entropy formulation of IRL has been applied to autonomous driving where the reward function is estimated based on a set of handcrafted features in .
Reinforcement learning (RL): RL learns the optimal driving behavior by interacting with the environment and optimizing a given reward function. RL is well-suited for handling the interaction between the agent and the environment in a sequential decision process. However, due to its learning by “trial-and-error” search nature, studies in RL rely on the driving simulation to provide the environment . These studies have demonstrated strong performance in simulations. The real-world applications of RL in autonomous driving are also reported in .
Hybrid solutions. Hybrid solutions are proposed to leverage the advantages of both the classical and learning based planning. Several studies use learning to improve the classical planning algorithm. These include using a learning model to guide the exploration for sampling-based path-planners , using a learning model to improve the efficiency of sampling-based motion-planners in high dimensional setting , applying optimizer to actively rectifies the learning model’s plan to satisfy the safety and comfort requirements . Other studies leverage the classical planning to generate the trajectory candidates, which are passed to an ML-based model to evaluate .
System design. While planning is the ultimate goal of the autonomous driving system, different system designs to assemble the perception and prediction introduce opportunities and challenges to improve the planning performance. Most autonomous driving systems use a multi-stage pipeline of independent tasks like perception, prediction and planning . The hope is that the performance gain on individual tasks translates to a better planning performance. In contrast, various studies consider multi-task learning (MTL) where they jointly train models to perform perception, prediction and planning simultaneously . These works have shown that MTL achieves better data utilization at lower computation cost. Recently, some studies leverage the query-based design in transformer architectures to integrate all tasks in a unified framework that is trainable end-to-end . Such a framework encourages better spatio-temporal feature learning from the sensor data and directly improves the planning performance.
III Dataset
In this section, we describe how we collect the nuPlan dataset. We enhance it by auto-labeling the object tracks of other agents and traffic lights, as well as mining for scenarios that are relevant for tracking.
We collected data from 4 cities (Boston, Pittsburgh, Las Vegas, and Singapore) to build a benchmark dataset for ML-based planning. In total, we have 1282 hours of challenging and real-world driving scenarios. For example, double parking in Boston, custom precedence patterns for left turns in Pittsburgh, crowded pick-up and drop-off points (PUDOs) in casinos in Las Vegas, and left-hand traffic in Singapore. We exclude heavy rain and night data, as these would impact the quality of our perception system (see Sec. III-B).
Manual driving. We use Chrysler Pacifica Plug-in Hybrid Electric Vehicles (PHEV) to drive in these cities. See Fig. 2 for the sensor setup. Our vehicle operators (VOs) are instructed to use a natural driving style and drive safely. Since our focus is on planning, it is crucial that we drive manually, while most other datasets use a combination of manual and automated driving, which may lead the planner to imitate less desirable driving behavior. The VOs drive from a predefined starting point to a goal using a known route. For example, we drive between various hotels and casinos on the Las Vegas strip, which are typical routes for our robotaxi and are known and mapped beforehand.
Sensor data. Sensor data include lidar point clouds and camera images. Due to the vast scale of the full sensor dataset (200+ TB), we only release a subset of the sensor data which totals 128 hours. This subset was selected to satisfy all stratification constraints as described below.
Maps. Similar to nuScenes , nuPlan provides detailed human-annotated 2D high-definition semantic maps of the driving locations. We release rasterized and vectorized maps. While rasterized maps are useful for simplicity and efficient lookup, vectorized maps provide more precise geometric information and metadata. Examples of semantic map layers are lanes, car parks, crosswalks and stop lines (see Fig. 4).
III-B Auto-labeling
In order to faithfully reconstruct various driving scenarios, we develop an auto-labeling system. It first generates the tracks for all the objects in the scene; then traffic light statuses are inferred from these tracks. Based on the above labels, we can reliably mine different driving scenarios.
Offline perception. We build an offline perception system to label the objects in the scene automatically. Compared with the online perception systems used in many other datasets , the offline version is not constrained by latency and causality. Therefore, the fidelity of the generated tracks is drastically higher, enabling us to evaluate planning performance under very limited perception noise.
Inspired by , our offline perception system contains three stages: 3D object detection, offline tracking, and global track refinement. The detector in the first stage takes the point clouds from both the top lidar and side lidars as input and detects the bounding boxes through a large neural network . The offline tracker leverages both past and future detections in an extended time window to generate tracks. In the last step, a novel network is developed to load both tracks and points clouds within the tracks to refine the attributes of all the bounding boxes of vehicle class, such as positions, headings, sizes and velocity.
Traffic light status. To create a realistic simulation of the environment, it is crucial to capture the traffic light statuses. Existing datasets lack traffic lights or use online vision-based systems to detect their statuses . In contrast, we develop a novel offline system to automatically label the statuses of traffic lights by inferring them from the motion of the actors present in the scene. Our labeling system is able to cover all lanes with observed agents.
We make use of the detections and tracks produced by our offline perception system, as well as map information. To infer a green traffic light status within a given intersection, we determine if there are agents moving within the intersection in the direction controlled by the particular traffic light. To infer a red traffic light status, we check for agents slowing down or being stationary in the lane that approaches the intersection and controlled by the particular traffic light.
Scenario mining. Traditional approaches to evaluate planning performance are dominated by monotonous lane following scenarios. To get a fine-grained understanding of the planning performance, we develop a scenario taxonomy and scenario mining algorithms using low-level attributes like vehicle speed and state transitions. The attributes can be inferred from offline perception tracks and traffic light statuses. In total, we have 73 unique scenario types.
IV Simulation
nuPlan provides a simulation framework (Fig. 1) that is modular and flexible to work with different datasets and setups. The simulation is initialized with the real-world observations captured in the dataset, namely raw sensor data or object tracks. Given these environment observations, an agent model can be used to predict the future trajectories of all agents. Observations and agent trajectories are passed to a planner that predicts the best route for the ego vehicle given the other agents’ routes. Finally, a controller converts the intended route into a feasible trajectory. The simulation can either playback the actions recorded in the dataset (open-loop) or allow the simulation to deviate from the recording by incorporating the ego’s actions (closed-loop). Below are the simulation components in detail.
Of the 6 object classes, vehicle, pedestrian, generic object, traffic cone, barrier, bicycle, construction zone sign, 3 are moving object classes simulated as agents: vehicles, pedestrians, and cyclists. Agents are dynamic objects that can move as the scenario evolves.
A well-known method is to simply propagate agents according to the logged data. We refer to these as non-reactive log-replay agents. Log-replay agents are used to simulate scenarios both in open-loop and closed-loop, a near-perfect recreation of the recorded data. However, closed-loop simulation quickly diverges if the planner decides to take different actions from what is recorded in the log. Thus, in closed-loop simulation agents can be also simulated with the aim to interact with the ego vehicle and with each other by producing novel simulation states that resemble real agent behaviors. We refer to these as reactive agents. Reactive agents are only relevant in closed-loop simulation and by definition open-loop simulation uses non-reactive agents.
We develop reactive agents following the Intelligent Driver Model (IDM) policy. IDM agents are initialized with the initial pose and velocity of the logged agents. The agents follow the lane center line of the underlying map. The longitudinal control is dictated by the IDM policy. This allows the agents to react to the ego’s actions, as well as other reactive agents. In turn, this reduces false collisions and lets scenarios play out for longer. Note that we only apply this policy to vehicles, while Vulnerable Road Users (VRUs) are replayed from the log. We choose not to model VRUs as reactive agents as their behavior is often uncooperative and thus hard to model.
IV-B Controller
In nuPlan, planners provide a trajectory as a sequence of poses in , without any kinematic feasibility requirement. This trajectory is assumed to be sampled at specific times in the future according to the simulation configuration. To assert kinematic feasibility and prevent users from cheating, we require the use of a controller. nuPlan provides the flexibility to use any controller, such as perfect tracking, which simply interpolates poses along the planned trajectory.
We developed a two-stage controller to propagate the simulation in closed-loop. This controller consists of two parts, a trajectory tracker and a motion model to forward-integrate the simulation. We implement a Linear Quadratic Regulator (LQR) as the tracker. The found optimal control policy is then fed to the second part of the simulation controller, a kinematic bicycle model which is forward-integrated to propagate the simulation state. Alternatively, we also support different trackers in the two-stage controller, such as an iterative-LQR tracker.
IV-C Evaluation
Different metrics and frameworks have been explored for scoring models in prediction and motion planning benchmarks . In this paper, we select a set of metrics and design an aggregation method to compare the performance of planners. In open-loop, we only evaluate the closeness of the planner generated trajectory to the human-driven trajectory. The open-loop metrics are modified Average Displacement Error (ADE), Final Displacement Error (FDE), Average Heading Error (AHE), Final Heading Error (FHE), and Miss Rate (MR) in which we calculate the metrics over different horizons and report their average score. For closed-loop, we use a combination of metrics to evaluate lawfulness and compliance with traffic rules consisting of no at-fault collisions, trajectories inside drivable area, no trajectories in lanes belonging to oncoming traffic, not driving above the speed limit and maintaining enough Time To Collision (TTC) with other road users, metrics to evaluate progress towards the goal and measure the rider comfort. All metrics’ scores are normalized to the range $$ using thresholds that are selected based on legal requirements and natural human driving. A higher score indicates a better performance.
The final score of a planner is computed by averaging the scores for its generated trajectories across all scenarios. The score of a trajectory in a scenario is given by a hybrid weighted average of all metrics’ scores.
The rest of the metrics are weighted according to their importance (See Tab. II) and then averaged to compute the scenario score as:
We define the score for each challenge (open-loop, closed-loop non-reactive, and closed-loop reactive) as the average scenario score across all scenarios for that challenge.
V Experiments
Here we present a number of planning baselines and their results when evaluated on the nuPlan benchmark. We analyze how the planning performance is impacted by lower quality perception inputs, as well as how it generalizes to other cities. Finally, we discuss the new state-of-the-art set by the submissions to the first nuPlan challenge.
We implement several planning methods that are representative of the literature.
The Simple planner has little planning capability. The planner plans a straight line at a constant speed. The only logic of this planner is to decelerate if the current velocity exceeds the max velocity.
IDM Planner
The Intelligent Driver Model (IDM) planner is essentially an Adaptive Cruise Control (ACC) policy . The planner consists of two parts: path planning and longitudinal control. The path planning component is a breadth-first search algorithm. It finds a center-line path toward the mission goal extracted from the underlying map structure. The longitudinal control follows the IDM policy. The policy describes how fast the planner should go based on the distance between itself and the closest leading agent.
Raster ML planner
Similar to the encoder in , the raster planner uses ResNet-50 as the backbone to encode features from an ego-centric multi-channel raster representing the ego, the agents and the map. The model directly outputs the final ego trajectory. The planner does not perform any post-processing on the predicted ego trajectory.
UrbanDriver ML Planner
We adopted an open-loop training variant of the UrbanDriver model as a representative machine learning planner baseline. The model processes vectorized agents and map inputs into local feature descriptors that are passed to a global attention mechanism for yielding a predicted ego trajectory. We train the model using imitation learning to match expert trajectories available in the nuPlan dataset. Data augmentation is additionally performed on the agents and expert trajectory provided during training to mitigate data distribution drift encountered during closed-loop simulation. This version was used for the challenge. We also implemented a multi-step prediction baseline variant as discussed in and originally proposed in to further address the distribution shift, for the experiments in this work but do not open-source this implementation.
V-B Main results
Tab. III shows the planning results for the proposed baselines in each of the three challenge setups. Supervised learning-based planners excel in an open-loop setting. This is unsurprising as the task is akin to the traditional motion forecasting challenge. This suggests that an ML planner can choose to make similar decisions to a human driver in open-loop settings. However, ML planners still struggle to overcome the distribution shift in closed-loop. A closed-loop scenario can develop into a new situation that was never present in the training dataset. Even techniques such as data augmentation and closed-loop training fail to overcome this domain gap. This is evident in both the literature and our experiments. Rule-based planners, on the other hand, face no such issues. Policies like IDM can produce decent driving behavior. This is confirmed by the metrics as it achieved the highest scores for closed-loop. It should be noted though that the reactive agents are also modelled with a similar IDM. The use of similar assumptions on the vehicle behavior may result in giving the IDM planner an unfair advantage over other planners in closed-loop evaluation. It is evident that sufficiently sophisticated rule-based planners still outperform purely learned planners in closed-loop settings.
V-C Perturbation
As the nuPlan dataset is created with offline perception, to capture the original probability distribution of data collected online, we injected uniform noise on the detections. Noise was added in the dimensions and the pose of the detected agents, with variance extracted by comparing offline and online detections. The scores of planners under nominal and noise-injected simulations in closed-loop reactive mode are presented in Tab. IV. A version of UrbanDriver trained on the perturbed data is called UrbanDriverOnline, which shows a performance deterioration compared to the nominal model on both nominal and injected data. This indicates the value of high-quality offline annotations in the dataset and the learning pipeline.
V-D Generalization
The location generalization experiment is shown in Tab. V. The experiment aims to test a model’s generalization capabilities. The UrbanDriver model was trained purely on data from Las Vegas. The model was tested separately on scenarios from Singapore, Boston, Pittsburgh, and Las Vegas. The open-loop performance dropped by 53.8%, while closed-loop non-reactive and closed-loop reactive performance dropped by 35.1% and 41.5% respectively. The worst-performing location is Singapore. This can be explained by the left-hand traffic, while the model was trained on right-hand traffic. One insight is that the correlation between the model’s open-loop and closed-loop performance is relatively weak. The difference between open and closed-loop scores across Singapore, Boston, and Pittsburgh is only 16.3%, while for Las Vegas it is more than double at 37.3%. This indicates that a good motion forecasting model does not translate to closed-loop capabilities. Thus a major challenge is to overcome the domain gap between open-loop and closed-loop before tackling larger generalization problems.
V-E nuPlan challenge
In the nuPlan motion planning challenge contestants create a planner to traverse a set of diverse and challenging scenarios across all four cities. Tab. VI shows the Overall Score, which is the average across Open-loop, Closed-loop Non-reactive, and Closed-loop Reactive challenges of the top four planners. In the open-loop challenge, planners that incorporated supervised learned methods scored relatively well. In the closed-loop challenges, planners employed a combination of learned and handcrafted components. A common theme was the use of a learned model to first predict the ego’s planned trajectory. uses a raster-based model that outputs a spatial-temporal heatmap for the ego and an occupancy map for the surrounding agents. are vector-based using transformers as a backbone. Once the trajectory is obtained, the planner has a further refinement stage used to ensure kinematic feasibility and collision avoidance. The highest-scoring planner in closed-loop was mostly rule-based . It generates a handful of trajectories by perturbing the center line laterally at different velocities. Trajectories are selected with a heuristic that considers factors such as collision, drivable areas, traffic laws, and comfort. An ML-generated trajectory is fused to correct the long-term planned horizon. This limited the influence of the learned model. We draw two conclusions from the challenge results. First, ML-based methods require additional post-processing for closed-loop driving. Second, hybrid methods appear to be the most effective approach, combining traditional and data-driven methods.
VI Conclusion
We presented nuPlan, the first real-world driving benchmark and the largest existing labeled autonomous driving dataset. The dataset consists of 1282 hours of diverse driving scenarios across 4 cities as well as an unprecedented 128 hours of raw sensor data and is accompanied by an evaluation framework powered by a closed-loop simulator; the dataset and the evaluation framework are publicly available. We investigated the state of current rule-based and learned-based planners by evaluating multiple approaches on the nuPlan dataset across challenging driving scenarios. The first public nuPlan challenge demonstrated that rule-based planners outperform purely ML-based ones, but hybrid planners with learned-based components show the most promise in handling difficult scenarios. In the future, we plan to mine for richer long-tail driving scenarios, design scenario-based metrics, provide ML-based planning and agent baselines and explore end-to-end planner training directly from sensor data.
Acknowledgement
We would like to thank the numerous current and former colleagues at Motional who contributed to this work, especially Juraj Kabzan, Edouard Capellier, Abhinav Rai, Xiaoli Meng, Jiong Yang, Lubing Zhou, Abirami Srinivasan, Bing Jui Ho, Hiok Hian Ong, Mitchell Spryn, Taufik Tirtosudiro, Michael Noronha, Vijay Govindarajan, Jeff Wang, Nelly Lyu, Samson Hang, Shashank Chaudhary, Patrick Weygand, Vern Jensen, Abhimanyu Singh, Qiang Xu and Sammy Omari. Furthermore, many external advisors gave useful feedback, including Kashyap Chitta and Oscar de Groot.
References
Supplementary Material
In this supplementary material we provide additional information about the creation of the nuPlan dataset, the simulation framework and evaluation protocol, as well as additional experiments.
VII Dataset
Here we describe the sensor setup, synchronization, dataset splits, semantic map layers and object annotation statistics.
Sensor setup. The nuPlan sensor setup is collected by a fleet of 31 vehicles with identical sensor setup (Tab. VII). The vehicles are equipped with 5 lidars, 8 cameras, as well as GNSS & IMU. Thus both lidars and cameras cover the full 360-degree environment and minimize blindspots (Fig. 5). The lidars can generate lidar point clouds with up to 360k and 720k points for 20-channel and 40-channel lidars, respectively.
Synchronization. To ensure high-quality data alignment between multiple sensors, lidar sensors are synchronized with the system time on the car in multiple lidar periods. The exposure of each camera is triggered in the same way plus a camera-specific offset, where the offset is an optimal value to ensure images are acquired right when the lidars are sweeping through each camera’s FOV. Given there are multiple lidar sweeps in a spin, we merge a full sweep from each lidar sensor into a merged sweep since they complete a revolution at nearly the same instance. After they are merged, the merged point clouds are transformed to the same frame coordinate and timestamp. We also perform ego motion compensation to take ego position into account when the points are acquired.
Dataset splits. The dataset splits (train, val, and test) are geographically overlapping, which means they can cover the same region in a city. However, to minimize temporal data leakage, data from the same day and city is not shared across splits. In addition, the dataset is stratified across days, cities, and driving scenarios to ensure all splits have similar balances across these dimensions.
Map layers. We define the semantic map layers in nuPlan. The map layers are designed with planning in mind. A graph can be constructed from the network of lane and lane connectors. Additionally, the map is annotated with important road features (stop lines, crosswalks, car parks) that impact the decision-making of the AV. See Tab. VIII for more details.
Annotation statistics. Here we present more statistics for the annotations in nuPlan. Fig. 6 shows the number of tracks for each of the 6 object classes in nuPlan. While there are about 2 orders of magnitude between the most (generic object) and least (construction zone sign) common class, the dataset is nevertheless less long-tailed than other datasets with more fine-grained classes . Furthermore, even the rarest class still has more than tracks, which shows the advantage of such a large dataset.
We also present size statistics for different classes in Fig. 7. These show that our boxes are statistically stable in their sizes except for the vehicle class that includes construction vehicles with irregular shapes.
Fig. 8 shows the absolute velocities for three classes (vehicle, pedestrian, and bicycle), since other classes are static most of the time. We observe that most of the objects are slow-moving, which represents primarily urban scenarios, while a small number of pedestrians and bicycles are moving at unrealistically high speeds, possibly due to noisy annotations.
Fig. 9 shows the object (box) orientations relative to the ego vehicle orientation. Due to the grid-like road layout, most vehicles have orientations that are multiples of .
Fig. 10 shows the spatial coverage of our ego vehicles across all maps. In Las Vegas, most routes start and end in PUDOs, which leads to these areas being visited more often. In other cities, the distribution is more uniform, with key intersections being visited the most.
VII-B Autolabeling
As described in the main paper, We developed an offboard perception system to generate the bounding boxes and tracks for the objects in the scene. It consists of three stages: object detection, offline tracking, and global track refinement. The system is deployed on the raw sensor data from nuPlan dataset.
Object detection. For the first stage, we extend the state-of-the-art MVF++ as our lidar-based 3D object detector. We select past frames () and compensate the points based on ego-motion, which significantly densifies the point clouds in the scene. We find that there is no improvement in incorporating future frames as well. To further boost the performance, we make several changes to the architecture. First, different from other works , we propose to include the points from four side lidars in addition to the top lidar, so the objects closer to the ego vehicle will not be missed. To use these points properly, we adopt the cylindrical view in MVF++, instead of the perspective view, to mitigate the collision of points from different lidars. Second, we enlarge the model capacity by using a RegNet backbone. The RegNet backbone processes voxelized outputs from MVF++ and generates expressive feature maps for final detection heads. Third, a CenterPoint detection head is applied to predict the bounding boxes. Since the resulting object boxes are often flipped, we perform majority voting using past and future detections and adjust the heading accordingly.
Offline tracking. Given the instantaneous bounding boxes generated from the first stage, a Kalman filter based multi-object tracker is applied to associate the boxes across timestamps. Since the tracker is running in offline mode, we leverage more information from both past and future frames to manage tracks: First, after frames of detections are used to confirm a track, the track is initiated from the first frame of frames, instead of the last one . As a result, the number of False Negatives is reduced. Second, compared with online tracking , we increase to maintain a longer memory before coasting a track. While a higher slows down the method, the number of ID switches is reduced significantly.
Global track refinement. Even though the detection network takes frames of lidar point clouds as input, it still perceives the scene in a short time window. The resulting points are limited to a certain perspective, posing challenges to estimating the size and heading accurately. Moreover, although the bounding boxes from the detection network have been smoothed to some extent by a Kalman Filter based tracker, the smoothing does not happen at a global scale. To mitigate this effect, we introduce a novel neural network. Inspired by recent works , our network loads tracks along with the aggregated point clouds within the tracks and refines the bounding boxes in the tracks in terms of position, heading, size and velocity. Notably, instead of using two networks for static and dynamic objects , or for different purposes (size, position or heading) , we design a single network, which achieves the above goals in one shot. In practice, the proposed method reduces the deployment overhead compared to other works , because the order of the two networks is not clear and the classification of dynamic/static is arbitrary for some slow-moving objects. Similar to these works, we also only apply this network to the vehicle class.
Implementation Details. For simplicity, we consider objects on a 2D bird’s eye view plane only and denote as the object’s position, as size, as velocity and as heading. An object at timestamp in a track is denoted as . It has a bounding box that tightly contains the points . We select the boxes from the past and future frames. These bounding boxes are concatenated as the input for the trajectory encoding branch. Meanwhile, the point clouds are cropped in both past and future frames and then transformed to the current frame . The aggregated point cloud is processed by dynamic voxelization and taken as input by the point encoding branch. Convolutional layers and global average pooling are applied for each branch to generate expressive features, which are concatenated for the final MLP. In training, we use a smooth L1 loss to regress the residuals between the ground truth and input . In deployment, after getting the refined bounding boxes for the whole track, we choose the median value of the size for all the bounding boxes. We update the position and velocity of each box according to the outputs of the network.
VII-C Evaluation of the autolabeling system
Dataset. To evaluate the proposed autolabeling system we use an internal dataset that follows a similar data distribution to nuPlan and contains 3098 human-labeled scenes collected from Singapore, Boston, Pittsburgh and Las Vegas. Among them, 2958 scenes are used for training, 49 for validation and 91 for testing.
Metrics. The performance of the detection network and the global track refinement are evaluated using the maximum F1 scores for all the classes. The performance of the offline tracker is evaluated using AMOTA, ID switches, recall and F1 score. Vehicles are divided into two categories depending on whether their lengths are larger than 7m. We use Birds Eye View 2D IOU as the matching criterion. The IOU threshold is for vehicles, for cyclists and pedestrians, for barriers and generic objects, for traffic cones.
3d object detection. The detection network proposed in this paper is compared with a modified CenterPoint detector, termed mCenterPoint. Our goal is to compare our proposed network with a detection network often deployed onboard for autonomous vehicles. We choose mCenterPoint as a baseline because it is lightweight and meets the real-time requirement. It differs from our offline detection network in three aspects. First, it aggregates past 3 sweeps of point clouds as input to the network. Second, mCenterPoint uses the vanilla PointPillars for voxelization. Third, a more compact backbone is used to extract features for the CenterPoint detection heads. In contrast, the offline model is applied on 7 sweeps of point clouds with a multi-view encoder and a heavy backbone . The max F1 scores for each class are shown in Tab. IX. From the table, we can see the proposed detection network can outperform mCenterPoint by a large margin. It clearly shows that using more sweeps of point clouds as well as a more complex network architecture improves the performance drastically.
Offline tracker. We also implemented a modified version of AB3DMOT as a baseline in the tracking experiments. Similar to mCenterPoint, mAB3DMOT is configured for real-time deployment onboard for autonomous vehicles. In particular, and are set to be and for mAB3DMOT, respectively. In comparison, the offline tracker uses and back-traces the first 2 frames with . From Tab. X we see a significant improvement on all tracking metrics. In particular, the number of ID switches has been reduced by , which is due to the significantly enlarged . The fewer number of ID switches indicates that the tracks in nuPlan dataset are less fragmented and thus reflect the agents’ movement in the real world.
Global track refinement. To evaluate global track refinement, we compare it against the outputs of the offline tracker. Because the offline tracker will interpolate and suppress some bounding boxes from the detection network, the F1 scores reported here are different from Tab. IX. The results are presented in Tab. XI. It is shown that the vehicle F1 score is consistently improved by global track refinement. In particular, when a stricter matching criterion is applied (BEV IOU = ), we see a relative boost of , which demonstrates the efficacy and necessity of global track refinement in producing high-quality bounding boxes.
VII-D Traffic lights
To automatically label traffic light statuses within nuPlan, we make use of the tracks produced by our autolabeling system, as well as the human-annotated map information. The map indicates the intersections with traffic lights. Within each traffic light intersection, the individual traffic lights control the flow of vehicles from a lane on one side of the intersection to another lane. Each such pair of lanes is connected via a lane connector in our map. Hence, we encode the status of the traffic lights in the corresponding lane connectors.
Green and amber statuses. We consider both green and amber statues of traffic lights to be the same, i.e. green, since vehicles are allowed to move under both light statuses. To infer the presence of a green traffic light on a particular lane, we determine if there are agents moving along the lane connector corresponding to the traffic light. We do this by comparing the directed Hausdorff distance between the trajectories of all the agents within the traffic light intersection, and the lane connector. Note that as the traffic light labeling system is offline, we are able to make use of both the past and future trajectories of each agent. If the directed Hausdorff distance is small, we surmise that there are agents moving along the lane connector, and thus the traffic light controlling that lane connector is likely green.
Red statuses. To infer the presence of a red traffic light on a particular lane, we check for the minimum speed of the agents that are on that lane. Note that we only take into account agents which are of a certain distance from the traffic light intersection. If the minimum speed of the agents on that lane is low, we surmise that the agents are stopped. We also consider the deceleration of the agents on the various lanes. When a traffic light is red, it is common for drivers to begin decelerating as they approach the intersection. Thus, we set a threshold for the deceleration magnitude above which we consider the agent in a lane to be decelerating. If we find that either the agents on a given lane are stopped or decelerating, then we infer that the traffic light controlling that lane connector is likely red. For lanes and lane connectors for which there are no observable agents, we set the status of the corresponding traffic light to ‘unknown‘.
Post-processing. After inferring the per-frame green and red statuses, we perform post-processing to refine the inferences. First, we perform grouping. Our map stores information on lane connectors that go in a ”parallel” direction, and therefore share the same traffic light statuses. We use this information to set all lane connectors in the parallel direction to have the same status. This helps to reduce false negatives, especially in situations when there are no observed agents on some of the lanes going in the same direction.
Second, we perform back-filling. When a traffic light changes from red to green, drivers often have a certain reaction time before moving off. Similarly, when a traffic light changes from yellow to red, there may be drivers still crossing the intersection. This might lead to the system inferring a green traffic light for the particular frame, even though the lights have changed in reality. Hence there is a slight lag in the transitions identified via motion inference with respect to the actual transition. To account for this, for all green statuses, we go back a specified time horizon into the past and override the statuses of each lane connector by setting them to green. We do the same for the red statuses and override the past statuses within the specified time horizon with red.
Evaluation. To perform a more quantitative evaluation of our traffic light labeling system, we select scenes where the ego vehicle is moving through traffic light intersections and manually label the traffic light statuses for a subset of the data. We select scenes from various cities and various traffic light behaviors (e.g. constant, transition). We manually label about 1000 frames. We seek to compare our system against a more conventional traffic light detection system. Such a system is usually deep-learning based and vision-only. Consequently, for the baseline, we use YOLOv3 , as implemented by and trained on the LISA Traffic Light Dataset , which contains 113,888 annotated traffic lights. We do not fine-tune YOLOv3 on any nuPlan data. By labeling the statuses of the traffic lights via motion inferences, we are able to recover x more traffic light statuses than YOLOv3. We find that YOLOv3 misses several traffic lights at long distances. It also performs poorly at very close distances to the traffic light due to the perspective warping of the cameras. Using YOLOv3 , the traffic light classification accuracy among visible traffic lights is . In contrast, our system performs slightly better with an accuracy of . We find that the transitions of the traffic light statuses are difficult to infer for our system, due to the lag in the agents moving off, slowing down or stopping whenever the traffic lights change. A data-driven alternative could provide better results here.
VII-E Scenario mining
We propose the following approach to mine scenarios from the nuPlan dataset. First, we compute a large number of atomic primitives. These primitives model an attribute (vehicle speed) or state transition (a vehicle being in two lanes simultaneously) and can be extracted from the entire dataset in a single pass. Second, we combine several primitives into an SQL query that can be run efficiently on a database. Third, we post-process the query results by including a sequence of seconds before and after the returned time step. Finally, we manually QA examples for each scenario. If the false positive rate is below , we either refine the query by adding more attributes or tuning hyperparameters or discard the scenario. This approach results in high-precision scenario labels, while recall is less relevant for us. We have a total of unique scenario types in the dataset, and their distributions per city are shown in Fig. 11.
The details and parameters of the 14 scenarios that were used to grade the nuPlan planning challenge can be found in Tab. XII.
VIII Simulation
The differential equation behind the IDM policy operates based on a focus agent - the vehicle that the IDM policy is controlling - and a lead object. The policy is parameterized by the distance to a lead object. However, in closed-loop simulation identifying the correct lead object is not trivial. Fig. 12 shows how the closest leading object is identified. Searching for the closest object often returns another agent in the adjacent lane. Hence, the search space is reduced to only the other objects in the scene that are or can potentially intersect with the agent’s planned path. The euclidean distance between the two closest points of the focus agent and the lead object is computed.
Merging situations can be tricky because IDM does not inherently account for multi-lane interactions. One way to simulate this is to extend all other objects’ footprints in the scene proportional to their speed. Fig. 13 shows an example of a situation where an extended footprint can help in lane-merge situations. The lead agent search will identify the oncoming agent. The agent will know to slow down sooner, giving way to the oncoming objects.
Breadth-first search path planning for a set distance.
Perform a one-step forward euler numerical integration on the IDM differential equations.
Propagate the agent along the planned path according to the solution of step 3.
Repeat for all simulation propagation steps.
The attempt to re-purpose prevailing motion forecasting models such as LaneGCN for traffic simulation proved less trivial than initially thought. The model independently predicts agents, resulting in a lack of scene cohesion, for example, predicting colliding trajectories. The model is also susceptible to distribution shift issues that arise in closed-loop simulation. When applied to traffic simulation, the model induces unrealistic driving scenarios . Furthermore, the number of simulated agents had to be limited to maintain acceptable simulation runtime. It can be concluded that motion forecasting alone cannot achieve realistic, scene-coherent traffic simulation.
VIII-B Evaluation
Metrics. As explained in the main paper, we design a set of metrics along with a scoring function to compare the performance of planners. The previously mentioned open-loop metrics are described in detail in Tab. XIII. For each metric, we evaluate an aggregated error/MR considering different time horizons within the planning horizon (i.e., ) and with the same sampling frequency of . This method allows us to have a fair comparison across planners with different planning horizons or sampling rates. Additionally, by taking the mean across the selected horizons, the errors at the beginning of the horizon will have more impact on the averaged value. The ”within bound” metric score used in the cost structure is found by comparing the average error/MR value to a maximum acceptable threshold ( for distance errors, for heading errors, and for MR). It is if the average value is more than the threshold, and otherwise. For example, ‘MR within bound’ from open-loop metrics, and ‘drivable area compliance’ from the closed-loop metrics. Multiplier metrics are assigned a score of 1 or 0, except for ‘no at-fault collision’ which takes 0 (if there is an at-fault collision with a vehicle, bicycle, or pedestrian, or there are multiple at-fault collisions with objects), 0.5 (if there’s an at-fault collision with a single object), and 1 (if there’s no at-fault collision).
In the following, we include additional information about closed-loop metrics mentioned in the main paper:
When identifying at-fault collisions, we only penalize the planner when the ego vehicle could be responsible for the collision, which includes collisions with stopped agents, collisions with agents in front of ego and collisions with agents in adjacent lanes while making a lane change. On the other hand, the ego is not penalized for rear-end collisions or other agents colliding with the ego when it is stopped. To further emphasize the importance of different agent types, at-fault collisions are grouped into vulnerable road users (including pedestrians and bicyclists), vehicles and objects (traffic cones, barriers and generic objects).
For drivable area violation, we measure the maximum distance of the corners of the ego bounding box from the nearest drivable area.
For driving direction, the movement of the ego during a 1s time horizon is calculated along the driving direction of its lane.
Speed limit violation is defined based on the magnitude and duration of the violation.
Time to collision is defined as the time required for ego and another object to collide if they continue at their present speed and heading. We only compute time to collision for objects in front of the ego, cross-traffic objects and lateral objects on the sides, when the ego is making a lane change or is in the intersection.
Rider comfort is measured based on jerk, acceleration and steering rate which are compared to those observed in human driving.
Progress of the planner trajectory towards the goal is evaluated by comparing its progress along expert’s route in the same scenario. The metric quantifies the progress as the ratio of overall ego progress to the overall expert’s progress during the scenario.
Final score structure. The final score of a planner is computed by averaging its scores across all scenarios as defined in the main paper. What follows helps explain the equation: For open-loop planners, the driven trajectory in a scenario is assigned a zero score if the miss rate is above the selected threshold (0.3), otherwise, a weighted average of other metrics’ scores is used as the score. All weights were tuned with the objective to maximize the overall performance of the planner against human-driven future trajectories across multiple scenarios. The weights can be found in the main paper. For closed-loop planners, the driven trajectory in a scenario is assigned:
A zero score if 1) there is an at-fault collision with a vehicle or a VRU, or 2) there are multiple at-fault collisions with objects (e.g. a cone), or 3) there is a drivable area violation, or 4) ego drives into oncoming traffic more than (driving distance), or 5) ego progress towards the destination is smaller than a threshold.
The weighted average of other metrics’ scores is multiplied with if there is one at-fault collision with an object (e.g. a cone), or if ego drives into oncoming traffic more than , but less than .
A weighted average of other metrics’ scores, otherwise.
Closed-loop metric scores are summarized in Tab. XIV. is a function that returns 1 if there are no speed limit violations and approaches 0 as the violation increases. Furthermore, the comfort metric accounts for jerk amplitude, lateral and longitudinal acceleration, and jerk, and yaw rate and acceleration. It is assigned a zero score if one of the comfort bounds is violated. Our scoring heuristic described above is an initial proposal that accounts for the natural importance of each metric and is hand-tuned for our dataset and simulation framework.
IX Experiments
Detailed scenario-stratified metrics for all four planner baselines (rule-based and learned) across the three challenges (open-loop, closed-loop non-reactive, and closed-loop reactive) can be found in Tab. XV. The metric breakdown corroborates the statement in the main paper that metrics reward conservative driving. Planners that do not collide and stay within the drivable area may score better than planners that attempt to mimic human drivers. Hence, finer grain and scenario-based metrics are required to distinguish between simple rules-abiding driving from desirable human-like driving behaviors.
IX-B nuPlan Challenge
The distribution of planners’ scores across the nuPlan challenge can be seen in Fig. 14 for open-loop, Fig. 15 for closed-loop non-reactive and Fig. 16 for closed-loop reactive. Most submitted planners were able to score well in the open-loop challenge. The most common scores lie between 0.8 - 0.85. The scores significantly dropped for closed-loop challenges. The most common scores lie between 0.65 - 0.7. A 0.15 drop from the open-loop challenge. This further supports the argument that most purely learned models fail to generalize to closed-loop scenarios. For an overview of the leaderboard, please refer to https://eval.ai/web/challenges/challenge-page/1856/leaderboard/4360.