nuScenes: A multimodal dataset for autonomous driving

Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, Oscar Beijbom

Introduction

Autonomous driving has the potential to radically change the cityscape and save many human lives . A crucial part of safe navigation is the detection and tracking of agents in the environment surrounding the vehicle. To achieve this, a modern self-driving vehicle deploys several sensors along with sophisticated detection and tracking algorithms. Such algorithms rely increasingly on machine learning, which drives the need for benchmark datasets. While there is a plethora of image datasets for this purpose (Table 1), there is a lack of multimodal datasets that exhibit the full set of challenges associated with building an autonomous driving perception system. We released the nuScenes dataset to address this gapnuScenes teaser set released Sep. 2018, full release in March 2019..

Multimodal datasets are of particular importance as no single type of sensor is sufficient and the sensor types are complementary. Cameras allow accurate measurements of edges, color and lighting enabling classification and localization on the image plane. However, 3D localization from images is challenging . Lidar pointclouds, on the other hand, contain less semantic information but highly accurate localization in 3D . Furthermore the reflectance of lidar is an important feature . However, lidar data is sparse and the range is typically limited to 50-150m. Radar sensors achieve a range of 200-300m and measure the object velocity through the Doppler effect. However, the returns are even sparser than lidar and less precise in terms of localization. While radar has been used for decades , we are not aware of any autonomous driving datasets that provide radar data.

Since the three sensor types have different failure modes during difficult conditions, the joint treatment of sensor data is essential for agent detection and tracking. Literature even suggests that multimodal sensor configurations are not just complementary, but provide redundancy in the face of sabotage, failure, adverse conditions and blind spots. And while there are several works that have proposed fusion methods based on cameras and lidar , PointPillars showed a lidar-only method that performed on par with existing fusion based methods. This suggests more work is required to combine multimodal measurements in a principled manner.

In order to train deep learning methods, quality data annotations are required. Most datasets provide 2D semantic annotations as boxes or masks (class or instance) . At the time of the initial nuScenes release, only a few datasets annotated objects using 3D boxes , and they did not provide the full sensor suite. Following the nuScenes release, there are now several sets which contain the full sensor suite (Table 1). Still, to the best of our knowledge, no other 3D dataset provides attribute annotations, such as pedestrian pose or vehicle state.

Existing AV datasets and vehicles are focused on particular operational design domains. More research is required on generalizing to “complex, cluttered and unseen environments” . Hence there is a need to study how detection methods generalize to different countries, lighting (daytime vs. nighttime), driving directions, road markings, vegetation, precipitation and previously unseen object types.

Contextual knowledge using semantic maps is also an important prior for scene understanding . For example, one would expect to find cars on the road, but not on the sidewalk or inside buildings. With the notable exception of , most AV datasets do not provide semantic maps.

From the complexities of the multimodal 3D detection challenge, and the limitations of current AV datasets, a large-scale multimodal dataset with coverage across all vision and range sensors collected from diverse situations alongside map information would boost AV scene-understanding research further. nuScenes does just that, and it is the main contribution of this work.

nuScenes represents a large leap forward in terms of data volumes and complexities (Table 1), and is the first dataset to provide sensor coverage from the entire sensor suite. It is also the first AV dataset to include radar data and captured using an AV approved for public roads. It is further the first multimodal dataset that contains data from nighttime and rainy conditions, and with object attributes and scene descriptions in addition to object class and location. Similar to , nuScenes is a holistic scene understanding benchmark for AVs. It enables research on multiple tasks such as object detection, tracking and behavior modeling in a range of conditions.

Our second contribution is new detection and tracking metrics aimed at the AV application. We train 3D object detectors and trackers as a baseline, including a novel approach of using multiple lidar sweeps to enhance object detection. We also present and analyze the results of the nuScenes object detection and tracking challenges.

Third, we publish the devkit, evaluation code, taxonomy, annotator instructions, and database schema for industry-wide standardization. Recently, the Lyft L5 dataset adopted this format to achieve compatibility between the different datasets. The nuScenes data is published under CC BY-NC-SA 4.0 license, which means that anyone can use this dataset for non-commercial research purposes. All data, code, and information is made available onlinegithub.com/nutonomy/nuscenes-devkit.

Since the release, nuScenes has received strong interest from the AV community . Some works extended our dataset to introduce new annotations for natural language object referral and high-level scene understanding . The detection challenge enabled lidar based and camera based detection works such as , that improved over the state-of-the-art at the time of initial release by 40%40\% and 81%81\% (Table 4). nuScenes has been used for 3D object detection , multi-agent forecasting , pedestrian localization , weather augmentation , and moving pointcloud prediction . Being still the only annotated AV dataset to provide radar data, nuScenes encourages researchers to explore radar and sensor fusion for object detection .

2 Related datasets

The last decade has seen the release of several driving datasets which have played a huge role in scene-understanding research for AVs. Most datasets have focused on 2D annotations (boxes, masks) for RGB camera images. CamVid , Cityscapes , Mapillary Vistas , D2D^{2}-City , BDD100k and Apolloscape released ever growing datasets with segmentation masks. Vistas, D2D^{2}-City and BDD100k also contain images captured during different weather and illumination settings. Other datasets focus exclusively on pedestrian annotations on images . The ease of capturing and annotating RGB images have made the release of these large image-only datasets possible.

On the other hand, multimodal datasets, which are typically comprised of images, range sensor data (lidars, radars), and GPS/IMU data, are expensive to collect and annotate due to the difficulties of integrating, synchronizing, and calibrating multiple sensors. KITTI was the pioneering multimodal dataset providing dense pointclouds from a lidar sensor as well as front-facing stereo images and GPS/IMU data. It provides 200k 3D boxes over 22 scenes which helped advance the state-of-the-art in 3D object detection. The recent H3D dataset includes 160 crowded scenes with a total of 1.1M 3D boxes annotated over 27k frames. The objects are annotated in the full view, as opposed to KITTI where an object is only annotated if it is present in the frontal view. The KAIST multispectral dataset is a multimodal dataset that consists of RGB and thermal camera, RGB stereo, 3D lidar and GPS/IMU. It provides nighttime data, but the size of the dataset is limited and annotations are in 2D. Other notable multimodal datasets include providing driving behavior labels, providing place categorization labels and providing raw data without semantic labels.

After the initial nuScenes release, followed to release their own large-scale AV datasets (Table 1). Among these datasets, only the Waymo Open dataset provides significantly more annotations, mostly due to the higher annotation frequency (10Hz10\text{Hz} vs. 2Hz2\text{Hz})In preliminary analysis we found that annotations at 2Hz2\text{Hz} are robust to interpolation to finer temporal resolution, like 10Hz10\text{Hz} or 20Hz20\text{Hz}. A similar conclusion was drawn for H3D where annotations are interpolated from 2Hz2\text{Hz} to 10Hz10\text{Hz}.. A*3D takes an orthogonal approach where a similar number of frames (39k) are selected and annotated from 55 hours of data. The Lyft L5 dataset is most similar to nuScenes. It was released using the nuScenes database schema and can therefore be parsed using the nuScenes devkit.

The nuScenes dataset

Here we describe how we plan drives, setup our vehicles, select interesting scenes, annotate the dataset and protect the privacy of third parties.

We drive in Boston (Seaport and South Boston) and Singapore (One North, Holland Village and Queenstown), two cities that are known for their dense traffic and highly challenging driving situations. We emphasize the diversity across locations in terms of vegetation, buildings, vehicles, road markings and right versus left-hand traffic. From a large body of training data we manually select 84 logs with 15h of driving data (242km travelled at an average of 16km/h). Driving routes are carefully chosen to capture a diverse set of locations (urban, residential, nature and industrial), times (day and night) and weather conditions (sun, rain and clouds).

Car setup.

We use two Renault Zoe supermini electric cars with an identical sensor layout to drive in Boston and Singapore. See Figure 4 for sensor placements and Table 2 for sensor details.

Front and side cameras have a FOV and are offset by . The rear camera has a FOV of .

Sensor synchronization.

To achieve good cross-modality data alignment between the lidar and the cameras, the exposure of a camera is triggered when the top lidar sweeps across the center of the camera’s FOV. The timestamp of the image is the exposure trigger time; and the timestamp of the lidar scan is the time when the full rotation of the current lidar frame is achieved. Given that the camera’s exposure time is nearly instantaneous, this method generally yields good data alignmentThe cameras run at 12Hz12\text{Hz} while the lidar runs at 20Hz20\text{Hz}. The 1212 camera exposures are spread as evenly as possible across the 2020 lidar scans, so not all lidar scans have a corresponding camera frame.. We perform motion compensation using the localization algorithm described below.

Localization.

Most existing datasets provide the vehicle location based on GPS and IMU . Such localization systems are vulnerable to GPS outages, as seen on the KITTI dataset . As we operate in dense urban areas, this problem is even more pronounced. To accurately localize our vehicle, we create a detailed HD map of lidar points in an offline step. While collecting data, we use a Monte Carlo Localization scheme from lidar and odometry information . This method is very robust and we achieve localization errors of ≤10cm\leq 10cm. To encourage robotics research, we also provide the raw CAN bus data (e.g. velocities, accelerations, torque, steering angles, wheel speeds) similar to .

Maps.

We provide highly accurate human-annotated semantic maps of the relevant areas. The original rasterized map includes only roads and sidewalks with a resolution of 10px/m10\text{px/m}. The vectorized map expansion provides information on 11 semantic classes as shown in Figure 3, making it richer than the semantic maps of other datasets published since the original release . We encourage the use of localization and semantic maps as strong priors for all tasks. Finally, we provide the baseline routes - the idealized path an AV should take, assuming there are no obstacles. This route may assist trajectory prediction , as it simplifies the problem by reducing the search space of viable routes.

Scene selection.

After collecting the raw sensor data, we manually select 10001000 interesting scenes of 20s20s duration each. Such scenes include high traffic density (e.g. intersections, construction sites), rare classes (e.g. ambulances, animals), potentially dangerous traffic situations (e.g. jaywalkers, incorrect behavior), maneuvers (e.g. lane change, turning, stopping) and situations that may be difficult for an AV. We also select some scenes to encourage diversity in terms of spatial coverage, different scene types, as well as different weather and lighting conditions. Expert annotators write textual descriptions or captions for each scene (e.g.: “Wait at intersection, peds on sidewalk, bicycle crossing, jaywalker, turn right, parked cars, rain”).

Data annotation.

Having selected the scenes, we sample keyframes (image, lidar, radar) at 2Hz2\text{Hz}. We annotate each of the 23 object classes in every keyframe with a semantic category, attributes (visibility, activity, and pose) and a cuboid modeled as x, y, z, width, length, height and yaw angle. We annotate objects continuously throughout each scene if they are covered by at least one lidar or radar point. Using expert annotators and multiple validation steps, we achieve highly accurate annotations. We also release intermediate sensor frames, which are important for tracking, prediction and object detection as shown in Section 4.2. At capture frequencies of 12Hz12\text{Hz}, 13Hz13\text{Hz} and 20Hz20\text{Hz} for camera, radar and lidar, this makes our dataset unique. Only the Waymo Open dataset provides a similarly high capture frequency of 10Hz10\text{Hz}.

Annotation statistics.

Our dataset has 23 categories including different vehicles, types of pedestrians, mobility devices and other objects (Figure 8-SM). We present statistics on geometry and frequencies of different classes (Figure 9-SM). Per keyframe there are 7 pedestrians and 20 vehicles on average. Moreover, 40k keyframes were taken from four different scene locations (Boston: 55%, SG-OneNorth: 21.5%, SG-Queenstown: 13.5%, SG-HollandVillage: 10%) with various weather and lighting conditions (rain: 19.4%, night: 11.6%). Due to the finegrained classes in nuScenes, the dataset shows severe class imbalance with a ratio of 1:10k for the least and most common class annotations (1:36 in KITTI). This encourages the community to explore this long tail problem in more depth.

Figure 5 shows spatial coverage across all scenes. We see that most data comes from intersections. Figure 10-SM shows that car annotations are seen at varying distances and as far as 80m from the ego-vehicle. Box orientation is also varying, with the most number in vertical and horizontal angles for cars as expected due to parked cars and cars in the same lane. Lidar and radar points statistics inside each box annotation are shown in Figure 14-SM. Annotated objects contain up to 100 lidar points even at a radial distance of 80m and at most 12k lidar points at 3m. At the same time they contain up to 40 radar returns at 10m and 10 at 50m. The radar range far exceeds the lidar range at up to 200m.

Tasks & Metrics

The multimodal nature of nuScenes supports a multitude of tasks including detection, tracking, prediction & localization. Here we present the detection and tracking tasks and metrics. We define the detection task to only operate on sensor data between [t−0.5,t][t-0.5,t] seconds for an object at time tt, whereas the tracking task operates on data between [0,t][0,t].

The nuScenes detection task requires detecting 10 object classes with 3D bounding boxes, attributes (e.g. sitting vs. standing), and velocities. The 10 classes are a subset of all 23 classes annotated in nuScenes (Table 5-SM).

We use the Average Precision (AP) metric , but define a match by thresholding the 2D center distance dd on the ground plane instead of intersection over union (IOU). This is done in order to decouple detection from object size and orientation but also because objects with small footprints, like pedestrians and bikes, if detected with a small translation error, give IOU (Figure 7). This makes it hard to compare the performance of vision-only methods which tend to have large localization errors .

True Positive metrics.

In addition to AP, we measure a set of True Positive metrics (TP metrics) for each prediction that was matched with a ground truth box. All TP metrics are calculated using d=2d=2m center distance during matching, and they are all designed to be positive scalars. In the proposed metric, the TP metrics are all in native units (see below) which makes the results easy to interpret and compare. Matching and scoring happen independently per class and each metric is the average of the cumulative mean at each achieved recall level above 10%10\%. If 10%10\% recall is not achieved for a particular class, all TP errors for that class are set to 11. The following TP errors are defined:

Average Translation Error (ATE) is the Euclidean center distance in 2D (units in metersmeters). Average Scale Error (ASE) is the 3D intersection over union (IOU) after aligning orientation and translation (1−IOU1-IOU). Average Orientation Error (AOE) is the smallest yaw angle difference between prediction and ground truth (radiansradians). All angles are measured on a full 360∘360^{\circ} period except for barriers where they are measured on a 180∘180^{\circ} period. Average Velocity Error (AVE) is the absolute velocity error as the L2 norm of the velocity differences in 2D (m/sm/s). Average Attribute Error (AAE) is defined as 1 minus attribute classification accuracy (1−acc1-acc). For each TP metric we compute the mean TP metric (mTP) over all classes:

We omit measurements for classes where they are not well defined: AVE for cones and barriers since they are stationary; AOE of cones since they do not have a well defined orientation; and AAE for cones and barriers since there are no attributes defined on these classes.

nuScenes detection score.

mAP with a threshold on IOU is perhaps the most popular metric for object detection . However, this metric can not capture all aspects of the nuScenes detection tasks, like velocity and attribute estimation. Further, it couples location, size and orientation estimates. The ApolloScape 3D car instance challenge disentangles these by defining thresholds for each error type and recall threshold. This results in 10×310\times 3 thresholds, making this approach complex, arbitrary and unintuitive. We propose instead consolidating the different error types into a scalar score: the nuScenes detection score (NDS).

2 Tracking

In this section we present the tracking task setup and metrics. The focus of the tracking task is to track all detected objects in a scene. All detection classes defined in Section 3.1 are used, except the static classes: barrier, construction and trafficcone.

Weng and Kitani presented a similar 3D MOT benchmark on KITTI . They point out that traditional metrics do not take into account the confidence of a prediction. Thus they develop Average Multi Object Tracking Accuracy (AMOTA) and Average Multi Object Tracking Precision (AMOTP), which average MOTA and MOTP across all recall thresholds. By comparing the KITTI and nuScenes leaderboards for detection and tracking, we find that nuScenes is significantly more difficult. Due to the difficulty of nuScenes, the traditional MOTA metric is often zero. In the updated formulation sMOTAr\text{sMOTA}_{r}Pre-prints of this work referred to sMOTAr\text{sMOTA}_{r} as MOTAR., MOTA is therefore augmented by a term to adjust for the respective recall:

This is to guarantee that sMOTAr\text{sMOTA}_{r} values span the entire $range.Weperform40−pointinterpolationintherecallrangerange. We perform 40-point interpolation in the recall range[0.1,1](therecallvaluesaredenotedas(the recall values are denoted as\mathcal{R}$). The resulting sAMOTA metric is the main metric for the tracking task:

Traditional metrics.

We also use traditional tracking metrics such as MOTA and MOTP , false alarms per frame, mostly tracked trajectories, mostly lost trajectories, false positives, false negatives, identity switches, and track fragmentations. Similar to , we try all recall thresholds and then use the threshold that achieves highest sMOTAr\text{sMOTA}_{r}.

TID and LGD metrics.

In addition, we devise two novel metrics: Track initialization duration (TID) and longest gap duration (LGD). Some trackers require a fixed window of past sensor readings or perform poorly without a good initialization. TID measures the duration from the beginning of the track until the time an object is first detected. LGD computes the longest duration of any detection gap in a track. If an object is not tracked, we assign the entire track duration as TID and LGD. For both metrics, we compute the average over all tracks. These metrics are relevant for AVs as many short-term track fragmentations may be more acceptable than missing an object for several seconds.

Experiments

In this section we present object detection and tracking experiments on the nuScenes dataset, analyze their characteristics and suggest avenues for future research.

We present a number of baselines with different modalities for detection and tracking.

To demonstrate the performance of a leading algorithm on nuScenes, we train a lidar-only 3D object detector, PointPillars . We take advantage of temporal data available in nuScenes by accumulating lidar sweeps for a richer pointcloud as input. A single network was trained for all classes. The network was modified to also learn velocities as an additional regression target for each 3D box. We set the box attributes to the most common attribute for each class in the training data.

Image detection baseline.

To examine image-only 3D object detection, we re-implement the Orthographic Feature Transform (OFT) method. A single OFT network was used for all classes. We modified the original OFT to use a SSD detection head and confirmed that this matched published results on KITTI. The network takes in a single image from which the full predictions are combined together from all 6 cameras using non-maximum suppression (NMS). We set the box velocity to zero and attributes to the most common attribute for each class in the train data.

Detection challenge results.

We compare the results of the top submissions to the nuScenes detection challenge 2019. Among all submissions, Megvii gave the best performance. It is a lidar based class-balanced multi-head network with sparse 3D convolutions. Among image-only submissions, MonoDIS was the best, significantly outperforming our image baseline and even some lidar based methods. It uses a novel disentangling 2D and 3D detection loss. Note that the top methods all performed importance sampling, which shows the importance of addressing the class imbalance problem.

Tracking baselines.

We present several baselines for tracking from camera and lidar data. From the detection challenge, we pick the best performing lidar method (Megvii ), the fastest reported method at inference time (PointPillars ), as well as the best performing camera method (MonoDIS ). Using the detections from each method, we setup baselines using the tracking approach described in . We provide detection and tracking results for each of these methods on the train, val and test splits to facilitate more systematic research. See the Supplementary Material for the results of the 2019 nuScenes tracking challenge.

2 Analysis

Here we analyze the properties of the methods presented in Section 4.1, as well as the dataset and matching function.

One of the contributions of nuScenes is the dataset size, and in particular the increase compared to KITTI (Table 1). Here we examine the benefits of the larger dataset size. We train PointPillars , OFT and an additional image baseline, SSD+3D, with varying amounts of training data. SSD+3D has the same 3D parametrization as MonoDIS , but use a single stage design . For this ablation study we train PointPillars with 6x fewer epochs and a one cycle optimizer schedule to cut down the training time. Our main finding is that the method ordering changes with the amount of data (Figure 6). In particular, PointPillars performs similar to SSD+3D at data volumes commensurate with KITTI, but as more data is used, it is clear that PointPillars is stronger. This suggests that the full potential of complex algorithms can only be verified with a bigger and more diverse training set. A similar conclusion was reached by with suggesting that the KITTI leaderboard reflects the data aug. method rather than the actual algorithms.

The importance of the matching function.

We compare performance of published methods (Table 4) when using our proposed 2m center-distance matching versus the IOU matching used in KITTI. As expected, when using IOU matching, small objects like pedestrians and bicycles fail to achieve above 0 AP, making ordering impossible (Figure 7). In contrast, center distance matching declares MonoDIS a clear winner. The impact is smaller for the car class, but also in this case it is hard to resolve the difference between MonoDIS and OFT.

The matching function also changes the balance between lidar and image based methods. In fact, the ordering switches when using center distance matching to favour MonoDIS over both lidar based methods on the bicycle class (Figure 7). This makes sense since the thin structures of bicycles make them difficult to detect in lidar. We conclude that center distance matching is more appropriate to rank image based methods alongside lidar based methods.

Multiple lidar sweeps improve performance.

According to our evaluation protocol (Section 3.1), one is only allowed to use 0.5s0.5s of previous data to make a detection decision. This corresponds to 10 previous lidar sweeps since the lidar is sampled at 20Hz20\text{Hz}. We device a simple way of incorporating multiple pointclouds into the PointPillars baseline and investigate the performance impact. Accumulation is implemented by moving all pointclouds to the coordinate system of the keyframe and appending a scalar time-stamp to each point indicating the time delta in seconds from the keyframe. The encoder includes the time delta as an extra decoration for the lidar points. Aside from the advantage of richer pointclouds, this also provides temporal information, which helps the network in localization and enables velocity prediction. We experiment with using 11, 55, and 1010 lidar sweeps. The results show that both detection and velocity estimates improve with an increasing number of lidar sweeps but with diminishing rate of return (Table 3).

Which sensor is most important?

An important question for AVs is which sensors are required to achieve the best detection performance. Here we compare the performance of leading lidar and image detectors. We focus on these modalities as there are no competitive radar-only methods in the literature and our preliminary study with PointPillars on radar data did not achieve promising results. We compare PointPillars, which is a fast and light lidar detector with MonoDIS, a top image detector (Table 4). The two methods achieve similar mAP (30.5% vs. 30.4%), but PointPillars has higher NDS (45.3% vs. 38.4%). The close mAP is, of itself, notable and speaks to the recent advantage in 3D estimation from monocular vision. However, as discussed above the differences would be larger with an IOU based matching function.

Class specifc performance is in Table 7-SM. PointPillars was stronger for the two most common classes: cars (68.4%68.4\% vs. 47.8%47.8\% AP), and pedestrians (59.7%59.7\% vs. 37.0%37.0\% AP). MonoDIS, on the other hand, was stronger for the smaller classes bicycles (24.5%24.5\% vs. 1.1%1.1\% AP) and cones (48.7%48.7\% vs. 30.8%30.8\% AP). This is expected since 1) bicycles are thin objects with typically few lidar returns and 2) traffic cones are easy to detect in images, but small and easily overlooked in a lidar pointcloud. 3) MonoDIS applied importance sampling during training to boost rare classes. With similar detection performance, why was NDS lower for MonoDIS? The main reasons are the average translation errors (5252cm vs. 7474cm) and velocity errors (1.55m/s1.55m/s vs. 0.32m/s0.32m/s), both as expected. MonoDIS also had larger scale errors with mean IOU 74%74\% vs. 71%71\% but the difference is small, suggesting the strong ability for image-only methods to infer size from appearance.

The importance of pre-training.

Using the lidar baseline we examine the importance of pre-training when training a detector on nuScenes. No pretraining means weights are initialized randomly using a uniform distribution as in . ImageNet pretraining uses a backbone that was first trained to accurately classify images. KITTI pretraining uses a backbone that was trained on the lidar pointclouds to predict 3D boxes. Interestingly, while the KITTI pretrained network did converge faster, the final performance of the network only marginally varied between different pretrainings (Table 3). One explanation may be that while KITTI is close in domain, the size is not large enough.

Better detection gives better tracking.

Weng and Kitani presented a simple baseline that achieved state-of-the-art 3d tracking results using powerful detections on KITTI. Here we analyze whether better detections also imply better tracking performance on nuScenes, using the image and lidar baselines presented in Section 4.1. Megvii, PointPillars and MonoDIS achieve an sAMOTA of 17.9%17.9\%, 3.5%3.5\% and 4.5%4.5\%, and an AMOTP of 1.50m1.50m, 1.69m1.69m and 1.79m1.79m on the val set. Compared to the mAP and NDS detection results in Table 4, the ranking is similar. While the performance is correlated across most metrics, we notice that MonoDIS has the shortest LGD and highest number of track fragmentations. This may indicate that despite the lower performance, image based methods are less likely to miss an object for a protracted period of time.

Conclusion

In this paper we present the nuScenes dataset, detection and tracking tasks, metrics, baselines and results. This is the first dataset collected from an AV approved for testing on public roads and that contains the full sensor suite (lidar, images, and radar). nuScenes has the largest collection of 3D box annotations of any previously released dataset. To spur research on 3D object detection for AVs, we introduce a new detection metric that balances all aspects of detection performance. We demonstrate novel adaptations of leading lidar and image object detectors and trackers on nuScenes. Future work will add image-level and point-level semantic labels and a benchmark for trajectory prediction .

The nuScenes dataset was annotated by Scale.ai and we thank Alexandr Wang and Dave Morse for their support. We thank Sun Li, Serene Chen and Karen Ngo at nuTonomy for data inspection and quality control, Bassam Helou and Thomas Roddick for OFT baseline results, Sergi Widjaja and Kiwoo Shin for the tutorials, and Deshraj Yadav and Rishabh Jain from EvalAI for setting up the nuScenes challenges.

References

Appendix A The nuScenes dataset

In this section we provide more details on the nuScenes dataset, the sensor calibration, privacy protection approach, data format, class mapping and annotation statistics.

To achieve a high quality multi-sensor dataset, careful calibration of sensor intrinsic and extrinsic parameters is required. These calibration parameters are updated around twice per week over the data collection period of 6 months. Here we describe how we perform sensor calibration for our data collection platform to achieve a high-quality multimodal dataset. Specifically, we carefully calibrate the extrinsics and intrinsics of every sensor. We express extrinsic coordinates of each sensor to be relative to the ego frame, i.e. the midpoint of the rear vehicle axle. The most relevant steps are described below:

Lidar extrinsics: We use a laser liner to accurately measure the relative location of the lidar to the ego frame.

Camera extrinsics: We place a cube-shaped calibration target in front of the camera and lidar sensors. The calibration target consists of three orthogonal planes with known patterns. After detecting the patterns we compute the transformation matrix from camera to lidar by aligning the planes of the calibration target. Given the lidar to ego frame transformation computed above, we compute the camera to ego frame transformation.

Radar extrinsics: We mount the radar in a horizontal position. Then we collect radar measurements by driving on public roads. After filtering radar returns for moving objects, we calibrate the yaw angle using a brute force approach to minimize the compensated range rates for static objects.

Camera intrinsic calibration: We use a calibration target board with a known set of patterns to infer the intrinsic and distortion parameters of the camera.

Privacy protection.

It is our priority to protect the privacy of third parties. As manual labeling of faces and license plates is prohibitively expensive for 1.4M images, we use state-of-the-art object detection techniques. Specifically for plate detection, we use Faster R-CNN with ResNet-101 backbone trained on Cityscapes https://github.com/bourdakos1/Custom-Object-Detection. For face detection, we use https://github.com/TropComplique/mtcnn-pytorch. We set the classification threshold to achieve an extremely high recall (similar to ). To increase the precision, we remove predictions that do not overlap with the reprojections of the known pedestrian and vehicle boxes in the image. Eventually we use the predicted boxes to blur faces and license plates in the images.

Data format.

Contrary to most existing datasets , we store the annotations and metadata (e.g. localization, timestamps, calibration data) in a relational database which avoids redundancy and allows for efficient access. The nuScenes devkit, taxonomy and annotation instructions are available onlinehttps://github.com/nutonomy/nuscenes-devkit.

Class mapping.

The nuScenes dataset comes with annotations for 23 classes. Since some of these only have a handful of annotations, we merge similar classes and remove classes that have less than 10000 annotations. This results in 10 classes for our detection task. Out of these, we omit 3 classes that are mostly static for the tracking task. Table 5-SM shows the detection classes and tracking classes and their counterpart in the general nuScenes dataset.

Annotation statistics.

We present more statistics on the annotations of nuScenes. Absolute velocities are shown in Figure 11-SM. The average speed for moving car, pedestrian and bicycle categories are 6.66.6, 1.31.3 and 44 m/s. Note that our data was gathered from urban areas which shows reasonable velocity range for these three categories.

We analyze the distribution of box annotations around the ego-vehicle for car, pedestrian and bicycle categories through a polar range density map as shown in Figure 12-SM. Here, the occurrence bins are log-scaled. Generally, the annotations are well-distributed surrounding the ego-vehicle. The annotations are also denser when they are nearer to the ego-vehicle. However, the pedestrian and bicycle have less annotations above the 100m range. It can also be seen that the car category is denser in the front and back of the ego-vehicle, since most vehicles are following the same lane as the ego-vehicle.

In Section 2 we discussed the number of lidar points inside a box for all categories through a hexbin density plot, but here we present the number of lidar points of each category as shown in Figure 13-SM. Similarly, the occurrence bins are log-scaled. As can be seen, there are more lidar points found inside the box annotations for car at varying distances from the ego-vehicle as compared to pedestrian and bicycle. This is expected as cars have larger and more reflective surface area than the other two categories, hence more lidar points are reflected back to the sensor.

Scene reconstruction.

nuScenes uses an accurate lidar based localization algorithm (Section 2). It is however difficult to quantify the localization quality, as we do not have ground truth localization data and generally cannot perform loop closure in our scenes. To analyze our localization qualitatively, we compute the merged pointcloud of an entire scene by registering approximately 800 pointclouds in global coordinates. We remove points corresponding to the ego vehicle and assign to each point the mean color value of the closest camera pixel that the point is reprojected to. The result of the scene reconstruction can be seen in Figure 15, which demonstrates accurate synchronization and localization.

Appendix B Implementation details

Here we provide additional details on training the lidar and image based 3D object detection baselines.

For all experiments, our PointPillars networks were trained using a pillar xy resolution of 0.25 meters and an x and y range of $meters.Themaxnumberofpillarsandbatchsizewasvariedwiththenumberoflidarsweeps.For1,5,and10sweeps,wesetthemaximumnumberofpillarsto10000,22000,and30000respectivelyandthebatchsizeto64,64,and48.Allexperimentsweretrainedfor750epochs.Theinitiallearningratewassettometers. The max number of pillars and batch size was varied with the number of lidar sweeps. For 1, 5, and 10 sweeps, we set the maximum number of pillars to 10000, 22000, and 30000 respectively and the batch size to 64, 64, and 48. All experiments were trained for 750 epochs. The initial learning rate was set to10^{-3}andwasreducedbyafactorofand was reduced by a factor of10$ at epoch 600 and again at 700. Only ground truth annotations with one or more lidar points in the accumulated pointcloud were used as positive training examples. Since bikes inside of bike racks are not annotated individually and the evaluation metrics ignore bike racks, all lidar points inside bike racks were filtered out during training.

OFT implementation details.

For each camera, the Orthographic Feature Transform (OFT) baseline was trained on a voxel grid in each camera’s frame with an lateral range of $meters,alongitudinalrangeofmeters, a longitudinal range of[0.1,50.1]metersandaverticalrangeofmeters and a vertical range of(-3,1)meters.Wetrainedonlyonannotationsthatwerewithin50metersofthecar’segoframecoordinatesystem’sorigin.Usingthe‘visibility’attributeinthenuScenesdataset,wealsofilteredoutannotationsthathadvisibilitylessthanmeters. We trained only on annotations that were within 50 meters of the car’s ego frame coordinate system’s origin. Using the ‘visibility’ attribute in the nuScenes dataset, we also filtered out annotations that had visibility less than40\%.Thenetworkwastrainedfor60epochsusingalearningrateof. The network was trained for 60 epochs using a learning rate of2\times 10^{-3}$ and used random initialization for the network weights (no ImageNet pretraining).

Appendix C Experiments

In this section we present more detailed result analysis on nuScenes. We look at the performance on rain and night data, per-class performance and semantic map filtering. We also analyze the results of the tracking challenge.

As described in Section 2, nuScenes contains data from 2 countries, as well as rain and night data. The dataset splits (train, val, test) follow the same data distribution with respect to these criteria. In Table 6 we analyze the performance of three object detection baselines on the relevant subset of the val set. We can see a small performance drop for Singapore as compared to the overall val set (USA and Singapore), particularly for vision based methods. This is likely due to different object appearance in the different countries, as well as different label distributions. For rain data we see only a small decrease in performance on average, with worse performance for OFT and PP, and slightly better performance for MDIS. One reason is that the nuScenes dataset annotates any scene with raindrops on the windshield as rainy, regardless of whether there is ongoing rainfall. Finally, night data shows a drastic performance relative drop of 36% for the lidar based method and 55% and 58% for the vision based methods. This may indicate that vision based methods are more affected by worse lighting. We also note that night scenes have very few objects and it is harder to annotate objects with bad visibility. For annotating data, it is essential to use camera and lidar data, as described in Section 2.

Per-class analysis.

The per class performance of PointPillars is shown in Table 7-SM (top) and Figure 17-SM. The network performed best overall on cars and pedestrians which are the two most common categories. The worst performing categories were bicycles and construction vehicles, two of the rarest categories that also present additional challenges. Construction vehicles pose a unique challenge due to their high variation in size and shape. While the translational error is similar for cars and pedestrians, the orientation error for pedestrians () is higher than that of cars (). This smaller orientation error for cars is expected since cars have a greater distinction between their front and side profile relative to pedestrians. The vehicle velocity estimates are promising (e.g. 0.240.24 m/s AVE for the car class) considering the typical speed of a vehicle in the city would be 1010 to 1515 m/s.

Semantic map filtering.

In Section 4.2 and Table 7-SM we show that the PointPillars baseline achieves only an AP of 1%1\% on the bicycle class. However, when filtering both the predictions and ground truth to only include boxes on the semantic map priorDefined here as the union of roads and sidewalks., the AP increases to 30%30\%. This observation can be seen in Figure 16-SM, where we plot the AP at different distances of the ground truth to the semantic map prior. As seen, the AP drops when the matched GT is farther from the semantic map prior. Again, this is likely because bicycles away from the semantic map tend to be parked and occluded with low visibility.

Tracking challenge results.

In Table 8 we present the results of the 2019 nuScenes tracking challenge. Stan use the Mahalanobis distance for matching, significantly outperforming the strongest baseline (+40%+40\% sAMOTA) and setting a new state-of-the-art on the nuScenes tracking benchmark. As expected, the two methods using only monocular camera images perform poorly (CeVi and MDIS). Similar to Section 4, we observe that the metrics are highly correlated, with notable exceptions for MDIS LGD and CeOp AMOTP. Note that all methods use a tracking-by-detection approach. With the exception of CeOp and CeVi, all methods use a Kalman filter .