Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting

Benjamin Wilson, William Qi, Tanmay Agarwal, John Lambert, Jagjeet Singh, Siddhesh Khandelwal, Bowen Pan, Ratnesh Kumar, Andrew Hartnett, Jhony Kaesemodel Pontes, Deva Ramanan, Peter Carr, James Hays

Introduction

In order to achieve the goal of safe, reliable autonomous driving, a litany of machine learning tasks must be addressed, from stereo depth estimation to motion forecasting to 3D object detection. In recent years, numerous high quality self-driving datasets have been released to support research into these and other important machine learning tasks. Many datasets are annotated “sensor” datasets in the spirit of the influential KITTI dataset . The Argoverse 3D Tracking dataset was the first such dataset with “HD maps” — maps containing lane-level geometry. Also influential are self-driving “motion prediction” datasets — containing abstracted object tracks instead of raw sensor data — of which the Argoverse Motion Forecasting dataset was the first.

In the last two years, the Argoverse team has hosted six competitions on 3D tracking, stereo depth estimation, and motion forecasting. We maintain evaluation servers and leaderboards for these tasks, as well as 3D detection. The leaderboards collectively contain thousands of submissions from four hundred teamsThis count includes private submissions not posted to the public leaderboards.. We also maintain the Argoverse API and have addressed more than one hundred issueshttps://github.com/argoverse/argoverse-api. From these experiences we have formed the following guiding principles to guide the creation of the next iteration of Argoverse datasets.

Bigger isn’t always better. Self-driving vehicles capture a flood of sensor data which is logistically difficult to work with. Sensor datasets are several terabytes in size, even when compressed. If standard benchmarks grow further, we risk alienating much of the academic community and leaving progress to well-resourced industry groups. For this reason, we match but do not exceed the scale of sensor data in nuScenes and Waymo Open .

Make every instance count. Much of driving is boring. Datasets should focus on the difficult, interesting scenarios where current forecasting and perception systems struggle. Therefore we mine for especially crowded, dynamic, and kinematically unusual scenarios.

Diversity matters. Training on data from wintertime Detroit is not sufficient for detecting objects in Miami — Miami has 15 times the frequency of motorcycles and mopeds. Behaviors differ as well, so learned pedestrian motion behavior might not generalize. Accordingly, each of our datasets are drawn from six diverse cities — Austin, Detroit, Miami, Palo Alto, Pittsburgh, and Washington D.C. — and different seasons, as well, from snowy to sunny.

Map the world. HD maps are powerful priors for perception and forecasting. Learning-based methods that found clever ways to encode map information performed well in Argoverse competitions. For this reason, we augment our HD map representation with 3D lane geometry, paint markings, crosswalks, higher resolution ground height, and more.

Self-supervise. Other machine learning domains have seen enormous success from self-supervised learning in recent years. Large-scale lidar data from dynamic scenes, paired with HD maps, could lead to better representations than current supervised approaches. For this reason, we build the largest dataset of lidar sensor data.

Fight the heavy tail. Passenger vehicles are common, and thus we can assess our forecasting and detection accuracy for cars. However, with existing datasets, we cannot assess forecasting accuracy for buses and motorcycles with their distinct behaviors, nor can we evaluate stroller and wheel chair detection. Thus we introduce the largest taxonomy to date for sensor and forecasting datasets, and we ensure enough samples of rare objects to train and evaluate models.

With these guidelines in mind we built the three Argoverse 2 (AV2) datasets. Below, we highlight some of their contributions.

The 1,000 scenario Sensor dataset has the largest self-driving taxonomy to date – 30 categories. 26 categories contain at least 6,000 cuboids to enable diverse taxonomy training and testing. The dataset also has stereo imagery, unlike recent self-driving datasets.

The 20,000 scenario Lidar dataset is the largest dataset for self-supervised learning on lidar. The only similar dataset, concurrently developed ONCE , does not have HD maps.

The 250,000 scenario Motion Forecasting Dataset has the largest taxonomy – 5 types of dynamic actors and 5 types of static actors – and covers the largest mapped area of any such dataset.

We believe these datasets will support research into problems such as 3D detection, 3D tracking, monocular and stereo depth estimation, motion forecasting, visual odometry, pose estimation, lane detection, map automation, self-supervised learning, structure from motion, scene flow, optical flow, time to contact estimation, and point cloud forecasting.

Related Work

The last few years have seen rapid progress in self-driving perception and forecasting research, catalyzed by many high quality datasets.

Sensor datasets and 3D Object Detection and Tracking. New sensor datasets for 3D object detection have led to influential detection methods such as anchor-based approaches like PointPillars , and more recent anchor-free approaches such as AFDet and CenterPoint . These methods have led to dramatic accuracy improvements on all datasets. In turn, these improvements have made isolation of object-specific point clouds possible, which has proven invaluable for offboard detection and tracking , and for simulation , which previously required human-annotated 3D bounding boxes . New approaches explore alternate point cloud representations, such as range images . Streaming perception introduces a paradigm to explore the tradeoff between accuracy and latency. A detailed comparison between the AV2 Sensor Dataset and recent 3D object detection datasets is provided in Table 1.

Motion Forecasting. For motion forecasting, the progress has been just as significant. A transition to attention-based methods has led to a variety of new vector-based representations for map and trajectory data . New datasets have also paved the way for new algorithms, with nuScenes , Lyft L5 , and the Waymo Open Motion Dataset all releasing lane graphs after they proved to be essential in Argoverse 1 . Lyft also introduced traffic/speed control data, while Waymo added crosswalk polygons, lane boundaries (with marking type), speed limits, and stop signs to the map. More recently, Yandex has released the Shifts dataset, which is the largest (by scenario hours) collection of forecasting data available to date. Together, these datasets have enabled exploration of multi-actor, long-range motion forecasting leveraging both static and dynamic maps.

Following upon the success of Argoverse 1.1, we position AV2 as a large-scale repository of high-quality motion forecasting scenarios - with guarantees on data frequency (exactly 10 Hz) and diversity (>2000 km of unique roadways covered across 6 cities). This is in contrast to nuScenes (reports data at just 2 Hz) and Lyft (collected on a single 10 km segment of road), but is complementary to Waymo Open Motion Dataset (employs a similar approach for scenario mining and data configuration). Complementary datasets are essential for these safety critical problems as they provide opportunities to evaluate generalization and explore transfer learning. To improve ease of use, we have also designed AV2 to be widely accessible both in terms of data size and format — a detailed comparison vs. other recent forecasting datasets is provided in Table 2.

Broader Problems of Perception for Self-Driving. Aside from the tasks of object detection and motion forecasting, new, large-scale sensor datasets for self-driving present opportunities to explore dozens of new problems for perception, especially those that can be potentially solved via self-supervision. A number of new problems have been recently proposed; real-time 3D semantic segmentation in video has received attention thanks to SemanticKITTI . HD map automation and HD map change detection have received additional attention, along with 3D scene flow and pixel-level scene simulation . Datasets exist with unique modalities such as thermal imagery . Our new Lidar Dataset enables large-scale self-supervised training of new approaches for freespace forecasting or point cloud forecasting .

The Argoverse 2 Datasets

The most similar sensor dataset to ours is the highly influential nuScenes – both datasets have 1,000 scenarios and HD maps, although Argoverse is unique in having ground height maps. nuScenes contains radar data while AV2 contains stereo imagery. nuScenes has a large taxonomy – twenty-three object categories of which ten have suitable data for training and evaluation. Our dataset contains thirty object categories of which twenty-six are well sampled enough for training and evaluation. nuScenes spans two cities, while our proposed dataset spans six.

Privacy. All faces and license plates, whether inside vehicles or outside of the driveable area, are blurred extensively to preserve privacy.

Sensor Dataset splits. We randomly partition the dataset with train, validation, and test splits of 700, 150, and 150 scenarios, respectively.

2 Lidar Dataset

Lidar Dataset splits. We randomly partition the dataset with train, validation, and test splits of 16,000, 2,000, and 2,000 scenarios, respectively.

3 Motion Forecasting Dataset

Motion forecasting addresses the problem of predicting future states (or occupancy maps) for dynamic actors within a local environment. Some examples of relevant actors for autonomous driving include: vehicles (both parked and moving), pedestrians, cyclists, scooters, and pets. Predicted futures generated by a forecasting system are consumed as the primary inputs in motion planning, which conditions trajectory selection on such forecasts. Generating these forecasts presents a complex, multi-modal problem involving many diverse, partially-observed, and socially interacting agents. However, by taking advantage of the ability to “self-label” data using observed ground truth futures, motion forecasting becomes an ideal domain for application of machine learning.

Building upon the success of Argoverse 1, the Argoverse 2 Motion Forecasting dataset provides an updated set of prediction scenarios collected from a self-driving fleet. The design decisions enumerated below capture the collective lessons learned from both our internal research/development, as well as feedback from more than 2,700 submissions by nearly 260 unique teamsThis count includes private submissions not posted to the public leaderboards. across 3 competitions :

Motion forecasting is a safety critical system in a long-tailed domain. Consequently, our dataset is biased towards diverse and interesting scenarios containing different types of focal agents (see section 3.3.2). Our goal is to encourage the development of methods that ensure safety during tail events, rather than to optimize the expected performance on “easy miles”.

There is a “Goldilocks zone” of task difficulty. Performance on the Argoverse 1 test set has begun to plateau, as shown in Figure 10 of the appendix. Argoverse 2 is designed to increase prediction difficulty incrementally, spurring productive focused research for the next few years. These changes are intended to incentivize methods that perform well on extended forecast horizons (3 s →\rightarrow 6 s), handle multiple types of dynamic objects (1 →\rightarrow 5), and ensure safety in scenarios from the long tail. Future Argoverse releases could continue to increase the problem difficulty by reducing observation windows and increasing forecasting horizons.

Within each scenario, we mark a single track as the “focal agent”. Focal tracks are guaranteed to be fully observed throughout the duration of the scenario and have been specifically selected to maximize interesting interactions with map features and other nearby actors (see Section 3.3.2). To evaluate multi-agent forecasting, we also mark a subset of tracks as “scored actors” (as shown in Figure 5), with guarantees for scenario relevance and minimum data quality.

3.2 Mining Interesting Scenarios

4 HD Maps

Each scenario in the three datasets described above shares the same HD map representation. Each scenario carries its own local map region, similar to the Waymo Open Motion dataset. This is a departure from the original Argoverse datasets in which all scenarios were localized onto two city-scale maps—one for Pittsburgh and one for Miami. In the Appendix, we provide examples. Advantages of per-scenario maps include more efficient queries and their ability to handle map changes. A particular intersection might be observed multiple times in our datasets, and there could be changes to the lanes, crosswalks, or even ground height in that time.

Experiments

Argoverse 2 supports a variety of downstream tasks. In this section we highlight three different learning problems: 3D object detection, point cloud forecasting, and motion forecasting — each supported by the sensor, lidar, and motion forecasting datasets, respectively. First, we illustrate the challenging and diverse taxonomy within the Argoverse 2 sensor dataset by training a state-of-the-art 3D detection model on our twenty-six evaluation classes including “long-tail” classes such as stroller, wheel chairs, and dogs. Second, we showcase the utility of the Argoverse 2 lidar dataset through large-scale, self-supervised learning through the point cloud forecasting task. Lastly, we demonstrate motion forecasting experiments which provide the first baseline for broad taxonomy motion prediction.

In Table 3, we provide a snapshot of submissions to the Argoverse 2 3D Object Detection Leaderboard.

2 Point Cloud Forecasting

We use three metrics to evaluate the performance of our forecasting model: mean IoU, l1l_{1}-norm, and Chamfer distance. The mean IoU evaluates the predicted range mask. The l1l_{1}-norm measures the average l1l_{1} distance between the pixel sets of predicted range image and the ground-truth image, which are both masked out by the ground-truth range mask. The Chamfer distance is obtained by adding up the Chamfer distances in both directions (forward and backward) between the ground-truth point cloud and the predicted scene point cloud which is obtained by back-projecting the predicted range image.

3 Motion Forecasting

We present several forecasting baselines which try to make use of different aspects of the data. Those which are trained using the focal agent only and do not capture any social interaction include: constant velocity, nearest neighbor, and LSTM encoder-decoder models (both with and without a map-prior). We also evaluate WIMP as an example of a graph-based attention method that captures social interaction. All hyper-parameters are obtained from the reference implementations.

Baseline approaches are evaluated according to standard metrics. Following , we use minADE and minFDE as the metrics; they evaluate the average and endpoint L2 distance respectively, between the best forecasted trajectory and the ground truth. We also use Miss Rate (MR) which represents the proportion of test samples where none of the forecasted trajectories were within 2.0 meters of ground truth according to endpoint error. The resulting performance illustrates both the community’s progress on the problem and the significant increase in dataset difficulty when compared with Argoverse 1.1.

Baseline Results. Table 5 summarizes the results of baselines. For K=1, Argoverse 1 showed that a constant velocity model (minFDE=7.89) performed better than NN+map(prior) (minFDE=8.12), which is not the case here. This further proves that Argoverse 2 is kinematically more diverse and cannot be solved by making constant velocity assumptions. Surprisingly, NN and LSTM variants that make use of a map prior perform worse than those which do not, illustrating the scope of improvement in how these baselines leverage the map. For K=6, WIMP significantly outperforms every other baseline. This emphasizes that it is imperative to train expressive models that can leverage map prior and social context along with making diverse predictions. The trends are similar to our past 3 Argoverse Motion Forecasting competitions : Graph-based attention methods (e.g. ) continued to dominate the competition, and were nearly twice as accurate as the next best baseline (Nearest Neighbor) at K=6. That said, some of the rasterization-based (e.g. ) methods also showed promising results. Finally, we also evaluated baseline methods in the context of transfer learning and varied object types, the results of which are summarized in the Appendix.

In Table 6, we provide a snapshot of submissions to the Argoverse 2 Motion Forecasting Leaderboard.

Conclusion

Discussion. In this work, we have introduced three new datasets that constitute Argoverse 2. We provide baseline explorations for three tasks – 3d object detection, point cloud forecasting and motion forecasting. Our datasets provide new opportunities for many other tasks. We believe our datasets compare favorably to existing datasets, with HD maps, rich taxonomies, geographic diversity, and interesting scenes.

Limitations. As in any human annotated dataset, there is label noise, although we seek to minimize it before release. 3D bounding boxes of objects are not included in the motion forecasting dataset, but one can make reasonable assumptions about the object extent given the object type. The motion forecasting dataset also has imperfect tracking, consistent with state-of-the-art 3D trackers.

References

Appendix

In Figure 8, we provide a diagram of the sensor suite used to capture the Argoverse 2 datasets. Figure 9 shows the speed distribution for annotated pedestrian 3D cuboids and the yaw distribution.

2 Additional Information About Motion Forecasting Dataset

Kinematic scoring selects for trajectories performing sharp turns or significant (de)accelerations. The map complexity program biases the data set towards trajectories complex traversals of the underlying lane graph. In particular, complex map regions, paths through intersections, and lane-changes score highly. Social scoring rewards tracks through dense regions of other actors. Social scoring also selects for non-vehicle object classes to ensure adequate samples from rare classes, such as motorcycles, for training and evaluation. Finally, the autonomous vehicle scoring program encourages the selection of tracks that intersect the ego-vehicle’s desired route.

3 Additional Information About HD Maps

In Figure 12, we display examples of local HD maps associated with individual logs/scenarios.

4 Additional 3D Detection Results

In Figure 13, we show additional evaluation metrics for our detection baseline.

True Positive Metrics Average Translation Error (ATE)

where X={mATEunit,mASEunit,mAOEunit}\mathcal{X}=\{mATE_{unit},mASE_{unit},mAOE_{unit}\}

5 Training Details of SPF2 baseline

We sample 2-second training snippets (representing 1 second of past and 1 second of future data) every 0.5 seconds. Thus, for a training log with 30 second duration, 59 training snippets would be sampled. We train the model for 16 epochs by using the Adam optimizer with the learning rate of 4e−34e-3, betas of 0.9 and 0.999, and batch size of 16 per GPU.

6 Additional Motion Forecasting Experiments

The results of transfer learning experiments are summarized in Table 8. WIMP was trained and tested in different settings with Argoverse 1.1 and Argoverse 2. As expected, the model works best when it is trained and tested on the same distribution (i.e. both train and test data come from Argoverse 1.1, or both from Argoverse 2). For example, when WIMP is tested on Argoverse 2 (6s), the model trained on Argoverse 2 (6s) has a minFDE of 2.91, whereas the one trained on Argoverse 1.1 (3s) has a minFDE of 6.82 (i.e. approximately 2.3x worse). Likewise, in the reverse setting, when WIMP is tested on Argoverse 1.1 (3s), the model trained on Argoverse 1.1 (3s) has a minFDE of 1.14 and the one trained on Argoverse 2 (6s) has minFDE of 2.05 (i.e. approximately 1.8x worse). This indicates that transfer learning from Argoverse 2 (Beta) to Argoverse 1.1 is more useful than the reverse setting, despite being smaller in the number of scenarios. However, the publicly released version of Argoverse 2 Motion Forecasting (the non-beta 2.0 version) has comparable size with Argoverse 1.1.

We note that it is a common practice to train and test sequential models on varied sequence length (e.g. machine translation). As such, it is still reasonable to expect a model trained with 3s to do well on 6s horizon. Several factors may contribute to distribution shift, including differing prediction horizon, cities, mining protocols, object types. Notably, however, these results indicate that Argoverse 2 is significantly more challenging and diverse than its predecessor.

6.2 Experiment with different object types

Table 9 shows the results on Nearest Neighbor baseline (without map prior) on different object types. As one would expect, the displacement errors in pedestrians are significantly lower than other object types. This occurs because they move at significantly slower velocities. However, this does not imply that pedestrian motion forecasting is a solved problem and one should rather focus on other object types. This instead means that we need to come up with better metrics that can capture that fact lower displacement errors in pedestrians can often be more critical than higher errors in vehicles. We leave this line of work for future scope.

Datasheet for Argoverse 2

For what purpose was the dataset created? Was there a specific task in mind? Was there a specific gap that needed to be filled? Please provide a description. Argoverse was created to support the global research community in improving the state of the art in machine learning tasks vital for self driving. The Argoverse 2 datasets described in this manuscript improve upon the initial Argoverse datasets. These datasets support many tasks, from 3D perception to motion forecasting to HD map automation.

The three datasets proposed in this manuscript address different gaps in this space. See the comparison charts in the main manuscript for a more detailed breakdown.

The Argoverse 2 Sensor Dataset has a richer taxonomy than similar datasets. It is the only dataset of similar size to have stereo imagery. The 1,000 logs in the dataset were chosen to have a variety of object types with diverse interactions.

The Argoverse 2 Motion Forecasting Dataset also has a richer taxonomy than existing datasets. The scenarios in the dataset were mined with an emphasis on unusual behaviors that are difficult to predict.

The Argoverse 2 Lidar Dataset is the largest Lidar Dataset. Only the concurrent ONCE dataset is similarly sized to enable self-supervised learning in lidar space. Unlike ONCE, our dataset contains HD maps and high frame rate lidar.

Who created this dataset (e.g., which team, research group) and on behalf of which entity (e.g., company, institution, organization)? The Argoverse 2 datasets were created by researchers at Argo AI.

What support was needed to make this dataset? (e.g.who funded the creation of the dataset? If there is an associated grant, provide the name of the grantor and the grant name and number, or if it was supported by a company or government agency, give those details.) The creation of this dataset was funded by Argo AI.

COLLECTION

PREPROCESSING / CLEANING / LABELING

USES

DISTRIBUTION

MAINTENANCE