INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps

Wei Zhan, Liting Sun, Di Wang, Haojie Shi, Aubrey Clausse, Maximilian Naumann, Julius Kummerle, Hendrik Konigshof, Christoph Stiller, Arnaud de La Fortelle, Masayoshi Tomizuka

I Introduction

In order to enable fully autonomous driving in complex scenarios, comprehensive understanding and accurate prediction of the behavior and motion of other road users are required. Moreover, autonomous vehicles need to behave like vehicles with human drivers to make themselves more predictable to others and thus, facilitate cooperation. These are two of the major challenges in the field of autonomous driving. To overcome these challenges, considerable amount of research efforts have been devoted to: i) predicting the future intention and motion of other road users , ii) modeling and analyzing driving behavior , iii) clustering the motion and finding representation of the motion primitives , iv) cloning and imitating human and expert behavior , and v) generating human-like and social behavior and motion .

All the aforementioned research areas require interactive vehicle motion data from real-world driving scenarios, which is the most fundamental and indispensable asset. NGSIM dataset is the most popular one used in the aforementioned areas, such as prediction , behavior modeling , social behavior generation and planning , and representation learning , since it is publicly available with decent scale and quality. The recently released highD dataset also greatly assists behavior-related research such as prediction . Public motion datasets such as NGSIM and highD facilitated, but also restricted behavior-related research due to limited diversity, complexity and criticality of the scenarios and behavior. Also, the importance of map information and completeness of interaction entities were under-addressed in most of the existing datasets. However, these missing points are crucial for behavior-related research, which will be discussed in the following.

1) Diversity of interactive driving scenarios: Recent behavior-related research using public datasets was mostly restricted to highway scenarios due to the data availability. There are many more highly interactive driving scenarios to explore, such as roundabouts with yield/stop signs, (unsignalized) one/two/all-way stop intersections (shown in Fig. 1), signalized intersections with unprotected left turn, zipper merge in cities, etc.

2) International diving cultures: Most of the existing datasets only contain driving data in one specific country. However, driving cultures in different countries and different continents can be distinct for very similar scenarios. Without motion data in similar scenarios from different countries, it is not possible to incorporate the impact of driving cultures in different countries, such as driving styles, preferences, risk tolerance, understanding of traffic rules, etc., for behavior modeling and analysis as well as the design of adaptive prediction and planning algorithms in different countries.

3) Complexity of the scenarios and behavior: Most of the scenarios in the existing public datasets are relatively simple and structured with explicit right-of-way. The behavior of the drivers is only occasionally impacted by others. There is very little social pressure (such as several vehicles waiting behind and even honking) on the drivers, so that their behavior is cautious without aggressive and irrational decisions. A motion dataset with much more complex and interactive behavior and scenarios is expected to facilitate the research tackling real and challenging problems.

4) Criticality of the situations: Critical situations (such as near-collision cases) are much more challenging and valuable than others for behavior-related research areas. For instance, proposed a fatality-aware prediction benchmark emphasizing prediction inaccuracies in critical situations. However, critical situations are too sparse in existing motion datasets, and can hardly be identified. Therefore, a motion dataset with denser critical situations is necessary to facilitate the research efforts on those difficult problems.

5) Map information: Map information with references and semantics such as lanelet connections and traffic rules, are crucial for behavior-related research areas such as motion planning and prediction. It provides key information on input (features), such as route and goal point , distance to the merging point , lateral position within the lane , etc., and makes the algorithms generalizable to other scenarios. Such semantic maps are currently missing for most of the existing public motion datasets.

6) Completeness of interaction entities: In order to accurately model, predict and imitate the interactive vehicle behavior, it is crucial to provide motions of all surrounding entities which may impact their behavior in the dataset. This requirement was often overlooked when using motion data collected by onboard sensors due to occlusions and limited field of view of the sensors. Although existing motion datasets collected from onboard sensors contain data collected from a wide range of areas for long time periods, complete and meaningful interaction pairs are relatively sparse.

In this paper, we will emphasize all the aforementioned aspects to construct an international motion dataset collected by drones and traffic cameras.

Diverse and international: It contains a variety of highly interactive driving scenarios from different countries, such as roundabouts, signalized/unsignalized intersections, as well as highway/urban merging and lane change.

Complex and critical: Part of the scenarios are relatively unstructured with inexplicit right-of-way. The driving behavior in the dataset are highly impacted by other drivers, whose behavior can be aggressive or irrational due to the social pressure. Near-collision or slight-collision scenes are contained in the dataset to facilitate the research for critical situations.

Semantic map and complete information: HD maps with semantics are provided to generated key features in the context. Motions of all entities which may influence the driving behavior are included in the dataset.

The proposed dataset can significantly facilitate behavior-related research such as motion prediction, imitation learning, decision-making and planning, representation learning, interaction extraction and social behavior generation. Results from exemplar methods in all these areas are provided utilizing the proposed dataset.

II Related Work

As mentioned in Section I, NGSIM dataset is the most popular vehicle motion dataset among the behavior-related research communities. The raw data was collected by cameras mounted on buildings and processed automatically . The accuracy of the dataset is mostly acceptable. However, there may be steady errors, and the image projection can significantly enlarge the size of the vehicles. Researchers proposed methods to rectify the errors, but it can only improve the quality of a small part of the dataset. In view of the problems in NGSIM, highD dataset was constructed by using a drone with more accurate vehicle motions and larger amount of high way driving data than NGSIM. Other datasets from bird’s eye view are more focused on pedestrian behavior without strong vehicle interactions.

The driving scenarios presented in NGSIM and highD are quite limited. NGSIM contains highway driving (including ramp merging and double lane change) and signalized intersection scenarios. In fact, signalized intersections are mostly controlled by the traffic lights and interactions are very rare and slight. A small amount of lane changes are interactive, but most of them are neither interactive nor critical. Ramp merging and double lane change can be highly interactive when the traffic is relatively dense, but the amount of interaction is still relatively limited in NGSIM. HighD only contains highway driving scenarios with car following and lane change. Urban scenarios which contain densely and highly interactive behavior, such as roundabouts and unsignalized intersections are not included in either of the two public datasets of vehicle motions.

II-B Datasets from Onboard Sensors

In addition to the bird’s-eye-view motion datasets, two types of onboard-sensor-based ones are also publicly available. One includes motion data of surrounding entities from onboard LiDARs and front-view cameras, such as Argoverse and HDD dataset . The other only contains motions of many data-collection vehicles from onboard GPS, such as 100-car study .

There are two major advantages for datasets from onboard sensors. One is that a variety of driving scenarios with relatively long data recording time are usually included in those datasets, such as urban driving at signalized/unsignalized intersections and highway driving with ramp merging, etc. The other is that the occlusions of LiDARs and cameras are recorded so that the actual occlusions from perspective of the ego vehicle can be partially recovered.

Completeness of interaction entities is a major problem when using datasets from onboard sensors for behavior-related research. For motion datasets with GPS-based fleets, it is hard to determine whether the vehicles in an ”interactive” motion segment was actually interacting with each other since there is no motion recording of other surrounding vehicles (or even pedestrians) without GPS devices installed. For motion datasets constructed from onboard LiDARs and cameras, it is hard to guarantee that all the surrounding objects impacting the behavior of other vehicles are included in the dataset when predicting the motions of others. Therefore, complete interactions are relatively sparse in such kind of datasets. If the sensors cannot cover the full field of view, it will be even impossible to guarantee the completeness of information for the surrounding entities of the ego data collection vehicle.

Also, the data collected in a large area may lead to very few repetitions at the same location. It is hard to learn multi-modal driving behavior for prediction or planning since only one sequence of motions can be found with similar features at the same location.

Map information is also missing in most of the motion datasets. To the best of our knowledge, Argoverse is the only motion dataset providing relatively rich map information. Physical layer (locations of curbs, road markings, etc.) is contained and semantic information (lane bounds and turn directions, etc.) required by prediction and planning is partially included.

Table I provides a comparison of the three most useful public vehicle motion datasets as well as the one presented in this article. The proposed dataset contains much more diverse, complex and critical scenarios and vehicle motions comparing to the other three. In addition, HD maps with full semantic information are provided, and the completeness of interaction entities is superior to datasets from onboard sensors.

III Features of the Dataset

In this section, we will illustrate the features of the proposed dataset by highlighting the diversity, internationality, complexity, criticality, and semantic map.

Fig. 2 illustrates a variety of highly interactive driving scenarios from traffic cameras and drones in our dataset, including zipper merging in a city (Fig. 2 (a)), ramp merging and lane change on a highway (Fig. 2 (b)), five roundabouts with yield and stop signs (Fig. 2 (c) - (g)), several unsignalized intersections with one/two/all-way stops (Fig. 2 (h) - (j)), and unprotected left turn at a signalized intersection (Fig. 2 (k)). In Fig. 2, the first two letters of the names represent the sources of the data (drone as DR and traffic camera as TC), while next three letters represent the corresponding country and the last two represent the scenario code in the dataset. The numbers in circles denote the branch ID for each scenario.

Fig. 2 (b) contains several subscenarios. The subscenario with the upper two lanes (that merge into one finally) is a zipper merging which is similar to the urban counterpart in Fig. 2 (a), where vehicles strongly interact with each other. It is also a ramp for the middle two lanes. The subscenario with the lower three lanes (that merge into two finally) is a forced merging and vehicles have to change their lanes.

The roundabout in Fig. 2 (f) is an extremely busy 7-way roundabout with one “yield” branch and six “stop” branches. Lots of vehicles enter the roundabout at the same time with intensive interactions and relatively high speeds. The branches of the roundabouts in Fig. 2 (c)-(e) are controlled by yield signs, while all branches of the roundabout in Fig. 2 (g) are controlled by stop signs.

Figure 2 (i) shows an extremely busy all-way-stop intersection with 9 lanes controlled by stop signs. Multiple vehicles are interactively inching to compete. The scenario shown in Fig. 2 (j) contains three branches (Branch 1, 2, 5) controlled by stop signs, while vehicles from Branch 3 and 6 have the right-of-way (RoW). Lots of vehicles are entering the intersections from all branches (except Branch 4), and vehicles holding RoW on the straight road are with relatively high speed. A busy all-way-stop T-intersection is shown in Fig. 2 (h), while three other branches (Branch 4-6) are also controlled by stop signs.

III-B Internationality

The motion data was collected from three continents (North America, Asia and Europe). Motion data collected by drones are from four countries, namely, the US, China, Germany and Bulgaria, as indicated in the names of the scenarios (USA/CHN/DEU/BGR). Vehicles in all these countries are driven on the right-hand side of the road. However, driving culture in these countries is with remarkable distinctions.

We provide motion data from three roundabouts with similar traffic rules, namely, SR from the US, OF from Germany and LN from China. All the three roundabouts do not have stop signs, and the nominal traffic rule is that the vehicles entering the roundabout should yield the ones which is already in the roundabout.

We also provide motion data from two zipper merging scenarios, those are, MT from Germany and ZS from China (the upper two lanes in Fig. 2 (b)). Although MT is urban road and ZS is the entrance of highway, the “zipper” rule remains the same, and the speeds are similar when the traffic is heavy.

III-C Complexity

In addition to regular driving behavior such as car-following, lane change, stop and left/right/U-turn, our dataset emphasizes highly interactive and complex driving behavior with cooperative and adversarial motions of the vehicles. By carefully choosing the locations and corresponding rush hours for the data collection, we were able to gather large amounts of strong interactions within relative short period of time. Strongly interactive pairs of vehicles can even appear every few seconds from time to time for scenarios such as the ramp in ZS, the entrance branches in FT, the all-way-stop intersections in EP and MA as well as the two-way-stop intersection in GL.

Also, scenarios in FT and GL are relatively unstructured since there is no explicit lane restrictions in the roundabout or intersections. Vehicles can exploit the space to achieve their goals, sometimes showing irrational and highly dangerous behavior. For instance, Fig. 3 shows a dangerous insertion of V0 between two vehicles (V1 and V2) stopping and making left turns from Branch 3 to Branch 5 in GL (refer to Fig. 2). The driver of V0 intended to drive from Branch 1 to Branch 4 but there was no explicit road structure for the driver.

Moreover, aggressive or irrational behavior can often be found due to inexplicit nominal or practical RoW. Vehicles may arrive at the stop bars almost at the same time and drivers may negotiate with each other by inching or even accelerating in MA and EP. The traffic in FT and GL can be very busy and it may take even minutes for the vehicle without nominal RoW to enter and pass, making the driver impatient. Also, there may be a queue of vehicles waiting behind and even honking to put social pressures to the one in the front of queue. Although there are explicit traffic rules on who goes first for roundabouts or 2-way-stop intersections, vehicles without nominal RoW may be aggressive, and vehicles with nominal RoW are mostly aware of such potential violations and are ready to react. For example, V0 in Fig. 4 was entering the roundabout in FT from Branch 3, while V1 was in the roundabout holding the RoW. However, V0 violated the rule and forced V1 to stop and yield.

Those factors significantly increase the complexity of the motions in the dataset and bring forward lots of challenging but valuable research topics for the community.

III-D Criticality

As discussed in Section III-C, vehicles holding the nominal RoW (in the roundabout of FT or on the straight road of GL) may often encounter slight violations from vehicles without nominal RoW (entering the roundabout or intersection from branches controlled by stop signs). Moreover, the vehicles holding the RoW may have relatively high speed (40 km/h or even higher). Therefore, critical situations can be observed in the dataset where time-to-collision-point (TTCP) can be extremely low. A slight collision can even be found in the dataset.

Fig. 5 shows a near-collision case in GL. V0 was making a left turn from Branch 5 (with a stop sign) to Branch 6, while V1 (with the RoW) was going straight forward from Branch 6 to Branch 3 with a relatively high speed. V1 had to execute emergency swerve to avoid the collision with V0, which was very dangerous.

Besides the critical, near-collision cases, a slight collision shown in Fig. 6 can also be found in the dataset in GL. V0 was making a right turn from Branch 5 (with a stop sign) to Branch 3, while V1 (with the RoW) was making a right turn from Branch 6 to Branch 4. In this situation, the driver of V0 might have predicted that V1 was going straight to Branch 3, so that V0 could accelerate in advance.

III-E Semantic Map

Map information is crucial for behavior-related research areas. The information required is twofold. The basic requirement is the physical layer containing a set of points or curves representing curbs, road markings (lane markings, stop bars, etc.) and other key features. In addition to the physical layer, semantic information is also necessary, which includes but is not limited to, 1) reference paths, 2) lanelets as well as their connections and turn directions, 3) traffic rules and RoW associated, etc. Moreover, such information needs to be organized with consistent format and toolkit to facilitate the users when utilizing the map. All the aforementioned requirements are met in our dataset, and more detailed information on map construction can be found in Section IV-C.

IV Construction of Motion Data and Maps

In this section, we will discuss the pipeline for constructing the motion data from both drones and traffic cameras, as well as the corresponding semantic maps.

Video stabilization and alignment: Due to gradual or sudden drift and rotation of drones, the collected videos need to be stabilized via video stabilization algorithms with transformation estimator. Also, similarity transformation is applied to project all the frames to the first one and aligned with the map.

Detection: In order to obtain accurate bounding boxes of the moving obstacles, Faster R-CNN is applied. The boxes are highly accurate, and very few inaccurate detections are rectified manually.

Data association, tracking and smoothing: Kalman filter is applied for data association and tracking. To obtain smooth motions of the vehicles, a Rauch-Tung-Striebel (RTS) smoother is also incorporated.

IV-B Motions from Traffic Camera Data

The data processing pipeline for motions from traffic camera data mainly contains the following steps, and more details, including the camera parameter estimation, can be found in .

Detection: To detect vehicles and pedestrians in each frame, we use a state-of-the-art object detector , which provides detections with 2D bounding box, instance mask and instance type.

Data association: Detections are grouped into tracks using a combination of an Intersection-over-Union tracker which associates detections with high mask overlap in successive frames, and a visual tracker to compensate for miss detections.

Tracking and smoothing: Once detections are grouped into tracks, trajectories on the ground plane are estimated using a RTS smoother. For the observation model, we use a pin-hole camera model . This allows to incorporate measurements and uncertainty directly in pixels, capturing the uncertainty due to the resolution, position and orientation of the camera. For vehicles, the RTS smoother uses a bicycle model as process model, allowing to capture the kinematics constraints of vehicles.

IV-C Construction of the High Definition Maps

As public roads are structured environments, the particular road layout of a certain area strongly affects the motion of all traffic participants. The structure for vehicles mostly starts by subdividing the road into lanes, and later combining them to create junctions, roundabouts, on ramps and so on. Further, movement within this structured area is guided by traffic rules, such as speed limits or prioritizing one road over another. In order to model such coherence, simply mapping center-lines of all lanes is not sufficient anymore.

Thus, in order to allow for a thorough analysis of the recorded trajectories, we provide centimeter-accurate high definition maps in the lanelet2 format . Within lanelet2, the physical layer of the road network, such as road borders, lane markings and traffic signs is stored. An exemplary physical layer is visualized in Figure 7. From this layer, atomic lane elements, called lanelets, are created. They describe the course of the lane and form the basis for so called regulatory elements, which determine traffic regulations such as the right of way or the speed limit.

When used alongside the recorded trajectories, these lanelet2 maps facilitate the reasoning about why some vehicles decelerate while approaching a junction, or why others do not, depending on the right of way but also on the presence of other traffic participants that potentially interact.

V Statistics of the Dataset

The dataset contains motion data collected in four categories of scenarios: roundabout, unsignalized intersection, signalized intersection, merging and lane change, as shown in Fig. 2. A detailed summary of the dataset is listed in Table II. In the roundabout scenarios, 10479 trajectories of vehicles from five different locations were recorded for around 365 minutes. Similarly, in the unsignalized intersection scenarios, three locations were included and 14867 trajectories were collected for around 433 minutes. In the merging and lane change scenarios, 10933 trajectories were recorded at two locations for around 133 minutes. Finally, one location was selected for the signalized intersection, which provided 3775 trajectories for around 60 minutes.

V-B Metrics for Interactive Behavior Identification

To represent the density of the interactive behavior of the proposed dataset, we use the metric - number of interaction pairs per vehicle (IPV) as in proposed in . To calculate the IPV, a set of rules were proposed in to extract the interactive behavior under different spatial representations of vehicle paths. The set of rules and metric are briefly reviewed below.

Minimum time-to-conflict-point difference (△TTCPmin⁡\triangle TTCP_{\min}): △TTCPmin⁡\triangle TTCP_{\min} is a metric to describe the relative states of two moving vehicles in a scenario where the paths of the two vehicles share a conflict point but without any forced stop. As shown in Fig. 8, such vehicle paths include two categories: (1) paths with static crossing or merging points such as intersections (Fig. 8 (a)-(b)), and (2) paths with dynamic crossing or merging points such as ramping and lane-changing, as shown in Fig. 8 (c)-(d). In such scenarios, merging can happen anywhere in the shaded area. We define △TTCPmin⁡\triangle TTCP_{\min} as

V-C Distribution of Interactivity

Based on the set of rules, there are 13375 interactive pairs of vehicles in the proposed dataset. We compare the interactivity among three datasets: the proposed INTERACTION dataset, the highD dataset, and the NGSIM dataset. Results are shown in Fig. 9, where the x-axis represents the length of △TTCPmin⁡\triangle TTCP_{\min} in seconds, and the y-axis are the number of vehicles (Fig. 9 (a)) and the density of vehiclesThe density is given by: density=number of vehicles with particular △TTCPmin⁡total number of vehicles in the dataset\text{density}=\dfrac{\text{number of vehicles with particular }\triangle TTCP_{\min}}{\text{total number of vehicles in the dataset}}. (Fig. 9 (b)), respectively. We can see that the INTERACTION dataset contains more intensive interactions with △TTCPmin⁡≤1\triangle TTCP_{\min}\leq 1s.

VI Utilization Examples

The proposed dataset is intended to facilitate researches related to driving behavior, as mentioned in Section I. In this section, we provide several utilization examples of the proposed dataset, including motion/trajectory prediction, imitation learning, motion planning and validation, motion clustering and representation, interaction extraction and human-like behavior generation.

Motion/trajectory prediction is of vital importance for autonomous vehicles, particularly in situations where intensive interaction happens. To obtain an accurate probabilistic prediction model of vehicle motion, both learning- and planning-based approaches have been extensively explored. By providing high-density interactive trajectories along with HD semantic maps, the proposed dataset can be used for both approaches.

For instance, proposed a deep latent variable model based on Wasserstein auto-encoder (WAE) to improve the interpretability. It incorporated the structure of recurrent neural network with vehicle kinematic model such that the output can be constrained. The motion data in FT was utilized to train and test the model in comparison with other state-of-the-art models such as variational auto-encoder (VAE), auto-encoder, and generative adversarial network (GAN). Quantitative results shown in Table III demonstrated that the proposed WAE-based method can outperform other state-of-the-art models, when comparing the root mean square error (RMSE) and mean absolute error (MAE) of the prediction for position and yaw angle.

On the other hand, took advantage of the HD semantic maps and combined the learning-based and the planning-based prediction methods. A deep learning model based on conditional variational auto-encoder (CVAE) and an optimal planning framework based on inverse reinforcement learning are dynamically combined to predict both irrational and rational behavior of the vehicles. Benefiting from the the HD semantic information, features for the deep learning model were defined in Frenet frame, which generated much better prediction performance in terms of generalization. Some exemplar results are given in Fig. 11.

VI-B Imitation Learning

The driving behavior in the proposed dataset can also be used for imitation learning which directly imitates how human drive in complicated scenarios. We extended the fast integrated learning and control framework proposed in in the FT roundabout scenario. As shown in Fig. 12, both the semantic HD map information and the states of surrounding vehicles (the red boxes) were included as the features. The grey box represents the current position of the ego vehicle. The green boxes and blue boxes, respectively, are the ground truth future positions and generated future positions of the ego vehicle via the imitation network.

VI-C Validation of Decision and Planning

Besides motion prediction and imitation, the motion data and maps in the dataset can also be used for testing different decision making and motion planning algorithms. The data-replay motions in the dataset are more suitable to test the performances of the decision-maker and planner when the motions of surrounding entities are independent of the ego motions. For example, the motion of the ego vehicle may not effect others when it does not have the RoW, or it has the RoW but others violate the rules or ignore the ego motion.

The environmental representation and motion planning methods proposed in were tested in the FT roundabout scenario. Fig. 13 is a bird’s-eye-view screen-shot of the simulation. The red rectangle represents the autonomous vehicle with the planner in . It was decelerating to avoid the collision with a vehicle entering the roundabout although it has the RoW.

We also combined the integrated decision and planning framework proposed in and the sample-based motion planner proposed in to design the decision-maker and planner under uncertainty. The predictor was designed according to based on dynamic Bayesian network (DBN) to provide the probabilities of the intentions of others.

VI-D Motion Clustering and Representation Learning

The X-means algorithm was employed to cluster the trajectories and obtain motion patterns with results shown in Fig. 15. We constructed a feature space with vehicle motions in Frenét Frame based on map information. Fig. 15 (a) shows the clustered trajectory segments in different colors with the map. Fig. 15 (b) and (d) demonstrate the cluster results with longitudinal positions and speeds of the two interacting vehicles as the coordinates. The clustering results with the first and second components of principle component analysis (PCA) for the feature space are shown in Fig. 15 (c). In the figures we can see that different interactive motions are separated and similar ones are clustered, which are desirable results to obtain motion patterns.

VI-E Extraction of Interactive Agents and Trajectories.

The proposed dataset can also be used to learn the interaction relationships between agents. We implemented the learning method and network structure proposed in to extract the interaction frames of two agents. Some example results are given in Fig. 16, where Fig. 16 (a) and (b) provide one exemplar pair of interacting cars in the FT scenario, while Fig. 16 (c) and (d) represent another pair. In Fig. 16 (a) and (c), the paths of both of the interacting cars are provided, and in Fig. 16 (b) and (d), the trajectories along longitudinal directions are shown. We can see that the extracted interaction frames (purple circle) align quite well with the ground truth frames (blue star).

VI-F Human-like Decision and Behavior Generation

We can also learn decision-making models that generate human-like decisions and behaviors with the proposed dataset. In , an interpretable human behavior model was proposed based on the cumulative prospect theory (CPT). As a non-expected utility theory, CPT can well explain some systematically biased or “irrational” behavior/decisions of human that cannot be explained by the expected utility theory. Parameters of three different models were learned and tested using the data in the FT roundabout scenario: a predefined model based on time-to-collision-point (TTCP), a learning-based model based on neural networks, and the proposed CPT-based model. The results (Fig. 17) showed that the CPT-based model outperformed the TTCP model and achieved similar performance as the learning-based model with much less training data and better interpretability.

VII Conclusion

In this paper, we presented a motion dataset in a variety of highly interactive driving scenarios from the US, Germany, China and other countries, including signalized/unsignalized intersections, roundabouts, ramp merging and lane change from cities and highway. Complex interactive motions were captured, featuring inexplicit right-of-way, relatively unstructured roads, as well as aggressive and irrational behavior caused by impatience and social pressure. Critical (near-collision and slight-collision) situations can be found in the dataset. We also included high-definition (HD) maps with semantic information for all scenarios in our dataset. The data was recorded from drones and traffic cameras and the data processing pipeline was briefly described. Our map-aided dataset with diversity, internationality, complexity and criticality of scenarios and behavior can significantly facilitate driving-behavior-related research such as motion prediction, imitation learning, decision-making and planning, representation learning, interaction extraction, and human-like behavior generation, etc. Results from various kinds of methods of these research areas were demonstrated utilizing the proposed dataset.

VIII Acknowledgement

The authors also would like to thank the Karlsruhe House of Young Scientists (KHYS) for their support of Maximilian’s research visit at MSC Lab.

References