Trust, but Verify: Cross-Modality Fusion for HD Map Change Detection

John Lambert, James Hays

Introduction

We live in a highly dynamic world, so much so that significant portions of our environment that we assume to be static are, in fact, in flux. Of particular interest to self-driving vehicle development is changing road infrastructure. Road infrastructure is often represented in an onboard map within a geo-fenced area. Geo-fenced areas have served as an operational design domain for self-driving vehicles since the earliest days of the DARPA Urban Challenge .

One way such maps could be used is to constrain navigation in all free space to a set of legal “rails” on which a vehicle can travel. Maps may also be used to assist in planning beyond the sensor range and in harsh environments. Besides providing routes for navigation, maps can ensure that the autonomous vehicle (AV) follows local driving laws when navigating through a city. They embody a representation of the world that the AV can understand, and contain valuable information about the environment. Building and validating maps represent an essential and general part of spatial artificial intelligence (AI) , embodied intelligence that enables safety-critical awareness of a robot’s surroundings.

However, maps assume a static world, an assumption which is violated in practice; although these changes are rare, they certainly occur and will continue to occur, and can have serious implications. Level 4 autonomy is defined as sustained performance by an autonomous driving system within an operational design domain, without any expectation that a user will respond to a request to intervene . Thus, constant verification that the represented world, expressed as a map, matches the real world, a task known as map change detection , is a clear requirement for L4 autonomy. This problem may be solved in the classification setting, reasoning globally over the map and scene, or additionally in a localization setting, where the spatial extent of changed map entities are locally identified. Because dedicated mapping vehicles cannot traverse the world frequently enough to keep maps up to date , high-definition (HD) maps become “stale,” with out of date information. If maps are used as hard priors, this could lead to confident but incorrect assumptions about the environment.

In this work, we present the first public dataset for urban map change detection based on actual, observed map changes, which we name TbV. While researchers could use paired sensor data and HD maps from other datasets such as Argoverse or nuScenes to hypothesize about map change detection performance in the real world based on synthetic data, these datasets include zero real map changes, meaning one could not know with any accuracy or degree of certainty about how well the system would actually operate in the real world. Concurrent work presents qualitative results on a handful of real-world map changes on a proprietary dataset, but relies upon synthetic test datasets for all quantitative evaluation. Not only does no comparable dataset to TbV exist, there also has not even been an attempt to characterize how often map changes occur and what form they take. Collecting data for map change detection is challenging since changes occur randomly and infrequently. In addition, in order to use data corresponding to real changes to train and evaluate models, identified changes must be manually localized in both space and time.

HD map change detection is a difficult task even for humans, as it requires the careful comparison of all nearby semantic entities in the real world with all nearby map elements in the represented world. In an urban scene, there can be dozens of such entities, many with extended shapes. The task is sufficiently difficult that several have even questioned the viability of HD maps for long-term autonomy, opting instead to pursue HD-map-free solutions . We concentrate on changes to two types of semantic entities – lane geometry and pedestrian crosswalks. We define the task as correctly classifying whether a change occurred at evenly spaced intervals along a vehicle’s trajectory.

The task itself is relatively new, especially since HD maps were not made publicly available until the release of the Argoverse, nuScenes, Lyft Level5, and Waymo Open Motion datasets . We present the first entirely learning-based formulation for solving the problem in either a bird’s eye view (BEV), as well as a new formulation for the ego-view (i.e. front camera frustum), eliminating several heuristics that have defined prior work. We pose the problem as learning a representation of maps, sensor data, or the combination of the two.

We present a novel AV dataset, with 799 vehicle logs in our train and synthetic validation splits, and over 200 vehicle logs with real-world map-changes in our real val and test splits.

We implement various learning-based approaches as strong baselines to explore this task for the first time with real data. We also demonstrate how gradients flowing through our networks can be leveraged to localize map changes.

We analyze the advantages of various data viewpoint by training both models operating on the ego-view and others on a bird’s eye view.

We show that synthetic training data is useful for detecting real map changes. At the same time, we identify a considerable domain gap between synthetic and real data, with significant performance consequences.

Related Work

HD Maps. HD maps include lane-level geometry, as well as other geometric data and semantic annotations . The Argoverse , nuScenes , Lyft Level 5 , Waymo Open Motion , and Argoverse 2.0 datasets are the only publicly available sources of HD maps today, all with different semantic entities. Argoverse includes a ground surface height map, rasterized driveable area, lane centerline geometry, connectivity, and other attributes. nuScenes followed by also releasing centerline geometry, pedestrian crossing polygons, parking areas, and sidewalk polygons, along with rasterized driveable area. Lyft Level 5 later provided a dataset with many more map entities, going beyond lane marking boundaries, crosswalks to provide traffic signs, traffic lights, lane restrictions, and speed bumps. The Waymo Open Motion Dataset released motion forecasting scenario data with associated HD maps. Their yet-richer HD map representation includes crosswalk polygons, speed bump polygons, lane boundary polylines with marking type, lane speed limits, lane types, and stop sign positions and their corresponding lane associations; their map data is most comparable with our HD maps. HD maps are useful for a range of tasks, from perception , to motion forecasting, to motion planning , to traffic scene simulation . All state-of-the-art motion forecasting methods for self-driving today use HD maps .

HD map change detection is a recent problem, with limited prior work. Pannen et al. introduce one of the first approaches; two particle filters are run simultaneously, with one utilizing only Global Navigation Satellite System (GNSS) and odometry measurements, and the other filter using only odometry with camera lane and road edge detections. These two distributions and sensor innovations are then fed to weak classifiers. Other prior work in the literature seeks to define hand-crafted heuristics for associating online lane and road detections with map entities . These methods are usually evaluated on a single vehicle log .

Instead of comparing vector map elements, Ding et al. use 2d BEV raster representations of the world; first, IMU-motion-compensated LiDAR odometry is used to build a local online “submap”. Afterwards, the submap is projected to 2d and overlaid onto a prebuilt map; the intensity mean, intensity variance, and altitude mean of corresponding cells are compared for change detection. Rather than pursuing this approach, which requires creating and storing high-resolution reflectance maps of a city, we pursue the alignment of vector maps with sensor data. Vector maps can be encoded cheaply with low memory cost and are the more common representation, being used in all five public HD map datasets.

In concurrent work, Heo et al. introduce an adversarial metric learning-based formulation for HD map change detection, but access to their dataset is restricted to South Korean researchers and performance is measured on a synthetic dataset, rather than on real-world changes. They employ a forward-facing ego-view representation, and require training a second, separate U-Net model to localize changed regions in 2d, whereas we show changed entity localization can come for free via examination of the gradients of a single model.

Mapping Dynamic Environments.

While “HD maps” are a relatively new entity, dynamic map construction is a more mature field of study. Semi-static environments are not limited to urban streets; households, offices, warehouses, and parking lots are relatively fixed environments that a robot may navigate, with changing cars, furniture, and goods . Mapping dynamic environments has been an area of study within the SLAM community for decades . However, we focus purely on change detection, rather than map updates.

Recently, machine learning for online mapping has generated interest. An alternative to using an HD map prior is to rebuild the map on-the-fly during robot operation; however, such an approach cannot map occluded objects or entities. In addition, these methods are limited to producing raster map layers, such as a driveable area mask, with an output resembling semantic segmentation. Raster data is significantly less useful than vector data for path planning and generating vector map data with machine learning is generally an unsolved problem. Raster map layers may be generated from LiDAR , accumulated from networks operating on ego-view images over multiple cameras and timesteps , or from a single image paired with a depth map or LiDAR . They all show that automatic mapping is quite challenging.

Image-to-Image Change Detection.

Image-to-image scene change detection over the temporal axis is a well-studied problem . Scenes are dynamic over time in numerous ways, and those ways are mostly nuisance variables for our purposes. We wish to develop models invariant to season, lighting, the fading of road markings, and occlusion because these variables don’t actually change the lane geometry. Wang et al. introduced the CDnet benchmark, a collection of videos with frame pixels annotated as static, shadow, non-ROI, unknown, or moving. Alcantarilla et al. introduce the VL-CMU-CD street view change detection benchmark, from a subset of the Visual Localization CMU dataset.

The TbV Dataset

We curate a novel dataset of autonomous vehicle driving data comprising 1043 logs, over 200 of which contain map changes. The vehicle logs are on average 54 seconds in duration, collected in six North American cities: Austin, TX, Detroit, MI, Miami, FL, Palo Alto, CA, Pittsburgh, PA, and Washington, D.C.

Our training set consists of real data with accurate corresponding onboard maps (“positives”). Accordingly, synthetic perturbation of positives to create plausible “negatives” is required for training. We release the data, code and API to generate them. However, in the spirit of other datasets meant for testing only (i.e. not training) such as the influential WildDash dataset , we curate our validation and test splits from the real-world distribution. We do so since map changes are difficult to mine , thus we save their limited quantity for testing and evaluation. We provide a few examples from our 244 validation and test logs in Figure 1. Statistics of the train, validation, and tests split are described in Table 1.

In order to label map changes, we use three rounds of human filtering, where changes are identified, confirmed, and characterized by three independent reviewing panels. We assign spatial coordinates to each changed object within a city coordinate system. Crosswalk changes are denoted by a polygon, and lane geometry changes by polylines. We use egovehicle-to-changed-map-entity distance (point-to-polygon or point-to-line) to determine whether or not a sensor and map rendering should be classified as a map-sensor match or mismatch.

We use our annotated map changes, along with 5 months of fleet data, to analyze the frequency of map changes on a city-scale across several cities. Two particular questions are of interest: (1) how often will an autonomous vehicle encounter a map change as part of its day-to-day operation? and (2) what percentage of map elements in a city will change each month or each year? For our analysis, we subdivide a city’s map into square spatial units of dimension 30 meters ×\times 30 meters, often referred to as “tiles” in the mapping community. We find the probability pp of an encounter at any given time with a tile with changed lane geometry or crosswalk to be p≈5.5174×10−5p\approx 5.5174\times 10^{-5}. Although the probability of a single event is low, we cannot ignore rare events, as doing so would be reckless. Given the 3.225 trillion miles driven in the U.S. per year , this could amount to billions of such encounters per year. We determine that up to 7 of every 1000 map tiles may change in a 5-month span (see Table 2), a significant number. More details are provided in the Appendix.

2 Sensor Data

Our TbV dataset includes LiDAR sweeps collected at 10 Hz, along with 20 fps imagery from 7 cameras positioned to provide a fully panoramic field of view. In addition, camera intrinsics, extrinsics and 6 d.o.f. AV pose in a global coordinate system are provided. LiDAR returns are captured by two 32-beam LiDARs, spinning at 10 Hz in the same direction (“in phase”), but separated in time by 180∘180^{\circ}. The cameras trigger in-sync with both of them, leading to a 20 Hz framerate. The 7 global shutter cameras are synchronized to the LiDAR to have their exposure centered on when the LiDAR sweeps through the middle of their fields of view. The top LiDAR spins clockwise in its frame, while the bottom LiDAR spins counter-clockwise in its frame; in the ego-vehicle frame, they both spin clockwise.

3 Map Data

In the Appendix, we list the semantic map entities we include in the TbV dataset. Previous AV datasets have released sensor data localized within a single map per city . This is not a viable solution for TbV, since the maps change over our long period of data gathering. We instead release local maps with all semantic entities within 100 meters of the egovehicle featured. Accordingly, single, incremental changes can be identified and tested. We release many maps, one per vehicle log; the corresponding map is the map used on-board at time of capture. Lane segments within our map are annotated with boundary annotations for both the right and left marking (including against curbs) and are marked as implicit if there is no corresponding paint. We use the same release format for our maps as Argoverse 2.0 uses.

4 Dataset Taxonomy

Our dataset’s taxonomy is intentionally oriented towards lane geometry and crosswalk changes. In general, we focus on permanent changes, which are far less frequent in urban areas than temporary map changes. Temporary map changes often arise due to construction and road blockades.

We postulate that temporary map changes – temporarily closed lanes or roads, or temporary lanes indicated by barriers or cones, should be relegated to onboard object recognition and detection systems. Indeed, recent datasets such as nuScenes include 3d labeling for traffic cones and movable road barriers, such as Jersey barriers (see Appendix for examples). Even certain types of permanent changes are object-centered (e.g. changes to traffic signs). Accordingly, a natural division arises between “things” and “stuff’ in map change detection, just as in general scene understanding . We focus on the “stuff” aspect, corresponding to entities which are often spatially distributed in the BEV; we find lane geometry and crosswalk changes to be more frequent than other “stuff”-related changes.

Approach

We formulate the learning problem as predicting whether a map is stale by fusing local HD map representations and incoming sensor data. We assume accurate pose is known. At training time, we assume access to training examples in the form of triplets (x,x∗,y)(x,x^{*},y), where xx is a local region of the map, x∗x^{*} is an online sensor sweep, and yy is a binary label suggesting whether a “significant” map change occurred. (x,x∗)(x,x^{*}) should be captured in the same location.

We explore a number of architectures to learn a shared map-sensor representation, including early fusion and late fusion (see Figure 2). The late fusion model uses a siamese network architecture with two input towers, and then a sequence of fully connected layers. We utilize a two-stream architecture with shared parameters, which has been shown to still be effective even for multi-modal input . We also explore an early-fusion architecture, where the map, sensor, and/or semantic segmentation data are immediately concatenated along the channel dimension before being fed to the network. We take no credit for these convnet architectures, which are well studied.

2 Synthesis of Mismatched Data

Real negatives are difficult to obtain; because their location is difficult to predict a priori, they cannot be captured in a deterministic way by driving around an urban area on any particular day. Therefore, rather than using real negatives for training, we synthesize fake negatives. While sensor data is difficult to simulate, requiring synthesis of sensor measurements from the natural image manifold , manipulating vector maps is relatively straightforward.

Synthetic data generation via randomized rendering pipelines can be highly effective for synthetic-to-real transfer . In order to synthesize fake negatives from true positives, one must be able to trust the fidelity of labeled true positives. In other words, one must trust that for true positive logs, the map is completely accurate for the corresponding sensor data. We perturb the data in a number of ways (See Appendix). Once such fidelity is confirmed and assured, vector map manipulation is trivial because map elements are vector entities which can be perturbed, deleted, or added.

While synthesizing random vector elements is trivial, sampling from a realistic distribution requires conformance to priors, including the lane graph, drivable area, and intersection. We aim for synthetic map/sensor deviations to resemble real world deviations, and real world deviations tend to be subtle, e.g. a single lane is removed or painted a different color, or a single crosswalk is added, while 90% of the scene is still a match. In order to generate realistic-appearing synthetic map objects, we hand-design a number of priors that must be respected for a perturbed example to enter our training set as a valid training example (see Appendix). Figure 3 and Table 10 of the Appendix enumerate a full list of the 6 types of synthetic changes we employ.

3 Sensor Data Representation

We experiment with two sensor data representations – ego-view (the front center camera image) and bird’s eye view (BEV). Rather than using Inverse Perspective Mapping (IPM) , we generate the BEV representation (i.e. orthoimagery) by ray-casting image pixels to a ground surface triangle mesh. For ray-casting, we use a set of camera sensors with a panoramic field of view, mounted onboard an autonomous vehicle. The temporal nature of the data is exploited as pixel values from 7 streams of ego-view images are aggregated to render each BEV image in order to reduce sparsity (see Appendix). 70 images, capturing 10 timesteps from each of 7 frustums, are used to create each rendering.

4 Map Data Learning Representation

We render our map inputs as rasterized images; Entities are layered from the back of the raster to the front in the following order: driveable area, lane segment polygons, lane boundaries, pedestrian crossings (i.e. crosswalks). We release the API to generate and render these map images. Vector map entities are synthetically perturbed before rasterization for synthetic negative examples.

Experimental Results

We frame the map change detection task as follows: at timestamp tt, given a buffer of all past sensor data, including camera intrinsics and extrinsics, along with 6 d.o.f. egovehicle poses cityTegovehicle{}^{city}T_{egovehicle} which we denote as Ti=0…tT_{i=0\dots t}, image data {Ii=0…tc}c=1C\{I_{i=0\dots t}^{c}\}_{c=1}^{C} where cc is a camera index, lidar sweeps Li=0…tL_{i=0\dots t}, onboard map data MM, estimate whether the map is in agreement with the sensor data.

Our ego-view models that operate on front-center camera images leverage both LiDAR and RGB sensor imagery. We use LiDAR information to filter out occluded map elements from the rendering. We linearly interpolate a dense depth map from sparse LiDAR measurements, and then compare the depth of individual map elements against the interpolated depth map; elements found behind the depth map are not rendered. In our early fusion architecture, we experiment with models that also have access to semantic label masks from the semantic head of a publicly-available seamseg ResNet-50 panoptic segmentation model . For those models with access to the semantic label map modality, we append 224×224224\times 224 binary masks for 5 semantic classes (‘road’, ‘bike- lane’, ‘marking-crosswalk-zebra’, ‘lane-marking-general’, and ‘crosswalk-plain’) as additional channels in early fusion.

Bird’s Eye View Models.

Training.

We use a ResNet-18 or ResNet-50 backbone, with ImageNet-pretrained weights, where a corresponding weight parameter’s size is applicable. We use a crop size of 224×224224\times 224 from images resized to 234×234234\times 234 px. Please refer to the Appendix for additional implementation details and an ablation experiment on the influence of input crop size on performance.

2 Evaluation

We report results on an earlier, beta version of TbV, which we call TbV-Beta, which includes slightly fewer logs. The publicly released version of TbV is TbV-1.0.

We use a mean of per-class accuracies to measure performance on a two-class problem: predicting whether the real world is changed (i.e. map and sensor data are mismatched), or unchanged (i.e. a valid match). This accounts for both precision and recall. If a confusion matrix is computed with predicted entries on the rows and actual classes as the columns, and normalized by dividing by the sum of each column, 2-class accuracy can be simply calculated as the mean of the diagonal of the confusion matrix. More formally, let ncl=2n_{cl}=2 be the number of classes, y^i\hat{y}_{i} be the prediction for the ii’th test example, and yiy_{i} be the ground truth label for the ii’th test example. We define per-class accuracy (Accc\text{Acc}_{c}) and mean accuracy (mAcc) as:

Models that operate on an ego-view scene perspective are more effective than those operating in the bird’s eye view (5% more effective over their own respective field of view), achieving 72.3% mAcc (see Table 3). We found a simpler architecture (ResNet-18) to outperform ResNet-50 in the ego-view.

Which modality fusion method is most effective?

Early fusion. For both BEV and ego-view, the early fusion models significantly outperform the late fusion models (+22.8% mAcc in the ego-view and +9.7% mAcc in BEV). This may be surprising, but we attribute this to the benefit of early alignment of map and sensor image channels for the early fusion models. Instead, the late-fusion model performs alignment with greatly reduced spatial resolution in a higher-dimensional space, and is forced to make decisions about both data streams independently, which may be suboptimal. While the map and sensor images represent different modalities, a shared feature extractor is useful for both.

Which input modalities are most effective?

A combination of RGB sensor data, semantics, and the map. We compare validation and test set performance of various input modalities in Table 4. Early fusion of map and sensor data is compared with models that have access to only sensor or map data, or a combination of the two, with or without semantics. All models suffer a significant performance drop on the test set compared to the validation set. While a gap between validation and test performance is undesirable, better synthesis heuristics and better machine learning models can close that gap. We find semantic segmentation label map input to be quite helpful, although it places a dependency upon a separate model at inference time, increasing latency for a real-time system. Mean accuracy improves by 4% in the ego-view and 2% in the BEV when sensor and map data is augmented with semantic information in early fusion. In fact, early fusion of the map with the semantic stream alone (without sensor data) is 1% more effective than using corresponding sensor data for the BEV.

Is sensor data necessary?

Yes. The “map-only” model we trained couldn’t meaningfully identify map changes over random chance, achieving mean accuracies on the test set of just 0.5754 in the BEV and 0.533 in an ego-view (see Table 4). This model can also be thought of as a binary classifier which is trained to identify whether the HD map is real or synthetically manipulated (i.e. classify the source domain). Inspection via Guided GradCAM demonstrates that the map-only baseline attends to onboard map areas that are not in compliance with real-world priors, such as identifying asymmetry in crosswalk layout, paint patterns, and lane subdivisions.

Ablation on Modality Dropout.

We find random drop-out of certain combinations of modalities to regularize the training in a beneficial way, improving accuracy by more than 1% of our best model (See Table 5). Given the wide array of modalities available to solve the task, from RGB sensor data, semantic label maps, rendered maps, and LiDAR reflectance data, we experiment with methods to force the network to learn to extract useful information from multiple modalities. Specifically, we perform random dropout of modalities, an approach developed in the self-supervised learning literature .

Perhaps the most intuitive approach would be to apply modality dropout to one of the sensor or semantic streams, forcing the network to extract useful features from both modalities during training. However, we find this is in fact detrimental. More effective, we discover, is to randomly drop out either the map or semantic streams. In theory, meaningful learning should be impossible without access to the map; however, since we drop-out each example in a batch with 50% probability, in expectation 50% of the examples should yield useful gradients in each batch. This approach improves accuracy by more than 1% of our best model. We zero out all data from a specific modality as our drop out technique.

Interpretability and Localizing Changes.

While accurately perceiving changes is important, the ability to localize them would also be helpful. Bounding boxes are often unsuitable for a compact localization of a map change because most changed map regions in TbV relate to “stuff” classes, for example, linear, extended lane boundary markings. Instead, pixelwise localization is often more appropriate. We use Guided Grad-CAM to identify which regions of the sensor, map, and semantic input are most relevant to the prediction of the ‘is-changed’ class. In Figure 4, we show qualitative results on frames for which our best model predicts real-world map changes have occurred.

Computational Runtime.

In order to demonstrate that our models can be used in practice without introducing a heavy computational burden, we report the time to complete a feedforward pass through a network, for our various models (see Table 6). The network architectures we employ are computationally lightweight, using ResNet-18 or ResNet-50 backbones. We report feedforward runtime (in milliseconds) for each network, averaged over 100 forward passes, after a warm-up period of 20 forward passes. The hardware used for the analysis is a Quadro P5000 GPU and Intel(R) Xeon(R) W-2145 CPU @ 3.70GHz processor, running the Ubuntu 20.04 operating system. Our best performing map change model (operating in the ego-view) requires just 2.24 milliseconds to complete a forward pass. However, this model does assume access to a readily-available semantic labelmap produced by a semantic segmentation network, which would increase the system latency. In Table 4, we show that for a 4% drop in accuracy for the ego-view models, the semantic labelmap can be excluded from the inputs, in which case the total runtime is just 2.19 milliseconds, well within real-time performance, introducing minimal latency for perception or planning modules which might utilize this information.

Conclusion

Discussion. In this work, we have introduced the first dataset for a new and challenging problem, map change detection. Our dataset is one of the largest AV datasets at the present time, featuring 1043 logs with an average duration of 54 seconds. We implement various approaches as strong baselines to explore this task for the first time with real data. Perhaps surprisingly, we find that comparing maps in a metric representation (a bird’s eye view) is inferior to operating directly in the ego-view. We attribute this to a loss of texture during the projection process, and to a more difficult task of reasoning about a much larger spatial area (85∘85^{\circ} f.o.v. instead of 360∘360^{\circ} f.o.v.). In addition, we provide a new method for localizing changed map entities, thereby facilitating efficient updates to HD maps.

We identify a significant gap between validation accuracy and test accuracy – 10-20% less on the test split – which supports the importance of testing on real data. If performance is only measured on fake changes that resemble one’s training distribution, performance can appear much better than what occurs in reality. Real changes can be subtle, and we hope the community will use this dataset to further push the state-of-the-art we introduce. We make publicly available our data, models, and code to generate our dataset and reproduce our results.

Rendering time. A second limitation of our work is that real-time rendering requires GPU hardware; in the ego-view, map entity tesselation and rasterization are costly, whereas in the BEV, ray-casting is computationally intensive. Perturbation diversity. In our work, we introduce just 6 types of possible map perturbations, of which far more types are possible; nonetheless, we prove that they are surprisingly useful. Accuracy. Perhaps last of all, although our baselines have reasonable performance and by inspection we demonstrate they are learning to attend to meaningful regions, a large gap still exists before such a model would be accurate enough to be used on-vehicle.

References

Appendix

In this appendix, we provide additional details about our dataset and experiments. In Section (A), we provide an ablation study on the influence of input crop size on model performance. In Section (B), we discuss additional implementation details about our training, data augmentation, and occlusion-based map rendering process. In Section (C), we discuss the paired positive-negative logs we include. In Section (D), we describe our evaluation metric. In Section (E), we provide additional experimental analysis of different models and rendering viewpoints. In Section (F), we provide additional details about how we generate orthoimagery. In Section (G), we offer additional examples from our test set. In Section (H), we give examples of other types of temporary map changes which we do not annotate or evaluate within our dataset. In Section (I) we provide further analysis of the frequency of map changes. In Section (J), we give additional details about our synthetic map perturbation protocol. In Section (K), we provide a datasheet for the dataset.

Appendix A: Influence of Input Crop Size

In this section, we perform an ablation on input crop size, as discussed in Section 5.1 of the main text. In the main paper, we set our input crop size to 224×224224\times 224 px for all experiments mentioned therein. In this section, we present an ablation to measure the influence of input crop size. Again, we find the ego-view model is the best-performing model, as measured on its own field of view. Perhaps surprisingly, we find that an RGB image at 234×234234\times 234 px resolution (∼\sim 164K pixel values/image) is sufficient to capture significant detail. In Table 7, we present an ablation where we find that for BEV models, higher resolution (i.e. 468×468468\times 468 px) does improve mAcc by 2%2\% mAcc, although requiring almost 4x the GPU memory during training and significantly longer training times. However, for ego-view models, a higher crop size is quite detrimental, reducing visibility-based mAcc by around 7%.

Appendix B: Additional Implementation Details

We train our models for 90 epochs with the Adam optimizer. We use a polynomial learning rate decay strategy, starting at 1×10−31\times 10^{-3}. We use a batch size of 1024 examples. We start with pretrained ImageNet weights for ResNet-18 or ResNet-50 .

We train with multiple negative examples per sensor image, which we found to be more beneficial than randomly sampling a single negative example (i.e. a synthetically perturbed map). In other words, we perform multiple types of perturbations for a given scene, and feed them to the network as separate negative examples (not necessarily in the same mini-batch).

B.2. Data Augmentation

We employ a number of data augmentation techniques to improve the generalization of our models and prevent overfitting. Input images are of dimension 2048×15502048\times 1550 for the front-center camera, and 1550×20481550\times 2048 for all other 6 cameras. For the ego-view models, we first take a square crop from the bottom 1550×15501550\times 1550 of an ego-view image. Afterwards, we resize to 234×234234\times 234, perform a random horizontal flip with 50% probability, take a random 224×224224\times 224 crop, divide pixel intensities by 255, and then normalize both sensor and map RGB channels by the ImageNet mean (μr,μg,μb)=(0.485,0.456,0.406)(\mu_{r},\mu_{g},\mu_{b})=(0.485,0.456,0.406) and standard deviation (σr,σg,σb)=(0.229,0.224,0.225)(\sigma_{r},\sigma_{g},\sigma_{b})=(0.229,0.224,0.225)

For BEV models, we resize input images from 2000×20002000\times 2000 px to 234×234234\times 234 px, perform a random horizontal and/or vertical flip with 50% probability each (independently), choose a random 224×224224\times 224 crop, and normalize as described above.

We find other traditional data augmentation techniques from the semantic segmentation literature , such as applying a random rotation to the input or randomly blurring the input with a small kernel, to be ineffective.

B.3. Occlusion Reasoning

As discussed in Section 5.1 of the main text, we use map occlusion reasoning when generating the input for our ego-view models. Occluded map elements and map elements that have been removed in the real world (“deleted”) are both not visible in camera imagery. While the former is an expected everyday occurrence, and the latter is of interest to us, we use occlusion reasoning in order to separate the two phenomena. We generate a dense depth map from sparse LiDAR returns (see Figure 5) and the depth of map entities is compared against the corresponding depth of its projection in the depth map.

B.4. Details about Semantic Label Map Input

As discussed in Section 5.1 of the main text, we use semantic label maps generated from the semantic head of a publicly-available seamseg ResNet-50 panoptic segmentation model Available at https://github.com/mapillary/seamseg.. We create 5 binary mask channels from the semantic label map, for the ‘road’, ‘bike-lane’, ‘marking-crosswalk-zebra’, ‘lane-marking-general’, and ‘crosswalk-plain’ classes. These are optionally provided as additional channels to the 3 RGB sensor channels and 3 RGB map channels via early fusion. Seamseg’s semantic label maps on their own do not capture sufficient granularity for the map change detection task we define, since the Mapillary Vistas public dataset’s taxonomy does not differentiate between lane color and or different marking types (e.g. double-solid, solid, dashed-solid), which are of interest to autonomous vehicle operation.

Unsuitability of Per-Pixel Semantic Comparison. Directly comparing rendered map and semantic label maps at a per-pixel level is not always useful since our HD map representation does not provide paint annotation for every single dashed longitudinal lane marking, but rather provides a description lane marking pattern, polyline boundary, and other corresponding attributes (See Table 8 of the main text). Thus, we can simulate the pattern of dashed lane markings, but not their exact, pixel-perfect location. As the main text shows, the network can abstract away the per-pixel details to provide more meaningful features.

Appendix C: Data Selection

For a subset of the ‘negative’ logs in our TbV dataset, we provide a corresponding ‘positive’ log captured before the change occurred. Example images from pair positive-negative logs are provided in Figure 6. This allows for non-learning based approaches (e.g. based upon comparison of 3d reconstructed world models) for a limited amount of the test set.

Appendix D: Evaluation

As our primary accuracy metric, we use a mean of class accuracies over two classes. This accounts for both precision and recall. If a confusion matrix is computed with predicted entries on the rows and actual classes as the columns, and normalized by dividing by the sum of each column, 2-class accuracy can be simply calculated as the mean of the diagonal of the confusion matrix.

More formally, let ncl=2n_{cl}=2 be the number of classes, y^i\hat{y}_{i} be the prediction for the ii’th test example, and yiy_{i} be the ground truth label for the ii’th test example. We define per-class accuracy (Accc\text{Acc}_{c}) and mean accuracy (mAcc) as:

Appendix E: Additional Experimental Analysis

In principle, the bird’s eye view (BEV) representation (orthoimagery) offers two main advantages: a single, dense, accumulated metrically-accurate representation for a single pass through a network, rather than passing in 7 images through 7 separate networks, trained on each frustum, in order to detect changes to the sides and rear of the vehicle. This approach can be costly at inference time given the number of camera frustums required to achieve a panoramic view with traditional cameras. Second, the BEV is generally free of distortion, compared to the ego-view. The ego-view can be seen as “spoiling” the map data’s metric nature.

Advantages of Ego-view.

However, an ego-view perspective also presents clear advantages over the BEV. Rendering data in the BEV can be seen as “spoiling” the sensor data’s texture. Importantly, there is less distraction and less overall content to reason about in the egoview. Therefore, the ego-view task is arguably easier than the BEV task, needing only to detect changes in a 85∘85^{\circ} f.o.v. instead of 360∘360^{\circ} f.o.v.

Analysis of Map-Only Baseline.

The map-only baseline performs quite poorly when predicting real-world lane geometry changes, slightly over random chance (2% or 3% over random chance in the ego-view and 7% over random change in the BEV). While the map-only stream may seem doomed to fail without access to real-world sensor information, we observe that a certain number of map changes exist to bring the real world into compliance with certain priors, which are already encapsulated in the map. For example, we find that upgrading a 4-way intersection from a single crosswalk to 4 crosswalks, or from a single crosswalk to 0 crosswalks (after repaving) is a common map change, which would agree with priors. Indeed, our experimental results suggest that the map-only baseline, which is completely blind to the real-world, can occasionally succeed at predicting real-world crosswalk changes by learning powerful priors. Inspection via Guided GradCAM demonstrates that the map-only models attends to asymmetric paint patterns along the left and right boundaries of a road, or asymmetric lane subdivisions along two sides of a road; modifications to such map asymmetry which are common real-world map updates.

Analysis of Sensor-Only Baseline.

The sensor-only model (see Table 4 of the main paper) sees randomly perturbed labels, with only “positive” training data, and therefore is not a meaningful baseline.

Appendix F: Orthoimagery Generation Implementation Details

In this section, we provide additional details about the orthoimagery generation process described in Sections 4.3 and 5.1 of the main text. In order to create a metrically-accurate sensor data representation that is free of perspective distortion, we generate orthoimagery using ray-casting. Orthoimagery from LiDAR suffers from extreme sparsity, leading to an impoverished representation. To generate dense panoramic orthoimagery, we use a set of high-definition camera sensors with a panoramic field of view, mounted onboard an autonomous vehicle. We generate the BEV representation (i.e. orthoimagery) by ray-casting image pixels to a ground surface triangle mesh. Our ground height maps exploit LiDAR offline, and in this way our ego-view method incorporates the strengths of LiDAR.

We tesselate quads from a ground surface mesh with 1 meter resolution to triangles; rays are cast to triangles up to 25 m away from the egovehicle. For acceleration, we cull triangles outside of the left and right cutting planes of each camera’s view frustum. We implement the Moller-Trombore ray-triangle intersection routine in CUDA.

Density.

Ray-casting yields a vastly more dense set of image rays than LiDAR, on the order of 2 orders of magnitude greater density; for a 1550×20481550\times 2048 image, one can obtain ∼3.17\sim 3.17 million rays per image, and across 7 camera frustums, this translates to over 22.19 million rays with available RGB values per second. With 20 fps imagery per camera frustum, this amounts to 440 million rays per second. Most conventional 10 Hz LiDAR sensors can provide little more than 100k returns per sweep, and thus at most 1 million rays per second.

Aggregation.

In order to prevent holes in the orthoimagery in the area underneath the egovehicle, we aggregate pixels in ring buffer of length 10 sweeps, and wait 10 sweeps before starting rendering. Future sensor data is not used to render the sensor data representation. We use linear interpolation to account for sparsity at range.

Comparison with IPM.

While Inverse Perspective Mapping (IPM) is the dominant approach in the literature, it is inaccurate as it cannot account for ground surface variation. Geiger model the image-to-ground plane mapping as a homography (IPM) and mosaics together monocular images, but requires scenes with an approximately-planar ground surface. Zhang et al. generate orthophoto ground imagery using fisheye cameras and IPM. Rapo explored the use of dashboard-mounted cell phones without access to LiDAR or known calibration, instead relying upon SfM, optical flow, and vanishing point estimation for online calibration and also use IPM for pixel-to-world correspondence.

Appendix G: Additional Examples from Test Set

In Figure 7, we show additional examples from our test set, as seen from a bird’s eye view.

Appendix H: Map Changes from Construction

In Figure 8, we show examples of object-centric map changes inside our TbV dataset, which we do not annotate and are not the focus of our work.

Appendix I: Additional Analysis of Map Change Frequency

In Section 3.1 and Table 2 of the main paper, we present an analysis of map change frequency. In this section, we provide additional analysis, an extended table, and derivations of our estimates. Map changes occur at random as part of a stochastic process. While some changes are coordinated at a city-administration level, it is still difficult to predict to a specific date or time when construction crews will complete changes. As discussed in the main text, we reason about square spatial areas of size 30\mboxm×30\mboxm30\mbox{ m}\times 30\mbox{ m}, which we refer to as tiles, which cover 900\mboxm2900\mbox{ m}^{2} each.

We consider the probability of entering a spatial area that has undergone a crosswalk or lane geometry within it. In other words, it is the probability of encountering a changed area, and thus we name it pecap_{eca}. In order to estimate the probability of encountering a changed area, rather than computing the ratio \Big{(}\frac{\text{num. change-discovery miles}}{\text{num. fleet miles}}\Big{)}, we compute the ratio \Big{(}\frac{\text{num. tiles where change is observed}}{\text{num. tiles entered by fleet}}\Big{)}. We do not require that the autonomous vehicle directly drove over the changed tile, as an observed change can very well still affect driving behavior. We model the probability as a Bernoulli(pp) r.v., with p≈5.517×10−5p\approx 5.517\times 10^{-5} across the more than 5 North American cities we analyze. A visit would occur once per every 18,124 times a vehicle enters such areas.

While the change percentage may seem inconsequential, one must consider that drivers in the United States are estimated to drive 3.225 trillion miles per year, according to the U.S. Department of Transportation . If one were to consider our rate of change equal to the rate of change of any stretch of road within the United States, this would amount to an upper bound of 9B encounters of spatial areas with changed lane geometry or crosswalks, per year:

This derivation assumes that all roads (including highways) are changed as often as urban roads (a generous estimate).

Derivation: Probability per Spatial Area We next estimate the probability of each unique tile in a city seeing a crosswalk or lane geometry change, which we also model as a Bernoulli(pp) random variable, with pp estimated as:

where the numerator and denominator are both measured over kk months.

In Table 9, we analyze the probability of change for a 30m×30m30m\times 30m spatial area across six particular cities. Since we can likely only catch changes for spatial areas that are somewhat frequently visited, we require that an area is visited by fleet at least n=5n=5 times over k=5k=5 months.

Appendix J: Synthetic Map Perturbation Technique

In Section 4.2, Table 10, and Figure 3 of the main text, we enumerate a number of hand-designed priors we use to generate realistic-appearing synthetic maps. In this section, we provide detailed descriptions of the generation process.

Next, in order to determine how many total lane segments the crosswalk must cross in order to span the entire road, we must determine the road extent. We approximate it as the union of all nearby lane segment polygons. The line representing the principal axis of the crosswalk may intersect with this road polygon in more than two locations, since it is often non-convex. We choose the shortest possible length segment that spans the road polygon to be valid, and thus find the closest two intersections to the sampled centerline waypoint. We randomly sample a crosswalk width ww in meters from a normal distribution w∼N(μ=3.5,σ=1)w\sim\mathcal{N}(\mu=3.5,\sigma=1), but clip to the range w∈w\in meters afterwards, in accordance to our empirical observations of the real-world distribution.

If the rendered synthetic crosswalk has overlap with any other real crosswalk above a threshold of IoU=0.05\text{IoU}=0.05, we continue to sample until we succeed. The crosswalk is rendered as a rectangle, bounded between two long edges both extending along the principal axis of the crosswalk. We use alternating parallel strips of white and gray to color the object. Crosswalks are deleted by simply not rendering them in the rasterized image.

J.2. Lane Geometry Perturbation Procedure

Our main observations from studying real-world map changes are that lane changes generally occur over a chain of lane segments, with combined length often over tens or hundreds of meters, although at times the combined length is far shorter. Accordingly, we use the directed lane graph to sample random connected sequences of lane segments, respecting valid successors. We then manipulate either the left or the right boundary only (not both) of this lane sequence.

When deleting lane boundaries, we sample only painted yellow or white lane boundary markings. When changing the color or structure of lane boundaries, we sample lane boundary markings of any color (including those that are implicit). When adding a bike lane, we sample a sequence of 5 lane segments. For marking deletion and changes to lane marking color and structure, we sample a sequence of length 3.

We render these boundaries as colored polylines; we use red for implicit boundaries, and yellow and white for lane markings of their respective color. Lane boundary markings are deleted by simply not rendering them in the rasterized image.

Bike lanes generally represent the rightmost lane in the United States. Accordingly, we synthesize a valid location for a new bike lane by iterating through the lane graph until there is no right neighbor; by dividing this rightmost lane into half, we can create two half-width lanes in place of one. We use solid white lines to represent their boundaries.

Appendix K: Datasheet for TbV

In this appendix, we answer the questions laid out in Datasheets for Datasets by Gebru et al. .

For what purpose was the dataset created? Was there a specific task in mind? Was there a specific gap that needed to be filled?

TbV was created to allow the community to improve the state of the art in machine learning tasks related to mapping, that are vital for self-driving.

To our knowledge, no prior datasets has ever been publicly released for HD map change detection. It is also one of the largest sensor datasets ever released, paired with HD maps, allowing for new exploration of the synergies between the sensor data and map data.

Who created this dataset (e.g., which team, research group) and on behalf of which entity (e.g., company, institution, organization)?

The TbV dataset was created by researchers who were employed at Argo AI.

The creation of this dataset was funded by Argo AI.

What do the instances that comprise the dataset represent (e.g., documents, photos, people, countries)? Are there multiple types of instances (e.g., movies, users, and ratings; people and interactions between them; nodes and edges)?

The core instances for TbV are brief “scenarios” or “logs” of that represent a continuous observation of a scene around a self-driving vehicle. On average, each scenario is 54 seconds in duration, although some capture as little as 4 seconds, while others last for up to 117 seconds.

Each scenario has an HD map representing lane boundaries, crosswalks, drivable area, etc. They also contain a raster map of ground height at 0.3 meter resolution.

How many instances are there in total (of each type, if appropriate)?

Does the dataset contain all possible instances or is it a sample (not necessarily random) of instances from a larger set? If the dataset is a sample, then what is the larger set? Is the sample representative of the larger set (e.g., geographic coverage)? If so, please describe how this representativeness was validated/verified. If it is not representative of the larger set, please describe why not (e.g., to cover a more diverse range of instances, because instances were withheld or unavailable).

The scenarios in the dataset are a sample of the set of observations made by a fleet of self-driving vehicles. The data is not uniformly sampled.

The “negative” instances in the dataset were chosen to include specific examples where an HD map has become out-of-date, due to real-world changes.

The “positive” instances in the dataset were chosen to include interesting behavior (e.g. cars making unexpected maneuvers), to contain interesting weather (e.g. rain and snow), and to be geographically diverse (spanning 6 cities – Pittsburgh, Detroit, Austin, Palo Alto, Miami, and Washington D.C.).

What data does each instance consist of? “Raw” data (e.g., unprocessed text or images) or features? In either case, please provide a description.

Each scenario has 20 fps video from 7 ring cameras, 20 fps video from two forward-facing stereo cameras, and 10 Hz LiDAR returns from two out-of-phase 32-beam LiDARs. The ring cameras are synchronized to fire when either LiDAR sweeps through their field of view. Each scenario contains vehicle pose over time and calibration data to relate the various sensors.

The HD map associated with each scenario contains polylines describing lanes, crosswalks, and drivable area. Lanes form a graph with predecessors and successors, e.g. a lane that splits can have two successors. Lanes have precisely localized lane boundaries that include paint type (e.g. double solid yellow). Drivable area, also described by a polygon, is the area where it is possible (but not necessarily legal) to drive without damaging the vehicle. It includes areas such as road shoulders.

Is there a label or target associated with each instance?

Yes. For the logs found in the train and synthetic validation splits, an up-to-date HD map serves as a label, as these are “positive” logs, where the map and sensor data are in agreement.

For the logs found in the “real” validation and test splits, 3d coordinates of polygons or polylines are manually annotated for areas where the map has changed, for lane paint and crosswalks, specifically.

In addition, the LiDAR depth estimates can act as ground truth for monocular depth estimation. The vehicle pose data could be considered ground truth labels for visual odometry. The evolving point cloud itself can be considered ground truth for point cloud forecasting.

Is any information missing from individual instances? If so, please provide a description, explaining why this information is missing (e.g., because it was unavailable). This does not include intentionally removed information, but might include, e.g., redacted text.

No. To our knowledge, all instances should be complete.

Are relationships between individual instances made explicit (e.g., users’ movie ratings, social network links)? If so, please describe how these relationships are made explicit.

Each instance of the dataset (a vehicle “log”) is disjoint. Each carries their own HD map for the region around their scenario. These HD maps may overlap spatially, though. For example, they may be captured at the same intersection, but separated in time by several months. If a user of the dataset wanted to recover the spatial relationship between scenarios, they could do so through our development kit.

Are there recommended data splits (e.g., training, development/validation, testing)? If so, please provide a description of these splits, explaining the rationale behind them.

We define splits of the TbV dataset. The train, validation, and test set include 799 / 111 / 133 logs each.

Are there any errors, sources of noise, or redundancies in the dataset?

Every sensor used in the dataset – ring cameras and lidar – has noise associated with it. Pixel intensities, lidar intensities, and lidar point 3D locations all have noise. Lidar points are also quantized to float16 which leads to roughly a centimeter of quantization error. 6 degree of freedom vehicle pose also has noise. The calibration specifying the relationship between sensors can be imperfect.

The HD map for each scenario can contain noise, both in terms of lane boundary locations and precise ground height.

Is the dataset self-contained, or does it link to or otherwise rely on external resources (e.g., websites, tweets, other datasets)? If it links to or relies on external resources, a) are there guarantees that they will exist, and remain constant, over time; b) are there official archival versions of the complete dataset (i.e., including the external resources as they existed at the time the dataset was created); c) are there any restrictions (e.g., licenses, fees) associated with any of the external resources that might apply to a future user? Please provide descriptions of all external resources and any restrictions associated with them, as well as links or other access points, as appropriate.

The data itself is self-hosted, and we will maintain public links to all previous versions of the dataset in case of updates.

Does the dataset contain data that might be considered confidential (e.g., data that is protected by legal privilege or by doctor-patient confidentiality, data that includes the content of individuals non-public communications)?

Does the dataset contain data that, if viewed directly, might be offensive, insulting, threatening, or might otherwise cause anxiety?

Yes, the dataset contains images and behaviors of thousands of people on public streets.

Does the dataset identify any subpopulations (e.g., by age, gender)?

Is it possible to identify individuals (i.e., one or more natural persons), either directly or indirectly (i.e., in combination with other data) from the dataset? If so, please describe how.

We do not believe so. Image data has been anonymized via blurring. Faces and license plates are obfuscated by replacing their corresponding bounding box with a 5×55\times 5 grid, where each grid cell is the average color of the original pixels in that grid cell. The anonymization is done manually. For example, a person sitting on their front porch 10 meters from the road would have their face obscured.

Does the dataset contain data that might be considered sensitive in any way (e.g., data that reveals racial or ethnic origins, sexual orientations, religious beliefs, political opinions or union memberships, or locations; financial or health data; biometric or genetic data; forms of government identification, such as social security numbers; criminal history)?

How was the data associated with each instance acquired? Was the data directly observable (e.g., raw text, movie ratings), reported by subjects (e.g., survey responses), or indirectly inferred/derived from other data (e.g., part-of-speech tags, model-based guesses for age or language)? If data was reported by subjects or indirectly inferred/derived from other data, was the data validated/verified? If so, please describe how.

The sensor data was directly acquired by a fleet of autonomous vehicles.

Over what timeframe was the data collected? Does this timeframe match the creation timeframe of the data associated with the instances (e.g., recent crawl of old news articles)? If not, please describe the timeframe in which the data associated with the instances was created.

The data was collected from May 2020 to March 2021.

What mechanisms or procedures were used to collect the data (e.g., hardware apparatus or sensor, manual human curation, software program, software API)? How were these mechanisms or procedures validated?

The Trust but Verify (TbV) data comes from Argo ‘Z1’ fleet vehicles. These vehicles use Velodyne lidars and traditional RGB cameras. All sensors are calibrated by Argo. HD maps are created and validated through a combination of computational tools and human annotations. Map change labels are created through human annotation.

If the dataset is a sample from a larger set, what was the sampling strategy (e.g., deterministic, probabilistic with specific sampling probabilities)?

The dataset scenarios were chosen from a larger set through manual review. The test set scenarios were selected to illustrate unambiguous map changes.

Who was involved in the data collection process (e.g., students, crowdworkers, contractors) and how were they compensated (e.g., how much were crowdworkers paid)?

Argo employees and Argo interns curated the data. Data collection and data annotation was done by Argo employees. Crowdworkers were not used.

Were any ethical review processes conducted (e.g., by an institutional review board)?

Did you collect the data from the individuals in question directly, or obtain it via third parties or other sources (e.g., websites)?

The data is collected from vehicles on public roads, not from a third party.

Were the individuals in question notified about the data collection? If so, please describe (or show with screenshots or other information) how notice was provided, and provide a link or other access point to, or otherwise reproduce, the exact language of the notification itself.

No, but the data collection was not hidden. The Argo fleet vehicles are well-marked and have obvious cameras and LiDAR sensors. The vehicles only capture data from public roads.

Did the individuals in question consent to the collection and use of their data? If so, please describe (or show with screenshots or other information) how consent was requested and provided, and provide a link or other access point to, or otherwise reproduce, the exact language to which the individuals consented.

No. People in the dataset were in public settings and their appearance has been anonymized. Drivers, pedestrians, and vulnerable road users are an intrinsic part of driving on public roads, so it is important that datasets contain people so that the community can develop more accurate perception systems.

If consent was obtained, were the consenting individuals provided with a mechanism to revoke their consent in the future or for certain uses? If so, please provide a description, as well as a link or other access point to the mechanism (if appropriate).

Has an analysis of the potential impact of the dataset and its use on data subjects (e.g., a data protection impact analysis) been conducted? If so, please provide a description of this analysis, including the outcomes, as well as a link or other access point to any supporting documentation.

Was any preprocessing/cleaning/labeling of the data done (e.g., discretization or bucketing, tokenization, part-of-speech tagging, SIFT feature extraction, removal of instances, processing of missing values)? If so, please provide a description. If not, you may skip the remainder of the questions in this section.

Yes. Images are reduced from their full resolution, and are JPEG compressed. 3D point locations are quantized to float16. Ground height maps are quantized to 0.3 meter resolution from their full resolution. HD map polygon vertex locations are quantized to 0.01 meter resolution.

Was the “raw” data saved in addition to the preprocessed/cleaned/labeled data (e.g., to support unanticipated future uses)? If so, please provide a link or other access point to the “raw” data.

Is the software used to preprocess/clean/label the instances available?

Has the dataset been used for any tasks already?

Yes, this manuscript benchmarks a novel HD map change detection method on the TbV dataset.

Is there a repository that links to any or all papers or systems that use the dataset?

Yes, at https://github.com/johnwlambert/tbv. We plan to add a leaderboard for the HD map change detection task using the test split of the TbV dataset.

What (other) tasks could the dataset be used for?

The TbV dataset could be used for research on visual odometry, lane detection, synthetic HD map generation, map automation, self-supervised learning, scene flow, point cloud forecasting, and more.

Is there anything about the composition of the dataset or the way it was collected and preprocessed/cleaned/labeled that might impact future uses? For example, is there anything that a future user might need to know to avoid uses that could result in unfair treatment of individuals or groups (e.g., stereotyping, quality of service issues) or other undesirable harms (e.g., financial harms, legal risks) If so, please provide a description. Is there anything a future user could do to mitigate these undesirable harms?

Are there tasks for which the dataset should not be used?

The dataset should not be used for tasks which depend on faithful appearance of faces or license plates since that data has been obfuscated. For example, running a face detector to try and estimate how often pedestrians use crosswalks will not result in meaningful data.

Will the dataset be distributed to third parties outside of the entity (e.g., company, institution, organization) on behalf of which the dataset was created?

Yes, the dataset is hosted on https://www.argoverse.org/. Our dataset requires no user registration for access. The dataset’s metadata page will include structured metadata.

In addition to long term hosting on Argoverse.org, the Creative Commons license enables rehosting by any repository. The authors will ensure that the dataset is accessible.

How will the dataset will be distributed (e.g., tarball on website, API, GitHub) Does the dataset have a digital object identifier (DOI)?

The TbV dataset is distributed as a series of tar.gz files. The files are broken up to make the process more robust to interruption (e.g. a single 1 TB file failing after 3 days would be frustrating) and to allow easier file manipulation (an end user might not have 1 TB free on a single drive, and if they do, they might not be able to decompress the entire file at once).

The dataset can be read with the Argoverse 2.0 API. See https://github.com/argoai/av2-api for details on usage.

The data is currently available for download, at the time of NeurIPS 2021.

Will the dataset be distributed under a copyright or other intellectual property (IP) license, and/or under applicable terms of use (ToU)? If so, please describe this license and/or ToU, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms or ToU, as well as any fees associated with these restrictions.

Yes, the dataset is released under the same Creative Commons license as Argoverse 1.0 (CC BY-NC-SA 4.0). The authors are responsible for the contents of the dataset and are responsible for any possible violation of rights.

Have any third parties imposed IP-based or other restrictions on the data associated with the instances?

Do any export controls or other regulatory restrictions apply to the dataset or to individual instances?

Who will be supporting/hosting/maintaining the dataset?

How can the owner/curator/manager of the dataset be contacted (e.g., email address)?

The TbV team will respond through the Github page https://github.com/johnwlambert/tbv/issues (where training code and pre-trained models have been made available). For privacy concerns, contact information may be found here: https://www.argoverse.org/about.html#privacy.

Will the dataset be updated (e.g., to correct labeling errors, add new instances, delete instances)? If so, please describe how often, by whom, and how updates will be communicated to users (e.g., mailing list, GitHub)?

It is possible that the TbV 1.0 Dataset will be updated to correct errors. Updates will be communicated on Github and through a mailing list we will create.

If the dataset relates to people, are there applicable limits on the retention of the data associated with the instances (e.g., were individuals in question told that their data would be retained for a fixed period of time and then deleted)? If so, please describe these limits and explain how they will be enforced.

Will older versions of the dataset continue to be supported/hosted/maintained? If so, please describe how. If not, please describe how its obsolescence will be communicated to users.

Yes. If we ever deprecate TbV 1.0, we will continue to host it, although we will declare it “deprecated.”

If others want to extend/augment/build on/contribute to the dataset, is there a mechanism for them to do so? If so, please provide a description. Will these contributions be validated/verified? If so, please describe how. If not, why not? Is there a process for communicating/distributing these contributions to other users? If so, please provide a description.

Yes. The Creative Commons license we use for TbV ensures that the community can do the same thing without needing Argo’s permission.

We do not have a mechanism for these contributions/additions to be incorporated back into the ‘base’ TbV Dataset. Our preference would generally be to keep the ‘base’ dataset as is, and to give credit to noteworthy additions by linking to them.

Environmental Impact Statement Amount of Compute Used: We estimate 5000 CPU hours and 3000 GPU hours for all of the data extraction, preparation and experiments.