SPG: Unsupervised Domain Adaptation for 3D Object Detection via Semantic Point Generation

Qiangeng Xu, Yin Zhou, Weiyue Wang, Charles R. Qi, Dragomir Anguelov

Introduction

A robust autonomous driving system requires its LiDAR-based detector to reliably handle different environmental conditions, e.g., geographic locations and weather conditions. While 3D detection has received increasing interest in recent years, most existing works have focused on the performance in a single domain, where training and test data are captured in similar conditions. It is still an open question how to generalize a 3D detector to different domains, where the environment varies significantly. In this paper, we address the domain gap caused by the deteriorating point cloud quality and aim to improve 3D object detection in the setting of unsupervised domain adaptation (UDA). We use the Waymo Domain Adaptation dataset to analyze the domain gap and introduce semantic point generation (SPG), a general approach to enhance the reliability of LiDAR detectors against domain shift. SPG is able to improve detection quality in both the target domain and the source domain and can be naturally combined with modern LiDAR-based detectors.

Waymo Open Dataset (OD) is mainly collected in California and Arizona, and Waymo Kirkland Dataset (Kirk) is collected in Kirkland. We consider OD as the source domain and Kirk as the target domain. To understand the possible domain gap, we take a PointPillars model trained on the OD training set and compare its 3D vehicle detection performance on OD validation set and those on Kirk validation set. We observe a drastic performance drop of 21.821.8 points in 3D average precision (AP) (see Table 1).

We first confirm that there is no significant difference in object size between two domains. Then by investigating the meta data in the datasets, we find that only 0.5%0.5\% of LiDAR frames in OD are collected under rainy weather, but almost all frames in Kirk share the rainy weather attribute. To rule out other factors, we extract all dry weather frames in Kirk training set and form a “Kirk Dry” dataset. Because the the rain drop changes the surface property of objects, there are twice amount of missing LiDAR points per frame in Kirk validation set than in OD or Kirk Dry (see Table 1). As a result, vehicles in Kirk receive around 27%27\% fewer LiDAR point observations than those in OD (see statistics and more details in the supplemental). In Figure 2, we visualize two range images from OD and Kirk, respectively. We can observe that in the rainy weather, a significant number of points are missing and the distribution of missing points is more irregular compared to the dry weather.

To conclude, the major domain gap between OD and Kirk is the deteriorating point cloud quality, which is caused by the rainy weather condition. In the target domain, we name this phenomenon as the “missing point” problem.

2 Previous Methods to Address the Domain Gap

Multiple studies propose to align the features across domains. Most of them focus on 2D tasks or object-level 3D tasks . Applying feature alignment requires a redesign of the model or loss of a detector. Our goal is to seek a general solution to benefit recently reported LiDAR-based detectors.

Another direction is to apply transformations to the data from one domain to match the data from another domain. A naive approach is to randomly down-sample the point cloud but this not only fails to satisfactorily simulate the pattern of missing points (Figure 2d) but also hurts the performance on the source domain. Another approach is to up-sample the point cloud in the target domain, which can increase point density around observed regions. However, those methods have a limited capability in recovering the 3D shape of very partially observed objects. Moreover, up-sampling the entire point cloud will lead to a significantly higher latency. A third approach is to leverage style transfer techniques: render point clouds as 2D pseudo images and enforce the renderings from different domains to be resemblant in style. However, these methods introduce an information bottleneck during rasterization and they are not applicable to modern point-based 3D detectors .

3 SPG for Closing the Domain Gap

The “missing point” problem deteriorates the point cloud quality and reduces the number of point observations, thus undermining the detection performance. To address this issue, we propose Semantic Point Generation (SPG). Our approach aims to learn the semantic information of the point cloud and performs foreground region prediction to identify voxels that are inside foreground objects. Based on the predicted foreground voxels, SPG generates points to recover the foreground regions. Since these points are discriminatively generated at foreground objects, we denote them by semantic points. These semantic points are merged with the original points into an augmented point cloud, which is then fed to a 3D detector.

The contributions of this paper are two-fold: 1. We present an in-depth analysis of unsupervised domain adaptation (UDA) for LiDAR 3D detectors across different geographic locations and weather conditions. Our study reveals that the rainy weather can severely deteriorate the quality of LiDAR point clouds and lead to drastic performance drop for modern detectors. 2. We propose semantic point generation (SPG). To our best knowledge, it is the first learning-based model that targets UDA for point cloud 3D detection. Specifically, SPG has the following merits:

SPG can generate semantic points that faithfully recover the foreground regions suffering from the “missing point” problem. SPG can significantly improve performance over poor-quality point clouds in the target domain while also benefiting source domain, for representative 3D detectors, including PointPillars and PV-RCNN .

SPG also improves the performance for the general 3D object detection task. We verify its effectiveness on KITTI for the aforementioned 3D detectors.

SPG is a general approach and can be easily combined with modern off-the-shelf LiDAR-based detectors.

Our approach is light-weight and efficient. Introducing less than 6%6\% additional points, SPG only adds a marginal complexity to a 3D detector.

Related Work

Unsupervised domain adaptation (UDA) aims to generalize a model to a novel (target) domain by using label information only from the source domain. The two domains are generally related, but there exists a distribution shift (domain gap). Most methods focus on learning aligned feature representations across domains. To reach this goal, proposes Maximum Mean Discrepancy (MMD) while proposes Transfer Component Analysis (TCA). designs a Joint Distribution Adaptation to close the distribution shift while utilize a shared Hilbert space. Without using explicit distance measures, deep learning models use adversarial training to get indistinguishable features between domains.

The object detection task is sensitive to local geometric features. hierarchically align the features between domains. Most of these works focus on UDA for 2D detection. With the current advances of unpaired style transfer methods , studies such as translate the image from source domain to target domain or vice versa.

Most of the UDA methods focus on 2D tasks, only a few studies explore the UDA in 3D. align the global and local features for object-level tasks. To reduce the sparsity, projects the point cloud to 2D view, while projects the point cloud to birds-eye view (BEV). creates a car model set and adapts their features to the detection object features. However, this study targets general car 3D detection on a single point cloud domain. is the first published study targeting UDA for 3D LiDAR detection. They identify the vehicle size as the domain gap between KITTI and other datasets. So they resize the vehicles in the data. In contrast, we identify the point cloud quality as the major domain gap between Waymo’s two datasets. We use a learning-based approach to close the domain gap.

2 Point Cloud Transformation

One way to improve point cloud quality is to suitably transform the point cloud. Studies of point cloud up-sampling can transfer a low density point cloud to a high density one. However, they need high density point cloud ground truth during training. These networks can densify the point cloud in the observed regions. But in our case, we also need to recover regions with no point observation, caused by “missing points”.

Point cloud completion networks aim to complete the point cloud. Specialized in object-level completion, these models assume a single object has been manually located and the input only consists of the points on this object. Therefore, these models do not fit our purpose of object detection. Point cloud style transfer models can transfer the color theme and the object-level geometric style for the point cloud. However, these models do not focus on preserving local details with high-fidelity. Therefore, their transformation cannot directly help 3D detection.

Semantic Point Generation

2 Model Structure

3 Foreground Region Recovery

The above pipeline supervises SPG to generate semantic points in the occupied voxels. However, it is also crucial to recover the empty voxels caused by the “missing points” problem. To generate semantic points in the empty areas, SPG employs two strategies:

“Hide and Predict”, which produces the “missing points” on the source domain during training and guides SPG to recover the foreground object shape in the empty space.

“Semantic Area Expansion”, which leverages the foreground/background voxel labels derived from the bounding boxes and encourages SPG to recover more unobserved foreground regions in each bounding box.

This strategy brings two benefits: 1. Hiding points region by region mimics the missing point pattern in the target domain; 2. The strategy naturally creates the training targets for semantic points in the empty space. Section 4.4 shows the effectiveness of this strategy. Here we set γ=25\gamma=25.

3.2 Semantic Area Expansion

In section 1.1, we find the poor point cloud quality leads to insufficient points on each object and substantially degrades the detection performance. To remedy this problem, we allow SPG to expand the generation area to the empty space. Figure 5 a and c show the examples of the generation area with and without the expansion, respectively.

Without the expansion, we can use the ground-truth knowledge of foreground points to supervise SPG only on the occupied voxels (Figure 5 b). However, with the expansion, there is no foreground point inside these empty voxels. Therefore, as shown in Figure 5 d, we design a supervision scheme as follows: 1. For both occupied and empty background voxels VobV_{o}^{b} and VebV_{e}^{b}, we impose negative supervision and set label yf=0y^{f}=0. 2. For the occupied foreground voxels VofV_{o}^{f}, we set yf=1y^{f}=1. 3. For the empty voxels inside a bounding box VefV_{e}^{f} , we set their foreground label yf=1y^{f}=1 and assign a weighting factor α\alpha, where α<1\alpha<1. 4. We only impose point features supervision ψ\psi at occupied foreground voxels VofV_{o}^{f}.

To investigate the effectiveness of the expansion, we train a model on the OD training set and evaluate it on the Kirk validation set. The expansion results in 510% more semantic points on foreground objects, which mitigates the “missing points” problem caused by environmental interference and occlusions. Figure 6 shows the generation results with and without the expansion. The supervision scheme encourages SPG to learn the extended shape of vehicle parts and enables SPG to fill in more foreground space with semantic points. We also conduct ablation studies (Section 4.4) to show the effectiveness of the proposed strategy.

4 Objectives

We use two loss functions, i.e., foreground area classification loss LclsL_{cls} and feature regression loss LregL_{reg}.

Please note that we are only interested in the LclsL_{cls} and LregL_{reg} on voxels inside the generation area. We find α=0.5\alpha=0.5 and β=2.0\beta=2.0 achieves the best result.

Experiments

In this section, we first evaluate the effectiveness of SPG as a general UDA approach for 3D detection, based on the Waymo Domain Adaptation Dataset . In addition, we show that SPG can also improve results for top-performing 3D detectors on the source domain. To demonstrate the wide applicability of SPG, we choose two representative detectors: 1) PointPillars , popular among industrial-grade autonomous driving systems; 2) PV-RCNN , a high performance LiDAR-based 3D detector . We perform two groups of model comparisons under the setting of unsupervised domain adaptation (UDA) and general 3D object detection: group 1, PointPillars vs. SPG + PointPillars; group 2, PV-RCNN vs. SPG + PV-RCNN. SPG can also be combined with range image-based detectors by applying ray casting to the generated points. However, we leave this as future work.

The Waymo Domain Adaptation dataset 1.0 consists of two sub datasets, the Waymo Open Dataset (OD) and the Waymo Kirkland Dataset (Kirk). OD provides 798 training segments of 158,361 LiDAR frames and 202 validation segments of 40,077 frames. Captured across California and Arizona, 99.40%99.40\% of its frames have dry weather. Kirk is a smaller dataset including 80 training segments of 15,797 frames and 20 validation segments of 3,933 frames. Captured in Kirkland, 97.99%97.99\% its LiDAR frames have rainy weather. To examine a detector’s reliability when entering a new environment, we conduct UDA experiments without using the data in Kirk during training.

KITTI contains 7481 training samples and 7518 testing samples. Following , we divide the training data into a train split and a val split containing 3721 and 3769 LiDAR frames, respectively.

We implement PointPillars following and use the PV-RCNN code provided by (the training settings on OD 1.0 are obtained via direct communication with the author). On the Waymo Domain Adaptation Dataset , we set the voxel dimensions to (0.32m, 0.32m, 0.4m) for PointPillars and (0.2m, 0.2m, 0.3m) for PV-RCNN. On KITTI, we set the voxel dimensions to (0.16m, 0.16m, 0.2m) and (0.2m, 0.2m, 0.3m) for PointPillars and PV-RCNN, respectively. By default, the generation area includes voxels within 6 steps of any occupied voxel. After probability thresholding, we preserve up to 80008000 semantic points for the Waymo Domain Adaptation Dataset and 60006000 for KITTI.

1 Evaluation on the Waymo Open Dataset

We perform two groups of model comparisons by training them on the OD training set and evaluating them on both the OD validation set and the Kirk validation set.

The Kirk 1.0 validation set only provides the evaluation labels for the vehicle and the pedestrian classes. We use the official evaluation tool released by . The IoU thresholds for vehicles and pedestrians are 0.7 and 0.5. In Table 2 we report both 3D and BEV AP on two difficulty levels. More results with distance breakdown are shown in the supplemental material.

On Kirk, we observe that SPG brings remarkable improvements over both detectors across all object types. Averaged over two difficulty levels, SPG improves PointPillars on Kirk vehicle 3D AP by 6.7%6.7\% and BEV AP by 8.8%8.8\%. For PV-RCNN, SPG improves Kirk pedestrian 3D AP by 5.6%5.6\% and BEV AP by 5.7%5.7\%.

Unlike most UDA methods that only optimize the performance on the target domain, SPG also consistently improves the results on the source domain. Averaged across both difficulty levels, SPG improves OD vehicle 3D AP for PointPillars by 5.4%5.4\% and improves OD pedestrian 3D AP for PV-RCNN by 1.6%1.6\%.

We compare SPG with alternative strategies that also target the deteriorating point cloud quality. We employ PointPillars as the baseline and choose LEVEL_1 vehicle 3D AP as the main metric on the Kirk validation set, during UDA. Three strategies are implemented: 1. RndDrop, where we randomly drop 17%17\% of the points in the source domain during training. This dropout ratio is chosen for the number of points in the source and target domain to match (see Table 1). 2. K-frames, where we use KK consecutive historical frames in both the source domain and the target domain. The points in the first K−1K-1 are transformed into the last frame according to the ground-truth ego-motion, so that the last frame has KK times the number of points. 3. Adversarial Domain Adaptation (ADA), where we follow and add a domain classification loss on the pillar features of PointPillars.

As shown in Table 3, although “RndDrop” enforces the quantity of missing points in the source domain to match with that in the target domain, the pattern of missing points still differs from the reality (see Figure 2), which limits the improvement to only 0.80%0.80\% in 3D AP. To remedy the “missing points” problem, “3-frames” contains real points from 3 frames and “5-frames” contains points from 5 frames. With around 800K points per scene, “5-frames” significantly improves the single-frame baseline. However, aggregating multiple frames inevitably increases the memory usage and the processing time. ADA improves 3D AP to 36.3436.34 on the target domain, but we observe an AP drop of 1.521.52 in the source domain. Remarkably, SPG can outperform “5-frames”, by adding only 8000 semantic points, which is less than 6%6\% of the points in a single frame.

2 Evaluation on the KITTI Dataset

In this section, we show besides the usefulness in UDA (Sec. 4.1) the proposed SPG can also boost performance in another popular 3D detection benchmark (i.e. KITTI ). We follow the training and evaluation protocols in .

As shown in Table 4, SPG significantly improves PV-RCNN on Car 3D detection. As of Mar. 3rd, 2021, our method ranks the 1st on KITTI car 3D detection among all published methods (4th among all submitted approaches). Moreover, SPG demonstrates strong robustness in detecting hard objects (truncation up to 50%). Specifically, SPG surpasses all submitted methods on the hard category by a big margin and achieves the highest overall 3D AP of 83.84%83.84\% (averaged over Easy, Mod. and Hard).

We summarize the results in Table 5. We train each group of models using the recommended settings of baseline detectors .

SPG remarkably improves both PointPillars and PV-RCNN on all object types and difficulty levels. Specifically, for PointPillars, SPG improves the 3D AP of car detection by 2.02%2.02\%, 2.97%2.97\%, 3.67%3.67\% on easy, moderate, and hard levels, respectively. For PV-RCNN, SPG improves the 3D AP of pedestrian detection by 5.40%5.40\%, 5.13%5.13\%, 4.48%4.48\% on easy, moderate and hard levels, respectively.

3 Model Efficiency

We evaluate the efficiency of SPG on the KITTI val split (Table 6). SPG contains 0.390.39 million parameters while adding less than 1717 milliseconds latency to the detectors. This indicates that SPG is highly efficient for industrial-grade deployment on a stringent computation budget.

4 Ablation Studies

In Table 8, we show the effect of choosing different thresholds during probability thresholding. While a higher PthreshP_{thresh} only keeps semantic points with high foreground probability, a lower PthreshP_{thresh} admits more points, but may introduce points to the background. We find the threshold of 0.50.5 achieves the best results.

Conclusions

In this paper, we investigate unsupervised domain adaptation for LiDAR-based 3D detectors across different geographic locations and weather conditions. We observe that rainy weather can severely deteriorate the point cloud quality and cause drastic performance drop for modern 3D detectors, based on the Waymo Domain Adaptation dataset. The proposed SPG method addresses this issue as a novel unsupervised domain adaptation (UDA) task without using any training data from the new domain. This setting allows us to rigorously test 3D detectors against real-world challenges autonomous vehicles may experience due to diverse conditions (e.g., different levels of fog/rain/snow beyond what one may effectively train for) during the trip.

Utilizing two strategies “Hide and Predict” and “Semantic Area Generation”, SPG generates semantic points to recover the shape of foreground objects with a negligible overhead (only adding 6%6\% extra points) and can be conveniently integrated with modern LiDAR-based detectors. We test SPG with two detectors: PointPillars and PV-RCNN. For unsupervised domain adaptation, SPG achieves significant performance gains on the challenging target domain. On Waymo Open dataset and KITTI, SPG also consistently benefits detection quality on the source domain.

Acknowledgement

We would like to thank Boqing Gong for the helpful discussions. We also thank Jingwei Ji for the careful proofreading.

References

Appendix A Statistics of the Waymo Domain Adaptation Dataset

We collect the statistics about the average number of points in a vehicle bounding box across different ranges. The range value is calculated as the euclidean distance between the LiDAR sensor and the center of a bounding box. We investigate four sets of point clouds:

The OD Validation set, in which 99.5%99.5\% of the frames are collected in the dry weather.

The Kirk Dry set, which consists of all the frames with the dry weather condition from the Kirk training set.

The Kirk Training Rainy set, which consists of all the frames with the rainy weather condition from the Kirk training set.

The Kirk Validation set, in which all the frames are collected in the rainy weather.

As shown in Figure 7, the point clouds with similar weather conditions share similar numbers of points per object, even though they are collected at different locations. Specifically, the vehicle objects of the two “dry datasets”, i.e., the Kirk Dry set and the OD Validation set, have similar numbers of points across all ranges. The vehicle objects of the two “rainy datasets” i.e., the Kirk Training Rainy set and the Kirk Validation set, share similar statistics.

In addition, the point clouds captured in the dry weather (the OD Validation set and the Kirk Dry set) have more points on each object than those collected in the rainy weather (the Kirk Training Rainy set and the Kirk Validation set). Please note that we have applied log10log_{10} to the number of points for better visualization. The difference in the number of points is substantial between two weather conditions across all ranges.

Appendix B The Robustness of the Foreground Voxel Classifier

In order to generalize detectors to different domains, it is crucial to correctly classify foreground voxels so that semantic points can be reliably generated. Table 9 lists the evaluation results of the foreground voxel classifier.

Appendix C Dropout Rate of the RndDrop Method

In the experiment section, we implement a baseline method RndDrop, where we randomly drop out 17%17\% of points for point clouds from the source domain during training. This dropout ratio is chosen to match the ratio of missing points in the target domain. We calculate (N‾src−N‾tgt)/N‾src=17%(\overline{N}_{src}-\overline{N}_{tgt})/\overline{N}_{src}=17\%, where N‾src=121.2K\overline{N}_{src}=121.2K is the average number of points per scene in the source domain and N‾tgt=100.4K\overline{N}_{tgt}=100.4K is the average number of points per scene in the target domain.

Appendix D More Results on the Waymo Domain Adaptation Dataset

The evaluation tool provides the average precision for three distance-based breakdowns: 0 to 30 meters, 30 to 50 meters, and beyond 50 meters. The AP is calculated using 100 recall thresholds.

We perform two groups of model comparisons in the setting of UDA: Group 1. PointPillars vs. SPG + PointPillars; Group 2. PV-RCNN vs. SPG + PV-RCNN. We train all models on the OD training set and evaluate them on both the OD validation set and the Kirk validation set. Table 10 and 11 show the comparisons on vehicle 3D AP and vehicle BEV AP, respectively. Table 12 and Table 13 show the comparisons in pedestrian 3D AP and pedestrian BEV AP, respectively. In most cases, SPG improves the detection performance across all ranges for both vehicles and pedestrians.

Appendix E More Results on KITTI

We provide more 3D object detection results on KITTI. There are two commonly used metric standards for evaluating the detection performance: 1) R11, where the AP is evaluated with 11 recall positions; 2) R40, where the AP is evaluated with 40 recall positions. In addition to the improvement on car and pedestrian detection, SPG also significantly boosts the performance in cyclist detection. Based on R11, Table 14 and Table 15 show the results in 3D AP and BEV AP for three object types, respectively. Based on R40, Table 16 and Table 17 show the results in 3D AP and BEV AP for three object types, respectively.

We show more comparisons on the KITTI test set in Table 18.

Appendix F More Visualization of Semantic Point Generation

In Figure 9, we illustrate more augmented point clouds, where the raw points are rendered in the grey color and the generated semantic points are highlighted in red.