2DPASS: 2D Priors Assisted Semantic Segmentation on LiDAR Point Clouds
Xu Yan, Jiantao Gao, Chaoda Zheng, Chao Zheng, Ruimao Zhang, Shenghui Cui, Zhen Li
Introduction
Semantic segmentation plays a crucial role in large-scale outdoor scene understanding, which has broad applications in autonomous driving and robotics . In the past few years, the research community has devoted significant effort to understanding natural scenes using either camera images or LiDAR point clouds as the input. However, these single-modal methods inevitably face challenges in complex environments due to the inherent limitations of the input sensors. Concretely, cameras provide dense color information and fine-grained texture, but they are ambiguous in depth sensing and unreliable in low light conditions. In contrast, LiDARs robustly offer accurate and wide-ranging depth information regardless of lighting variances but only capture sparse and textureless data. Since cameras and LiDARs complement each other, it is better to perceive the surrounding with both sensors.
Recently, many commercial cars have been equipped with both cameras and LiDARs. This excites the research community to improve the semantic segmentation by fusing the information from two complementary sensors . These approaches first establish the mapping between 3D points and 2D pixels by projecting the point clouds onto the image planes using the sensor calibrations. Based on the point-to-pixel mapping, the models fuse the corresponding image features into the point features, which are further processed to obtain the final semantic scores. Despite the improvements, fusion-based methods have the following unavoidable limitations: 1) Due to the difference of FOVs (field of views) between cameras and LiDARs, the point-to-pixel mapping cannot be established for points that are out of the image planes. Typically, the FOVs of LiDAR and cameras only overlap in a small portion (see Fig. 1), which significantly limits the application of fusion-based methods. 2) Fusion-based methods consume more computational resources since they process both images and point clouds (through multitask or cascade manners) at runtime, which introduces a great burden on real-time applications.
To address the above two issues, we focus on improving semantic segmentation by leveraging both images and point clouds through an effective design in this work. Considering the sensors are moving in the scenes, the non-overlap part of the 360-degree LiDAR point clouds corresponding to image in the same time-stamp (see the gray region of the right part in Fig. 1) can be covered by images from other time-stamp. Besides, the dense and structural information of images provides useful regularization for both seen and unseen point cloud regions. Based on these observations, we propose a “model-independent” training scheme, namely 2D Priors Assisted Semantic Segmentation (2DPASS), to enhance the representation learning of any 3D semantic segmentation networks with minor structure modification. In practice, on the one hand, for above-mentioned non-overlap regions, 2DPASS takes pure point clouds as the inputs to train the segmentation model. On the other hand, for subregions with well-aligned point-to-pixel mappings, 2DPASS adopts an auxiliary multi-modal fusion to aggregate image and point features in each scale, and then aligns the 3D predictions with the fusion predictions. Unlike previous cross-modal alignment apt to contaminate the modal-specific information, we design a multi-scale fusion-to-single knowledge distillation (MSFSKD) strategy to transfer extra knowledge to the 3D model as well as retaining its modal-specific ability. Compared with fusion-based methods, our solution has the following preferable properties: 1) Generality: It can be easily integrated with any 3D segmentation model with minor structural modification; 2) Flexibility: The fusion module is only used during the training to enhance the 3D network. After training, the enhanced 3D model can be deployed without image inputs. 3) Effectively: Even with only a small section of overlapped multi-modality data, our method can significantly boost the performance. As a result, we evaluate 2DPASS with a simple yet strong baseline implemented with sparse convolutions . The experiments show 2DPASS brings noticeable improvements even over this strong baseline. Equipped with 2DPASS using multi-modal data, our model achieves the top-1 results on the single and multiple-scan leaderboards of SemanticKITTI . The state-of-the-art results on the NuScenes dataset further confirm the generality of our method.
In general, the main contributions are summarized as follows.
We propose 2D Priors Assisted Semantic Segmentation (2DPASS) that assists 3D LiDAR semantic segmentation with 2D priors from cameras. To the best of our knowledge, 2DPASS is the first method that distills multi-modal knowledge to single point cloud modality for semantic segmentation.
Equipped with the proposed multi-scale fusion-to-single knowledge distillation (MSFSKS) strategy, 2DPASS achieves the significant performance gains on SemanticKITTI and NuScenes benchmarks, ranking the 1st on single and multiple tracks of SemanticKITTI.
Related Work
Camera-Based Methods. Camera-based semantic segmentation aims to predict the pixel-wise labels for input 2D images. FCN is the pioneer in semantic segmentation, which proposes an end-to-end fully convolutional architecture based on image classification networks. Recent works have achieved significant improvements via exploring multi-scale features learning , dilated convolution , and attention mechanisms. However, camera-only methods are ambiguous in depth sensing and not robust in low light conditions.
LiDAR-Based Methods. The LiDAR data is generally represented as point clouds. There are several mainstreams to process point clouds with different representations. 1) Point-based methods approximate a permutation-invariant set function using a per-point Multi-Layer Perceptron (MLP). PointNet is the pioneer in this field. Later on, many studies design point-wise MLP , adaptive weight and pseudo grid based methods to extract local features of point clouds or exploit nonlocal operators to learn long distance dependency. However, point-based methods are not efficient in the LiDAR scenario since their sampling and grouping algorithms are generally time-consuming. 2) Projection-based methods are very efficient approaches for LiDAR point clouds. They project point clouds onto 2D pixels so that traditional CNN can play a normal role. Previous works project all points scanned by the rotating LiDAR onto 2D images by plane projection , spherical projection or both . However, the projection inevitably causes information loss. And the projection-based methods currently meet the bottleneck of the segmentation accuracy. 3) Most recent works adopt voxel-based frameworks since they balance the efficiency and effectiveness, where sparse convolution (SparseConv) are most commonly utilized. Compared to traditional voxel-based methods (i.e., 3DCNN) directly transforming all points into the 3D voxel grids, SparseConv only stores non-empty voxels in a Hash table and conducts convolution operations only on these non-empty voxels in a more efficient way. Recently, many studies have used SparseConv to design more powerful network architectures. Cylinder3D changes original grid voxels to cylinder ones and designs an asymmetrical network to boost the performance. AF2-S3Net applies multiple branches with different kernel sizes, aggregating multi-scale features via an attention mechanism. 4) Very recently, there is a trend of exploiting multi-representation fusion methods. These methods combine multiple representations above (i.e., points, projection images, and voxels) and design feature fusion among different branches. Tang et.al. combines point-wise MLPs in each sparse convolution block to learn a point-voxel representation and uses NAS to search for a more efficient architecture. RPVNet proposes range-point-voxel fusion network to utilizes information from three representations. Nevertheless, these methods only take sparse and textureless LiDAR point clouds as inputs, thus appearance and texture in the camera images have not been fully utilized.
2 Multi-Sensor Methods
Multi-sensor methods attempt to fuse information from two complementary sensors and leverage the benefits of both camera and LiDAR . RGBAL converts RGB images to a polar-grid mapping representation and designs early and mid-level fusion strategies. PointPainting exploits the segmentation logits of images and projects them to the LiDAR space by bird’s-eye projection or spherical projection for LiDAR network performance improvement. Recently, PMF exploits a collaborative fusion of two modalities in camera coordinates. However, these methods require multi-sensor inputs in both training and inference phases. Moreover, the paired multi-modality data is usually computation-intensive and unavailable in practical application.
3 Cross-modal Knowledge Transfer
Knowledge distillation was initially proposed for compressing the large teacher network to a small student one . Over the past few years, several subsequent studies enhanced knowledge transferring through matching feature representations in different manners . For instance, aligning attention maps and Jacobean matrixes were independently applied. With the development of multi-modal computer vision, recent research apply knowledge distillation to transfer priors across different modalities, e.g., exploiting extra 2D images in the training phase and improving the performance in the inference . Specifically, introduces the 2D-assisted pre-training, inflates the kernels of 2D convolution to the 3D ones, and applies well-designed teacher-student framework. Inspired but different from the above, we transfer 2D knowledge through a multi-scale fusion-to-single manner, which additionally takes care of the modal-specific knowledge.
Method
This paper focuses on improving the LiDAR point cloud semantic segmentation, which aims to assign the semantic label to each point. To handle difficulties in large-scale outdoor LiDAR point clouds, i.e., sparsity, varying density, and lack of texture, we introduce the strong regularization and priors from 2D camera images through a fusion-to-single knowledge transferring.
The workflow of our 2D Priors Assisted Semantic Segmentation (2DPASS) is shown in Fig. 2. Since the camera images are pretty large (e.g., ), sending the original ones to our multi-modal pipeline is intractable. Therefore, we randomly sample a small patch () from the original camera image as the 2D input , accelerating the training processing without performance drop. Then the cropped image patch and LiDAR point cloud independently pass through independent 2D and 3D encoders, where multi-scale features from the two backbones are extracted in parallel. Afterwards, multi-scale fusion-to-single knowledge distillation (MSFSKD) is conducted to enhance the 3D network using multi-modal features, i.e., fully utilizing texture and color-aware 2D priors as well as retaining the original 3D-specific knowledge. Finally, all the 2D and 3D features at each scale are used to generate semantic segmentation predictions, which are supervised by pure 3D labels. During inference, the 2D-related branch can be discarded, which effectively prevents extra computational burden in real application compared with fusion-based approaches.
2 Modal-Specific Architectures
Multi-Scale Feature Encoders. As shown in Fig. 2, we use two different networks to independently encode multi-scale features from 2D image and 3D point cloud. We apply ResNet34 encoder with 2D convolution as the 2D network. For the 3D network, we adopt sparse convolution to construct the 3D network. One merit of sparse convolution lies in the sparsity, with which the convolution operation only considers the non-empty voxels. Specifically, we design a hierarchical point-voxel encoder as that used in the decoder of , and adopt the ResNet bottleneck in each scale while replacing the ReLU with Leaky ReLU . In both network, we extract feature maps from different scales, obtaining the 2D and 3D features, i.e., and .
Prediction Decoders. After processing the features from images and point clouds at each scale, two modal-specific prediction decoders are independently applied to restore the down-sampled feature maps to their original sizes.
For the 2D network, we adopt FCN decoder to up-sample the features from each encoder layer. Specifically, the feature map from the -th decoder layer can be gained by up-sampling the feature map from the -th encoder layer, where all the up-sampled feature maps will be merged through element-wise addition. Finally, the semantic segmentation of the 2D network is obtained by passing the fused feature map through a linear classifier.
For the 3D network, we do not adopt the U-Net decoder used in previous methods . In contrast, we up-sample the features from different scales to the original size and concatenate them together before feeding them into the classifier. We find out that such a structure can better learn hierarchical information while gaining the prediction in a more efficient way.
3 Point-to-Pixel Correspondence
Since the 2D features and 3D features are generally represented as pixels and points, respectively, it is difficult to directly transfer information between two modalities. In this section, we aim to generate paired features of two modalities for further knowledge distillation, using the point-to-pixel correspondence. The details of paired feature generation in two modalities are demonstrated in Fig. 3.
After the projection, the point-to-pixel mapping is represented as
3D Features. The process of 3D features is relatively straightforward (as shown in Fig. 3 (b)). Specifically, for the point cloud , we obtain a point-to-voxel mapping in the -th layer through
2D Ground Truths. Considering only 2D images is provided, the 2D ground-truths are obtained by projecting the 3D point labels to the corresponding image plane using above point-to-pixel mapping. Afterwards, the projected 2D ground truths can work as the supervision for the 2D branch.
Features Correspondence. Since both 2D and 3D feature use the same point-to-pixel mapping, 2D features and 3D features in arbitrary -th layer have the same number of point and point-to-pixel correspondence.
4 Multi-Scale Fusion-to-Single Knowledge Distillation (MSFSKD)
As the key of 2DPASS, MSFSKD aims at improving the 3D representation in each scale using auxiliary 2D priors through a fusion-then-distillation manner. The knowledge distillation (KD) design of MSFSKD is partially inspired by . However, conducts KD in a naive cross-modal manner, i.e., simply aligning the outputs from two sets of single modal features (i.e. either 2D or 3D), which inevitably pushes the features from two modals to their overlapped space. Therefore, such a manner actually discards the modal-specific information, which is crucial in multi-sensor segmentation. Although this issue can be relieved by introducing extra segmentation heads , it is inherent for the cross-modal distillation, resulting in biased predictions. To this end, we propose multi-scale fusion-to-single knowledge distillation (MSFSKD) module as shown in Fig. 4, which first fuses features of both images and point clouds and then conducts unidirectional alignment between the fused and the point cloud features. In our fusion-then-distillation manner, the fusion well retains the complete information from multi-modal data. Besides, the unidirectional alignment ensures boosted point cloud features from fusion without losing modal-specific information.
Modality Fusion. For each scale, considering the 2D and 3D feature gaps owing to different backbones, it is ineffective to directly fuse the raw 3D features into their 2D counterparts . Thus, we firstly transform to through a “2D learner” MLP, which struggles to narrow the feature gap. Afterwards, the not only flows into the subsequent concatenation with 2D features to gain the fused features through another MLP, but also goes back into the original 3D features via a skip connection to yield enhanced 3D features . Besides, similar to attention mechanism, the final enhanced fused features is obtained by:
where denotes Sigmoid activation function.
Modality-Preserving KD. Although the is generated from pure 3D features, it is influenced by the segmentation loss of the 2D decoder as well, which takes enhanced fused feature as inputs. Acting like a residual between fused and point features, the 2D learner feature well prevents the distillation from contaminating the modal-specific information in , achieving a Modality-Preserving KD. Finally, two independent classifiers (fully-connected layers) are respectively applied on top of and to obtain the semantic scores and . We choose KL divergence as the distillation loss as follows:
Through such an implementation, it enforces the uni-directional distillation by pushing closer to .
By taking such a knowledge distillation scheme, there are several advantages in our framework: 1) The 2D learner and the fusion-to-single distillation provides rich texture information and structural regularization to enhance the 3D feature learning without losing any modal-specific information in 3D. 2) The fusion branch is only adopted in the training phase. Therefore, the enhanced model can almost run without extra computational cost during the inference.
Experiments
Datasets. We extensively evaluate 2DPASS on two large-scale outdoor benchmarks: SemanticKITTI and Nuscenes . SemanticKITTI provides dense semantic annotations for each individual scan of sequences 00-10 in KITTI dataset . According to the official setting, sequence 08 is the validation split, while the remaining are the train split. SemanticKITTI uses sequences 11-21 in KITTI as the test set, whose labels are held on for blind online testinghttps://competitions.codalab.org/competitions/20331. NuScenes contains 1000 scenes which show a great diversity in inner cities traffic and weather conditions. It officially divides the data into 700/150/150 scenes for train/val/test. Similar to SemanticKITTI, the test set of NuScenes is used for online benchmarkinghttps://eval.ai/web/challenges/challenge-page/720/leaderboard/1967. For 2D sensors, KITTI has only two front-view cameras, while NuScenes has six cameras covering the full 360° fields of view.
Evaluation Metrics. We evaluate methods mainly using mean intersection over union (mIoU), which is defined as the average IoU over all classes. Additionally, we report the overall accuracy (Acc)/ frequency-weighted IOU (FwIOU) provided by the online leaderboard of two benchmarks. FwIoU is similar to mIoU except that each IoU is weighted by the point-level frequency of its class.
Network Setup. We apply ResNet34 encoder with 2D convolution as the 2D network, where features after each down-sampling layers are extracted to generate 2D features. The 3D encoder is a modified SPVCNN (voxel size 0.1) with fewer parameters, whose hidden dimensions are 64 for SemanticKITTI and 128 for NuScenes to speed up the network. The number of layers for MSFSKD is set to 4 and 6 for SemanticKITTI and NuScenes, respectively. In each scale of knowledge distillation, 2D and 3D features are reduced to 64 dimensions through deconvolution or MLPs. Similarly, the hidden size of MLPs and 2D learner in MSFSKD are identically 64.
Training and Inference Details. We employ the cross-entropy and Lovasz losses as for semantic segmentation. For the knowledge distillation, we set the proportion of segmentation loss and KL divergence as . Test-time augmentation is applied during the inference. Training details will be introduced in supplementary material.
2 Benchmark Results
SemanticKITTI. SemanticKITTI evaluates segmentation performance using two settings: single scan and multiple scans. For methods using a single scan as input, moving and non-moving are mapped to a single class. While methods using multiple scans as inputs should distinguish between moving and non-moving objects, which is more challenging. All the reported results are from the official blind test competition website of SemanticKITTI.
Tab. 1 shows our performance under the single scan setting. Our baseline without 2DPASS already performs on par with a strong model Cylinder3D while runs at a faster speed. Even so, the application of 2DPASS still brings a significant improvement over the baseline. Thanks to the auxiliary knowledge distillation, 2DPASS does not put any extra burden on the original model and thus does not sacrifice the running speed of the baseline. Overall, 2DPASS achieves the best result in terms of mIoU and running speed, outperforming the state-of-the-art (i.e., (AF)2-S3Net ) by 2.1%. The visualization results on SemanticKITTI single scan are shown in Fig. 5.
Tab. 2 reports the results under the multiple scans setting. The mIoU and overall accuracy are calculated over all 25 classes. Due to the limited space, we only report the per-class IOUs for dynamic objects with non-moving/moving properties. Under this challenge setting, 2DPASS surprisingly surpasses previous approaches with even larger margins, i.e., achieving better mIoU (5.5% improvement over (AF)2-S3Net ) and overall accuracy.
NuScenes. The results on NuScenes are reported in Tab. 3, where 2DPASS achieves the 1st place as well. Note that we only include published works in Tab. 3 and the results are directly taken from the official leaderboard of NuScenes, where our model also ranks the 3rd place with slight disadvantage when considering unpublished works. Besides surpassing all single-modal methods, 2DPASS surprisingly outperforms those fusion-based approaches (the last two rows in Tab. 3). Note that NuScenes provides images covering the whole FOV of the LiDAR, and fusion-based approaches achieve such results by using both point clouds and image features during the inference. In contrast, our method only takes point clouds as input.
3 Comprehensive Analysis
Comparing with Other Knowledge Distillation. To further verify the effectiveness of our fusion-to-single knowledge distillation paradigm upon common teach-student architecture and other cross-modal manners, we compare 2DPASS with typical approaches of knowledge transfer in Tab. 5, where we utilize these methods in each scale for fair comparison. Among all the methods, Hinton et.al. , Huang et.al. and Yang et.al. are pure knowledge distillation designs, where the former is the pioneer for the research field and the latter is newly proposed. As shown in the Tab. 5, pure knowledge distillation manners cannot be directly adopted on the LiDAR semantic segmentation, and their improvement upon the baseline model is limited. Recently, adopts cross-modal feature alignment technique in the task of domain adaptation on semantic segmentation. However, their improvement is still marginal. To the end, in the Tab. 5, 2DPASS significantly performs better, which illustrates the effectiveness of our multi-scale fusion-to-single knowledge distillation (MSFSKD).
Design Analysis of MSFSKD. Tab. 5 demonstrates the ablation study on SemanticKITTI validation set. As shown in the table, our baseline only achieves a lower result of 65.58 mIoU. Note that simply using feature alignment between two modalities cannot effectively improve the result, where the metric of mIoU will be only increased to 66.34. After using 2D-3D fusion in each knowledge distillation scale, there is a significant improvement to 69.13. This improvement mainly comes from the knowledge provided by the stronger fusion prediction. Finally, we find out that 2D learner design can slightly improve the performance by about 0.2%. Note that the results on SemanticKITTI validation set is lower than that on benchmark since small object category (i.e., motocyclist) only occupies a small proportion.
Distance-based Evaluation. We investigate how segmentation is affected by distance of the points to the ego-vehicle, and compare 2DPASS, current state-of-the-art and the baseline on the SemanticKITTI validation set. Fig. 6 (a) illustrates the mIoU of 2DPASS as opposed to the baseline and (AF)2-S3Net. The results of all the methods get worse by increasing the distance since points are relatively sparse in the long distance. 2DPASS improves the performance greatly within 10, i.e., from 61.2 to 89.1, which is the best distance for the camera to capture objects’ color and texture. There is also a significant improvement upon (AF)2-S3Net within this distance, i.e., 84.4 v.s. 89.1.
Generality. We show our 2DPASS can be a “model-independent” training scheme that boosts the performance of other networks. We additionally trained two open-sourced baselines, i.e., MinkowskiNet and SPVCNN implemented in with 2DPASS. During the experiment, we keep all the setups the same except for the 2D-related components. As shown in Fig. 6 (b), 2DPASS improves the former one from 63.1 to 66.2 and the latter from 63.8 to 66.9. These results sufficiently demonstrate the effectiveness and generality of 2DPASS.
Conclusion
This work proposes the 2D Priors Assisted Semantic Segmentation (2DPASS), a general training scheme, to boost the performance of LiDAR point cloud semantic segmentation via 2D prior-related knowledge distillation. By leveraging an auxiliary modal fusion and knowledge distillation in a multi-scale manner, 2DPASS acquires richer semantic and structural information from the multi-modal data, effectively enhancing the performance of a pure 3D network. Eventually, it achieves the state-of-the-arts on two large-scale benchmarks (i.e., SemanticKITTI and NuScenes). We believe that our work can be applied to a wider range of other scenarios in the future, such as 3D detection and tracking.
Acknowledgment. This work was supported in part by NSFC-Youth 61902335, by the Basic Research Project No. HZQB-KCZYZ-2021067 of Hetao Shenzhen HK S&T Cooperation Zone, by the National Key R&D Program of China with grant No.2018YFB1800800, by Shenzhen Outstanding Talents Training Fund, by Guangdong Research Project No. 2017ZT07X152 and No. 2019CX01X104, by the Guangdong Provincial Key Laboratory of Future Networks of Intelligence (Grant No. 2022B1212010001), by the NSFC 61931024&8192 2046, by NSFC-Youth 62106154, by zelixir biotechnology company Fund, by Tencent Open Fund, and by ITSO at CUHKSZ.
References
A Training and Inference Details
For the 3D input, we utilize the widely used data augmentation strategy for semantic segmentation, including global scaling with a random scaling factor sampled from [0.95, 1.05], and global rotation around the Z axis with a random angle. For the 2D input, we employ horizontal flipping and color jitter. Each 2D image is cropped to the size 480 320 (width height) for faster training. The 2DPASS is trained in an end-to-end manner with the SGD optimizer. For the SemanticKITTI validation set, our model was trained with batch size 8 and learning rate 0.24 for 64 epochs, which is kept the same as SPVCNN for fair comparison. For the SemanticKITTI online benchmark, we conduct instance CutMix as , and fine-tune the last checkpoint with additional 48 epochs. As for the NuScenes dataset, we trained the model with batch size 16 for 80 epochs since the number of points per scene in NuScenes is generally smaller. During the inference, following , we apply the voting test-time augmentation, i.e., rotating the input scene with 12 angles around the Z axis and averaging the prediction scores. All experiments are on Nvidia Tesla V100 GPUs.
B Additional Experiments
To further demonstrate the advantages of our 2DPASS upon multi-sensor methods, we set several multi-sensor baselines and compare against them.
PointPainting: We follow the setup of previous work , which exploits the segmentation logits of images and projects them to the LiDAR space by bird’s-eye projection or spherical projection . Here, we use several pre-trained backbones, i.e., FCN with ResNet34 and DeepLab_v3 , to achieve the 2D semantic segmentation logits. After that, we use outputs of 2D backbones as the inputs of our 3D network.
Multi-branch Baseline: As shown in Fig. 1 (a), we design an ensemble architecture through concatenating the output logits from the two modalities.
Multi-branch with Interaction: Instead of only concatenating the predictions, we also concatenate the 2D features from each layer into the corresponding layers in the 3D network, as illustrated in Fig. 1 (b).
2DPASS (light): Since above multi-sensor manners are trained with the entire 2D image as input, they are time-consuming and GPU memory cost expensive. So we set all of hidden dimensions as 64 in the 3D network due to GPU memory limitation. This design is different from our manuscript with hidden dimensions 128 due to our light memory cost.
The experiment results are shown in Table 1, where we illustrate the results on NuScenes validation set and inference time (speeds), respectively. As shown in Table 1, using naive combination such as PointPainting and concatenation (i.e., Multi-branch Baseline) of prediction cannot improve the segmentation results obviously while introducing huge computational burden (i.e., there are six camera images corresponding to each point cloud). Exploiting feature combination in each scale can slightly improve the performance, but leads to much slower network compared with the pure 3D network. On the contrary, 2DPASS (light) achieves the second-best performance in term of mIoU criterion while 60 speed faster than multi-sensor methods.
B.2 Concrete Results
In this section, we give our detailed results on the NuScenes dataset in Table 2 as a benchmark for future work.