EPMF: Efficient Perception-aware Multi-sensor Fusion for 3D Semantic Segmentation
Mingkui Tan, Zhuangwei Zhuang, Sitao Chen, Rong Li, Kui Jia, Qicheng Wang, Yuanqing Li
Introduction
Semantic scene understanding is a fundamental task for many applications, such as auto-driving and robotics . Specifically, in the scenes of auto-driving, it provides fine-grained environmental information for high-level motion planning and improves the safety of autonomous cars . One of the important tasks in semantic scene understanding is semantic segmentation, which assigns a class label to each data point in the input data, and helps autonomous cars to better understand the environment.
According to the sensors used by semantic segmentation methods, recent studies can be divided into three categories: camera-only methods , LiDAR-only methods and multi-sensor fusion methods . Camera-only methods have achieved great progress with the help of a massive amount of open-access data sets . Since images obtained by a camera are rich in appearance information (e.g., texture and color), camera-only methods can provide fine-grained and accurate semantic segmentation results. However, as passive sensors, cameras are susceptible to changes in lighting conditions and are thus unreliable .See Section 4.5 for more details. To address this problem, researchers conduct semantic segmentation on point clouds from LiDAR. Compared with camera-only approaches, LiDAR-only methods are more robust to different light conditions, as LiDAR provides reliable and accurate spatio-depth information on the physical world. Unfortunately, LiDAR-only semantic segmentation is challenging due to the sparse and irregular distribution of point clouds. In addition, point clouds lack texture and color information, resulting in high classification error in the fine-grained segmentation task of LiDAR-only methods. A straightforward solution for addressing both drawbacks of camera-only and LiDAR-only methods is to fuse the multimodal data from both sensors, i.e., multi-sensor fusion methods. Nevertheless, due to the large domain gap between RGB cameras and LiDAR, multi-sensor fusion is still a nontrivial task.
In multi-sensor fusion methods, fusing multimodal data from different sensors is an important problem. Existing fusion-based methods mainly project dense image features to the LiDAR coordinates using spherical projection and conduct feature fusion in the sparse LiDAR domain. However, these methods suffer from a critical limitation: as the point clouds are very sparse, most of the appearance information from the RGB images is missing after projecting it to the LiDAR coordinates. For example, as shown in Figure 1 (c), the car and motorcycle in the image become distorted with spherical projection. As a result, existing fusion-based methods have difficulty capturing the appearance information from the projected RGB images.
In this paper, we aim to exploit an effective multi-sensor fusion method. Unlike existing methods , we assume and highlight that the perceptual information from both RGB images and point clouds, i.e., appearance information from images and spatio-depth information from point clouds, is important in fusion-based semantic segmentation. Based on this intuition, we propose a perception-aware multi-sensor fusion (PMF) scheme that conducts collaborative fusion of perceptual information from two modalities of data in three aspects. First, we propose a perspective projection to project the point clouds to the camera coordinate system to obtain additional spatio-depth information for RGB images. Second, we propose a two-stream network (TSNet) that contains a camera stream and a LiDAR stream to extract perceptual features from multimodal sensors separately. Considering that the information from images is unreliable in an outdoor environment, we fuse the image features to the LiDAR stream by effective residual-based fusion (RF) modules, which are designed to learn the complementary features of the original LiDAR modules. Third, we propose perception-aware losses to measure the vast perceptual difference between the two data modalities and boost the fusion of different perceptual information. Specifically, as shown in Figure 2, the perceptual features captured by the camera stream and LiDAR stream are different. Therefore, we use the predictions with higher confidence to supervise those with lower confidence.
Our contributions are summarized as follows. First, we propose a perception-aware multi-sensor fusion (PMF) scheme to effectively fuse the perceptual information from RGB images and point clouds. Second, by fusing the spatio-depth information from point clouds and appearance information from RGB images, PMF is able to address segmentation with undesired light conditions and sparse point clouds. More critically, PMF is robust to adversarial samples of RGB images by integrating the information from point clouds. Third, we introduce perception-aware losses into the network and force the network to capture the perceptual information from two different-modality sensors. The extensive experiments on two benchmark data sets demonstrate the superior performance of our method. For example, on nuScenes , PMF outperforms Cylinder3D , a state-of-the-art LiDAR-only method, by 0.8% in mIoU.
Related Work
In this section, we revisit the existing literature on 2D and 3D semantic segmentation, i.e., camera-only methods, LiDAR-only methods and multi-sensor fusion methods.
Camera-only semantic segmentation aims to predict the pixel-wise labels of 2D images. FCN is a fundamental work in semantic segmentation, which proposes an end-to-end fully convolutional architecture based on image classification networks. In addition to FCN, recent works have achieved significant improvements via exploring multi-scale information , dilated convolution , and attention mechanisms . However, camera-only methods are easily disturbed by lighting (e.g., underexposure or overexposure) and may not be robust to outdoor scenes.
2 LiDAR-Only Methods
To address the drawbacks of cameras, LiDAR is an important sensor on an autonomous car, as it is robust to more complex scenes. According to the preprocessing pipeline, existing methods for point clouds mainly contains two categories, including direct methods and projection-based methods . Direct methods perform semantic segmentation by processing the raw 3D point clouds directly. PointNet is a pioneering work in this category that extracts point cloud features by multi-layer perception. A subsequent extension, i.e., PointNet++ , further aggregates a multi-scale sampling mechanism to aggregate global and local features. However, these methods do not consider the varying sparsity of point clouds in outdoor scenes. Cylinder3D addresses this issue by using 3D cylindrical partitions and asymmetrical 3D convolutional networks. However, direct methods have a high computational complexity, which limits their applicability in auto-driving. Projection-based methods are more efficient because they convert 3D point clouds to a 2D grid. In projection-based methods, researchers focus on exploiting effective projection methods, such as spherical projection and bird’s-eye projection . Such 2D representations allow researchers to investigate efficient network architectures based on existing 2D convolutional networks . In addition to projection-based methods, one can easily improve the efficiency of networks by existing neural architecture search and model compression techniques .
3 Multi-Sensor Fusion Methods
To leverage the benefits of both camera and LiDAR, recent work has attempted to fuse information from two complementary sensors to improve the accuracy and robustness of the 3D semantic segmentation algorithm . RGBAL converts RGB images to a polar-grid mapping representation and designs early and mid-level fusion strategies. PointPainting obtains the segmentation results of images and projects them to the LiDAR space by using bird’s-eye projection or spherical projection . The projected segmentation scores are concatenated with the original point cloud features to improve the performance of LiDAR networks. Unlike existing methods that perform feature fusion in the LiDAR domain, PMF exploits a collaborative fusion of multimodal data in camera coordinates.
Proposed Method
In this work, we propose a perception-aware multi-sensor fusion (PMF) scheme to perform effective fusion of the perceptual information from both RGB images and point clouds. Specifically, as shown in Figure 3, PMF contains three components: (1) perspective projection; (2) a two-stream network (TSNet) with residual-based fusion modules; (3) perception-aware losses. The general scheme of PMF is shown in Algorithm 1. We first project the point clouds to the camera coordinate system by using perspective projection. Then, we use a two-stream network that contains a camera stream and a LiDAR stream to extract perceptual features from the two modalities, separately. The features from the camera stream are fused into the LiDAR stream by residual-based fusion modules. Finally, we introduce perception-aware losses into the optimization of the network.
Existing methods mainly project images to the LiDAR coordinate system using spherical projection. However, due to the sparse nature of point clouds, most of the appearance information from the images is lost with spherical projection (see Figure 1). To address this issue, we propose perspective projection to project the sparse point clouds to the camera coordinate system.
Because the point cloud is very sparse, each pixel in the projected may not have a corresponding point . Therefore, we first initialize all pixels in to 0. Following , we then compute 5-channel LiDAR features, i.e., , for each pixel in the projected 2D image , where represents the range value of each point.
2 Architecture Design of PMF
As images and point clouds are different-modality data, it is difficult to handle both types of information from the two modalities by using a single network . Motivated by , we propose a two-stream network (TSNet) that contains a camera stream and a LiDAR stream to process the features from camera and LiDAR, separately, as illustrated in Figure 3. In this way, we can use the network architectures designed for images and point clouds as the backbones of each stream in TSNet.
where indicates the concatenation operation. is the convolution operation w.r.t. the -th fusion module.
where indicates the sigmoid function. indicates the convolution operation in the attention module w.r.t. the -th fusion module. indicates the element-wise multiplication operation.
3 Construction of Perception-Aware Loss
The construction of perception-aware loss is very important in our method. As demonstrated in Figure 2, because the point clouds are very sparse, the LiDAR stream network learns only the local features of points while ignoring the shape of objects. In contrast, the camera stream can easily capture the shape and texture of objects from dense images. In other words, the perceptual features captured by the camera stream and LiDAR stream are different. With this intuition, we introduce a perception-aware loss to make the fusion network focus on the perceptual features from the camera and LiDAR.
Following , we use to normalize the entropy to . Then, the perceptual confidence map w.r.t. the LiDAR stream is computed by . For the camera stream, the confidence map is computed by .
Note that not all information from the camera stream is useful. For example, the camera stream is confident inside objects but may make mistakes at the edge. In addition, the predictions with lower confidence scores are more likely to be wrong. Incorporating with a confidence threshold, we measure the importance of perceptual information from the camera stream by
Here indicates the confidence threshold.
Inspired by , to learn the perceptual information from the camera stream, we construct the perception-aware loss w.r.t. the LiDAR stream by
where and indicates the Kullback-Leibler divergence .
In addition to the perception-aware loss, we also use multi-class focal loss and Lovász-softmax loss , which are commonly used in existing segmentation work , to train the LiDAR stream.The details of the multi-class focal loss and Lovász-softmax loss can be found in the supplementary material.
The objective w.r.t. the LiDAR stream is defined by
where and indicate the multi-class focal loss and Lovász-softmax loss, respectively. and are the hyperparameters that balance different losses.
Similar to the LiDAR stream, we construct the objective for the optimization of the camera stream. Following Eq. (6), the importance of the information from the LiDAR stream is computed by
The perception-aware loss w.r.t. the camera stream is
Then the objective w.r.t. the camera stream is defined by
Experiments
In this section, we empirically evaluate the performance of PMF on the benchmark data sets, including SemanticKITTI and nuScenes . SemanticKITTI is a large-scale data set based on the KITTI Odometry Benchmark , providing 43,000 scans with pointwise semantic annotation, where 21,000 scans (sequence 00-10) are available for training and validation. The data set has 19 semantic classes for the evaluation of semantic benchmarks. nuScenes contains 1,000 driving scenes with different weather and light conditions. The scenes are split into 28,130 training frames and 6,019 validation frames. Unlike SemanticKITTI, which provides only the images of the front-view camera, nuScenes has 6 cameras for different views of LiDAR.
We implement the proposed method in PyTorch , and use ResNet-34 and SalsaNext as the backbones of the camera stream and LiDAR stream, respectively. Because we process the point clouds in the camera coordinates, we incorporate ASPP into the LiDAR stream network to adjust the receptive field adaptively. To leverage the benefits of existing image classification models, we initialize the parameters of ResNet-34 with the pretrained ImageNet models from . We also adopt hybrid optimization methods to train the networks w.r.t. different modalities, i.e., SGD with Nesterov for the camera stream and Adam for the LiDAR stream. We train the networks for 50 epochs on both the benchmark data sets. The learning rate starts at 0.001 and decays to 0 with a cosine policy . We set the batch size to 8 on SemanticKITTI and 24 on nuScenes. We set to 0.7,0.5, and 1.0, respectively.We investigate the effect of in the supplementary material. To prevent overfitting, a series of data augmentation strategies are used, including random horizontal flipping, color jitter, 2D random rotation, and random cropping. Our source code is available at https://github.com/ICEORY/PMF.
2 Results on SemanticKITTI
To evaluate our method on SemanticKITTI, we compare PMF with several state-of-the-art LiDAR-only methods including SalsaNext , Cylinder3D , etc. Since SemanticKITTI provides only the images of the front-view camera, we project the point clouds to a perspective view and keep only the available points on the images to build a subset of SemanticKITTI. Following , we use sequence 08 for validation. The remaining sequences (00-07 and 09-10) are used as the training set. We evaluate the release models of the state-of-the-art LiDAR-only methods on our data set. Because SPVNAS did not release its best model, we report the result of the best-released model (with 65G MACs). In addition, we reimplement two state-of-the-art fusion-based methods, i.e., RGBAL and PointPainting , on our data set.
From Table 1, PMF achieves the best performance among projection-based methods. For example, PMF outperforms SalsaNext by 4.5% in mIoU. However, PMF performs worse than the state-of-the-art 3D convolutional method, i.e., Cylinder3D, by 1.0% in mIoU. As long-distance perception is also critical to the safety of autonomous cars, we also conduct a distance-based evaluation on SemanticKITTI. From Figure 5, because the point clouds becomes sparse when the distance increases, LiDAR-only methods suffer from great performance degradation at long distances. In contrast, since the images provide more information for distant objects, fusion-based methods outperform LiDAR-only methods at large distances. Specifically, PMF achieves the best performance when the distance is larger than 30 meters. This suggests that our method is more suitable to address segmentation with sparse point clouds. This ability originates from our fusion strategy, which effectively incorporates RGB images.
3 Results on nuScenes
Following , to evaluate our method on more complex scenes, we compare PMF with the state-of-the-art methods on the nuScenes LiDAR-seg validation set. The experimental results are shown in Table 2. Note that the point clouds of nuScenes are sparser than those of SemanticKITTI (35k points/frame vs. 125k points/frame). Thus, it is more challenging for 3D segmentation tasks. In this case, PMF achieves the best performance compared with the LiDAR-only methods. Specifically, PMF outperforms Cylinder3D by 0.8% in mIoU. Moreover, compared with the state-of-the-art 2D convolutional method, i.e., SalsaNext, PMF achieves a 4.7% improvement in mIoU. These results are consistent with our expectation. Since PMF incorporates RGB images, our fusion strategy is capable of addressing such challenging segmentation under sparse point clouds.
4 Qualitative Evaluation
To better understand the benefits of PMF, we visualize the predictions of PMF on the benchmark data sets.More visualization results on SemanticKITTI and nuScenes are shown in the supplementary material. From Figure 6, compared with Cylinder3D, PMF achieves better performance at the edges of objects. For example, as shown in Figure 6 (d), the truck segmented by PMF has a more complete shape. More critically, PMF is robust to different lighting conditions. Specifically, as illustrated in Figure 7, PMF outperforms the baselines on more challenging scenes (e.g., night). In addition, as demonstrated in Figure 6 (e) and Figure 7 (c), PMF generates dense segmentation results that combine the benefits of both the camera and LiDAR, which is significantly different from existing LiDAR-only and fusion-based methods.
5 Adversarial Analysis
To investigate the robustness of PMF on adversarial samples, we first insert extra objects (e.g., a traffic sign) to the images and keeping the point clouds unchanged.More adversarial samples are shown in the supplementary material. In addition, we implement a camera-only method, i.e., FCN , on SemanticKITTI as the baseline. Note that we do not use any adversarial training technique during training. As demonstrated in Figure 8, the camera-only methods are easily affected by changes in the input images. In contrast, because PMF integrates reliable point cloud information, the noise in the images is reduced during feature fusion and imposes only a slight effect on the model performance.
6 Efficiency Analysis
In this section, we evaluate the efficiency of PMF on GeForce RTX 3090. Note that we consider the efficiency of PMF in two aspects. First, since predictions of the camera stream are fused into the LiDAR stream, we remove the decoder of the camera stream to speed up the inference. Second, our PMF is built on 2D convolutions and can be easily optimized by existing inference toolkits, e.g., TensorRT. In contrast, Cylinder3D is built on 3D sparse convolutions and is difficult to be accelerated by TensorRT. We report the inference time of different models optimized by TensorRT in Table 3. From the results, our PMF achieves the best performance on nuScenes and is faster than Cylinder3D (22.3 ms vs. 62.5 ms) with fewer parameters.
Ablation Study
We study the effect of the network components of PMF, i.e., perspective projection, ASPP, residual-based fusion modules, and perception-aware loss. The experimental results are shown in Table 4. Since we use only the front-view point clouds of SemanticKITTI, we train SalsaNext as the baseline on our data set using the officially released code. Comparing the first and second lines in Table 4, perspective projection achieves only a 0.4% mIoU improvement over spherical projection with LiDAR-only input. In contrast, comparing the fourth and fifth lines, perspective projection brings a 5.9% mIoU improvement over spherical projection with multimodal data inputs. From the third and fifth lines, our fusion modules bring 2.0% mIoU improvement to the fusion network. Moreover, comparing the fifth and sixth lines, the perception-aware losses improve the performance of the network by 2.2% in mIoU.
2 Effect of Perception-Aware Loss
To investigate the effect of perception-aware loss, we visualize the predictions of the LiDAR stream networks with and without perception-aware loss in Figure 9. From the results, perception-aware loss helps the LiDAR stream capture the perceptual information from the images. For example, the model trained with perception-aware loss learns the complete shape of cars, while the baseline model focuses only on the local features of points. As the perception-aware loss introduces the perceptual difference between the RGB images and the point clouds, it enables an effective fusion of the perceptual information from the data of both modalities. As a result, our PMF generates dense predictions that combine the benefits of both the images and point clouds.
Conclusion
In this work, we have proposed a perception-aware multi-sensor fusion scheme for 3D LiDAR semantic segmentation. Unlike existing methods that conduct feature fusion in the LiDAR coordinate system, we project the point clouds to the camera coordinate system to enable a collaborative fusion of the perceptual features from the two modalities. Moreover, by fusing complementary information from both cameras and LiDAR, PMF is robust to complex outdoor scene. The experimental results on two benchmarks show the superiority of our method. In the future, we will extend PMF to other challenging tasks in auto-driving, e.g., object detection.
This work was partially supported by Key-Area Research and Development Program of Guangdong Province 2019B010155001, Ministry of Science and Technology Foundation Project (2020AAA0106901), Guangdong Introducing Innovative and Enterpreneurial Teams 2017ZT07X183, Fundamental Research Funds for the Central Universities D2191240.