FusionPainting: Multimodal Fusion with Adaptive Attention for 3D Object Detection
Shaoqing Xu, Dingfu Zhou, Jin Fang, Junbo Yin, Zhou Bin, Liangjun Zhang
I Introduction
3D object detection, as a fundamental task in computer vision and robotics, has been extensively studied with the development of autonomous driving and intelligent transportation. Currently, LiDAR sensor is widely used for perception tasks because it can provide accurate range measurements of the surrounding environment, especially for the obstacles such as vehicles, pedestrians and cyclists, etc. With the development of the deep learning techniques on point-cloud based representation, many LiDAR-based 3D object detection approaches have been developed. Generally, these approaches can be categorized into point-based and voxel-based methods. LiDAR sensors have the superiority of providing distance information of the obstacles, even though the detailed geometry is often lost due to its sparse scanning and furthermore texture and color information can be not captured. Therefore, False Positive (FP) detection and wrong categories classification often happen for LiDAR-based object detection solutions.
On the contrary, the camera sensors can provide detailed texture and color information with high resolution, though the depth has been lost during the perspective projection based imaging procedure. The combination of the two different types of sensors of LiDAR and camera is a promising way for boosting the performance of autonomous driving perception. In literature, multi-modal based object detection approaches can be divided into as early fusion , deep fusion and late fusion approaches . Early fusion approaches aim at creating a new type of data by combining the raw data directly before sending them into the detection framework. Usually, these kinds of methods require pixel-level correspondence between each type sensor data. Different from the early fusion methods, late fusion approaches execute the detection for each type of data separately first and then fuse the detection results in the bounding box level. Different from the above two methods, deep fusion-based methods usually extract the features with different types of deep neural networks first and then fuse them at the features level. As a simple yet effective sequential fusion method, PointPainting has achieved superior detection results on different benchmarks. This approach employees the 2D image semantic segmentation results from an off-the-shelf neural network first and then adds them into a point-cloud-based 3D object detection based on the 2D-3D projection. The superiority of PointPainting suggests that 2D image segmentation approach can be used for providing the semantic results and it can be incorporated into any 3D object detectors even point-based or voxel-based approaches.
However, the boundary-blurring effect often happens in image-based semantic segmentation methods due to the relatively low resolution of the deep feature map. This effect becomes much more severe after re-projecting them into the 3D point cloud. An example of the reprojected 2D result into 3D is shown in sub-fig. 1-(a). Taking the big truck at the bottom of the image as an example, we can find that there is a large frustum area of the background (e.g, points in orange color) that has been miss-classified as foreground due to the inaccurate segmentation results in the 2D image. In addition, the correspondence of 3D points to 2D image pixels is not exactly a one-to-one projection due to the digital quantization problem, and many-to-one projection issues. An interesting phenomenon is that the segmentation from the 3D point cloud (e.g., sub-fig. 1-(b)) performs much better on the boundary of obstacles. However, compared to the 2D image, the category classification from the 3D point cloud often gives worse results(e.g.,point in blue color) due to the detailed texture information from the RGB images.
The painted points , with semantic information has been proved to be very effective for the object detection task even with some semantic errors. An intuition idea is that the detection performance can be further improved if 2D and 3D segmentation results can be fused together. Based on this idea, we propose a general multi-modal fusion framework “FusionPainting” to fuse different types of sensors at the semantic level to further boost the 3D object detection performance. First, any off-the-shelf semantic segmentors are employed directly for obtaining the semantic information. To further improve the performance, an attention module is proposed to fuse different kinds of semantic information in voxel-level adaptively based on the learned context features in a self-attention style. Finally, the points painted by the fused semantic labels are sent to 3D detectors for obtaining the final detection results. The main contributions of this work can be summarized as following:
A general multi-modal fusion framework “FusionPainting” has been firstly proposed to fuse the different types of information in the semantic level to improve the 3D object detection performance.
To further improve the performance, an attention module is proposed to fuse different kinds of semantic information in voxel-level by learning the context features.
The experimental results on the large-scale autonomous driving benchmark nuScenes show the superiority of the proposed fusion framework and achieve SOTA results compared to other methods.
II Related Work
The existing point cloud-based 3D object detection methods can be categorised as projection-based , voxel-based and point-based . RangeRCNN is a typical 3D object detection framework which is based on the range image, and generates anchors on the BEV (bird’s-eye-view) map. Different from hand-crafted features design in previous works, VoxelNet proposes a VFE (voxel feature encoding) layer, which can learn a unified feature representation for each 3D voxel. Based on VoxelNet, CenterPoint presents an anchor-free method which use a center-based framework based on CenterNet , and achieves state-of-the-art performance. SECOND leverages the sparse convolution operation to alleviate the burden of 3D convolution. PointPillars extracts the features from vertical columns (Pillars) with PointNet and encodes the features as a pseudo image, then 2D object detection pipeline can be applied. PointRCNN is a pioneer work that directly generates 3D proposals from a raw point cloud. Furthermore, PV-RCNN integrates the multi-scale 3D voxel CNN and PointNet-based network to learn more discriminative features.
II-B Multi-modal Fusion
Various methods exploit combining multiple types of sensory data to boost the detection performance. LiDAR and camera are the most used sensors. MV3D proposes a multi-view representation including BEV, camera view and front view image. The framework generates proposals from the BEV and deeply fuses the features from other views. AVOD takes LiDAR point clouds and RGB images as input, to generate features shared by RPN (region proposal network) and the refined network. F-PointNet uses the image to generate proposals and refines the final bounding box in cropped frustums with PointNet. Different from the above methods, PointPainting uses an easy but effective strategy: retrieving the class score vector for each point by projecting them into image semantic segmentation results, and feeding the concatenated attribution to 3D object detection framework. Multi-task fusion is also an effective technology, e.g., jointing the semantic segmentation with 3D object detection tasks to further improve the performance. Furthermore, in and , the HD Maps are also taken as an strong prior information for movable object detection in autonomous driving scenario.
III Proposed Approach
We describe the architecture of our “FusionPainting” framework in Fig. 2. We propose to leverage both the 2D images and 3D point clouds information for obtaining accurate 3D locations of obstacles in the autonomous driving scenarios. Here, the fusion process is achieved by adaptively integrating 2D and 3D information at the semantic level. As shown in Fig. 2, the proposed framework consists of three models as a multi-modal semantic segmentation module, an adaptive attention module, and a 3D detector module. First, any off-the-shelf 2D and 3D segmentors can be employed to obtain the semantic segmentation from the RGB image and LiDAR point clouds respectively. Then a simple but effective attention strategy is proposed to sufficiently benefit the merits and suppress the drawbacks of the 2D and 3D semantic results. Finally, painted 3D point clouds with enhanced semantic labels are sent to modern 3D object detectors for achieving the final detection results.
III-B Adaptive Attention Module
Though the 2D segmentation achieves impressive performance, the boundary-blurring effect is also inevitable due to the limited resolution of the feature map. Therefore, the painted point cloud from the 2D segmentation mask usually has misclassified regions around the boundary of the objects. For example, the frustum region is illustrated in the sub-fig. 1 (a) behind the big truck. On the contrary, the point cloud-based semantic segmentation methods often can produce a clear and accurate object boundary e.g., sub-fig. 1(b). In order to combine the merits of both modules and suppress their disadvantages, we propose an adaptive attention module that adaptively fuses the semantic labels with an attention mechanism. By doing this, the refined semantic labels can further improve the following 3D detection results.
For the local feature, a PointNet -like module is employed here to extract the voxel-wise information inside each non-empty voxel. For -th voxel, its local feature can be represented as
where and are the local muti-layer perception (Mlp) networks and max-pooling operation. Specifically, consists of a linear layer, a batch normalization layer, an activation layer and outputs local features with channels. To achieve global feature information, we aggregate information based on the voxels. In particular, we first use a to map each voxel features from dimension to dimension. Then, another PointNet-like module is applied on all the voxels as
Then, the fused feature is adopted to estimate an attention score for each voxel. This is achieved by applying another Mlp module on followed by a Sigmod activation function . Afterwards, we multiply the confidence score by corresponding one-hot semantic vectors for each point in a voxel, which is denoted as
III-C 3D Object Detection Module
After obtaining the painted point cloud, any off-the-shelf 3D object detectors can be directly employed for predicting 3D object detections. The 3D detector receives the painted voxels produced by the adaptive attention module, and achieve better results for all the categories. The detailed analysis of the performance can be found in the following experimental results part.
IV Experimental Results
We evaluate the effectiveness of the proposed “FusionPainting” on the large-scale autonomous driving 3D object detection dataset. First, the details of experiments are described and then the evaluation results on three baseline detectors are given.
Baselines: in addition, to verify its universality,we have implemented the proposed module on three different State-of-the-Art 3D object detectors as following,
SECOND , which is the first to employ the sparse convolution on the voxel-based 3D object detection task and the detection efficiency has been highly improved compared to previous work such as VoxelNet.
PointPillars , which divides the point cloud into pillars and “Pillar Feature Net” is applied to each pillar for extracting the point feature. Then 2d convolution network has been adopted on the bird-eye-view feature map for object detection. PointPillars is a trade-off between efficiency and performance.
CenterPoint , is the first anchor-free based 3D object detector which is very suitable for small object detection.
Dataset: nuScenes 3D object detection benchmark has been employed for evaluation here, which is a large-scale dataset with a total of 1,000 scenes. For a fair comparison, the dataset has been officially divided into train, val, and testing, which includes 700 scenes (28,130 samples), 150 scenes (6019 samples), 150 scenes (6008 samples) respectively. For each video, only the key frames (every 0.5s) are annotated with 360-degree view. With a 32 lines LiDAR used by nuScenes, each frame contains about 300,000 points. For object detection task, 10 kinds of obstacles are considered including “car”, “truck”, “bicycle” and “pedestrian” et al. Besides the point clouds, the corresponding RGB images are also provided for each keyframe. For each keyframe, there are 6 images that cover 360 fields-of-view.
Evaluation Metrics: For 3D object detection, proposes mean Average Precision (mAP) and nuScenes detection score (NDS) as the main metrics. Different from the original mAP defined in , nuScenes consider the BEV center distance with thresholds of {0.5, 1, 2, 4} meters, instead of the IoUs of bounding boxes. NDS is a weighted sum of mAP and other metric scores, such as average translation error (ATE) and average scale error (ASE).
Implementation Details: HTCNet and Cylinder3D are employed here as the 2D Segmentor and 3D Segmentor, respectively, due to their outstanding semantic segmentation ability. The 2D segmentor is pretrained on nuImageshttps://www.nuscenes.org/nuimages dataset, and the 3D segmentor is pretrained on nuScenes detection dataset while the point cloud semantic information can be generated by extracting the points inside each bounding box of the obstacles. For Adaptive Attention Module, , , respectively, For each baseline, all the experiments share the same setting, the voxel size for SECOND, PointPillar and CenterPoint are , and , respectively. We use AdamW with max learning rate 0.001 as the optimizer. Follow , 10 previous LiDAR sweeps are stacked into the keyframe to make the point clouds more dense.
IV-B Evaluation Results
We have evaluated the proposed framework on nuScenes benchmark for both validation and test splits.
Evaluation on Baseline Methods: First of all, we have integrated the proposed fusion module into three different baselines methods and experimentally we have achieved consistently improvements on both mAP and NDS compared to all baselines. Detailed results are given in Tab. I. From this table, we can obviously find that the proposed module gives more than 10 points improvements on mAP and 5 points improvements on NDS for all the three baselines. The improvements have reached 17.31 and 8.45 points for PointPillars method on mAP and NDS respectively. In addition, we find that the “Traffic Cone”, “Moto” and “Bicycle” have received the surprising improvements compared to other classes. For “Bicycle” category specifically, the AP has improved about 36.19%, 48.83% and 26.84% based on Second, PointPillar and Centerpoint respectively. Interestingly, all these categories are small objects which are hard to be detected because of a few LiDAR points on them. By adding the prior semantic information, the category classification becomes relatively easier. Especially, the experiments here we share the same setting, so the baseline figures may be subtle differences with official.
Comparison with other SOTA methods: To compare the proposed framework with other SOTA methods, we submit our best results (proposed module on the CenterPoint ) to the nuScenes evaluation severhttps://www.nuscenes.org/object-detection/. The detailed results are listed in Tab. II. To be clear, only these methods with publications are compared here due to space limitation. From this table, we can find that the “NDS” and “mAP” have been improved by 3.1 points and 6.0 points respectively compared with the baseline method CenterPoint . More importantly, our algorithm outperforms previous multi-modal method 3DCVF by a large margin, i.e., improving 8.1 and 13.6 points in terms of NDS and mAP metrics.
IV-C Qualitative Results
More qualitative detection results are illustrated in Fig. 4. Fig. 4 (a) shows the annotated ground truth, (b) is the detection results for CenterPoint based on raw point cloud without using any painting strategy. (c) and (d) show the detection results with 2D painted and 3D painted point clouds, respectively, while (e) is the results based on our proposed framework. As shown in the figure, there are false positive detection results caused 2D painting due to frustum blurring effect, while 3D painting method produces worse foreground class segmentation compared to 2D image segmentation. Instead, our FusionPainting can combine two complementary information from 2D and 3D segmentation, and detect objects more accurately.
IV-D Ablation Studies
To verify the effectiveness of different modules, a series of ablation studies have been designed here. All the experiments are executed on the validation split and the settings keep the same as in Sec. IV-A. All the results are given in Tab. III, from where we could figure out the impact of each module.
As demonstrated in Tab. III, 3D semantic segmentation information alone contributes about 9.81% mAP and 3.92% NDS averagely, while 2D semantic segmentation information contributes about 21.90% and 8.46% averagely. The higher improvement from 2D semantic information benefit from the better recall ability especially for objects in the distance where the point cloud is very sparse. The advantage for 3D semantic information is that the point clouds can well handle the occlusion among objects which is very common and hard to deal in 2D. But the most improvement comes from the combination of 3D Painting, 2D Painting and Adaptive Attention Module, 26.58% mAP and 10.90% averagely from 3 detectors we used.
V Conclusion and Future Works
We present a novel multi-modal fusion framework, “FusionPainting” which can aggregate rich semantic information from 2D/3D semantic segmentation networks. More important, we are the first to use 3D semantic segmentation results as additional information for improving the 3D object detection performance. To fuse both the 2D/3D segmentation results, an adaptive attention module has been proposed to learn an attention mask for each of them. Experimental results on nuScenes dataset illustrate the effectiveness of the proposed strategy. Furthermore, the proposed “FusionPainting” is detector independent, which can be freely used for other 3D object detectors. In the future, we plan to deeply integrate the 2D/3D semantic segmentation branches into the detection framework and simultaneously achieve instance segmentation and 3D object detection tasks.