Spatial Pruned Sparse Convolution for Efficient 3D Object Detection
Jianhui Liu, Yukang Chen, Xiaoqing Ye, Zhuotao Tian, Xiao Tan, Xiaojuan Qi
Introduction
3D object detection has always been a research field of great interest due to its wide applications in autonomous driving, virtual reality, and robotics. However, compared with the 2D detection task, LiDAR-based 3D detection is more challenging due to the inherent characteristics of LiDAR points such as disorders, irregularity, and non-uniformity. Recently, voxel-based methods with deep 3D sparse convolutional neural networks (CNNs) have become one of the major research streams to tackle 3D object detection problem due to its effectiveness and simplicity. However, the improved accuracy is often accompanied by increased computational costs , limiting its applicability in practical systems. This motivates us to investigate potential redundancies in the detection model that can be safely avoided to improve efficiency without sacrificing accuracy.
When delving into the task of 3D detection, we found that 3D data itself has high redundancy as shown in Fig. 1(a). Notably, compared with background points in each stage of the sparse CNN, the proportion of foreground points in the entire scene is extremely low (around 5%), which shows that less-informative background points occupy the major areas of a scene. However, the existing 3D sparse CNNs are applied uniformly to the whole scene, causing a considerable amount of computation on the background areas, which might be potentially redundant. Intuitively, if such areas can be identified and selectively skipped, the computational costs have the potential to be dramatically reduced without deteriorating the performance.
Besides the redundancy of the data, the model design itself also brings redundancy. To avoid aggressive down-sampling, 3D sparse CNNs adopt a dilation-like design when applied to perform the down-sampling operation. Specifically, it will compute features for adjacent empty voxels as shown in Fig. 2 (b): the green voxels which are initially empty will be assigned with computed features. Consequently, after convolutional (stride > 1) down-sampling, the number of non-empty voxels might be increased rather than decreased as shown in Fig. 1 (b): the number of non-empty voxels even doubled (see stage 2) compared to the input (see stage1). This will undoubtedly increase unnecessary computational costs for follow-up stages.
The above observations prompt us to ask whether there is a way to identify redundant computations that can be pruned to improve efficiency. A solution that has been attempted in the 2D image domain is to add an auxiliary learnable module to predict a soft mask that locates areas to be skipped for computational efficiency. The module often requires additional post-fine-tuning, auxiliary costs for integration, and incurs non-negligible computational overheads.
With the aforementioned considerations, we propose a new simple and efficient sparse convolution operator named Spatial Pruned Sparse Convolution (SPS-Conv), to alleviate redundancies caused by data and model design. The core idea of SPS-Conv is to find redundancy in the model dynamically. To avoid the overheads of learning-based methods for simplicity and efficiency, we investigate whether the features from the model itself contain useful cues that can be leveraged to identify redundancies. Fortunately, we find that the magnitude of features could be a robust signal to reflect the importance where locations with smaller magnitude are more likely to be redundant areas. This leads us to a magnitude-based sampling module which is incorporated into 3D convolution layers to reduce redundancy in data and model.
Specifically, to address data redundancy, the submanifold variant of SPS-Conv, termed as spatial pruned submanifold sparse convolution (SPSS-Conv) can adaptively calculate important positions in the light of the magnitude of features and perform convolutions on these locations, leaving features of redundant locations unchanged. As for model redundancy in the down-sampling process, another variant named spatial pruned regular sparse convolution (SPRS-Conv), which can dynamically determine the position that needs to be inflated. As Tab. 1(b) shows, our method effectively reduces the meaningless expansion caused by convolution (around 50% in stage2). In general, SPS-Conv only keeps the essence and discards the irrelevant ones for pursuing a higher efficiency without compromising performance.
In summary, we propose an efficient convolution operator for spatial redundancy pruning, which can be easily incorporated into existing sparse 3D CNNs for 3D detection. We conduct extensive experiments on both KITTI, Waymo, and nuScenes datasets. The result shows that we can achieve comparable performance with SOTA methods while enjoying 52.4% , 62.46%, and 46.5% GFLOPs reduction.
Related Work
Sparse Convolution. Depending on how it is used, sparse convolution can be subdivided into regular sparse convolution and submanifold sparse convolution. Regular sparse convolution is often used for the down-sample layer, which dilates all input features to its kernel-size neighbors, sacrificing computational efficiency in exchange for a wider interaction of information. In contrast, the submanifold convolution is more like a simplified version of the regular sparse convolution, frequently used in residual blocks, only calculates on valid positions, avoiding meaningless expansion and achieving efficient computation.
Voxel-based Detectors. In order to process irregular point cloud data, voxel-based methods first convert the point cloud into regular voxel, and then use mature CNNs for feature extraction. However, the computational cost and memory requirement both increase cubically along with the voxel resolution. Thus, it is infeasible to train a voxel-based model with high-resolution inputs. Benjamin et al. proposed a novel convolution operator named submanifold sparse convolution, by reducing calculation on invalid position, which greatly improves the computing efficiency and memory.
Approaches for voxel-based detection methods can be grouped into two categories, i.e., single-stage and two-stage detectors. Single-stage detectors are relatively simple, directly predict final bounding boxes based on the features extracted by sparse CNN. VoxelNet utilizes a Sparse CNN to extract voxel features from a dense grid. SECOND proposes 3D sparse convolutions to efficiently extract voxel features. HVNet designs a convolutional network that attentively aggregates and projects the multi-scale feature maps to achieve better performance. In contrast, two-stage detectors are more complex, but can get higher performance. Part-A2 proposed a part-aware and aggregation module to exploit the intra-object part location. PV-RCNN uses keypoints to extract voxel features for better boxes refinement.
Dynamic kernel shape design. Many literature work on dynamic kernel design for adapting various tasks. Deformable convolution predicts offets for feature sampling. For 3D scene understanding, KPConv constructs dynamic graph for kernel points. Deformable PV-RCNN applies offset prediction for feature sampling in 3D object detection. Focals Conv learns a dynamic spatially sparsity which enhances network for spatial modeling capabilities,
Background and Motivation
In this section, we will introduce the mechanism of sparse convolution, then analyze its redundancy and introduce our motivation.
Here, and refer to the input and output feature space, is the kernel offset that corresponds to all the valid locations in kernel space , denotes all non-empty neighbor voxels around center . Due to the discrete data distribution in 3d space, we define as a subset of , leaving out the empty position, which is constrained by position and input feature space .
There are two types of sparse convolutions, namely, regular sparse convolution and submanifold sparse convolution , their major difference lies in the output position of the convolution. For regular sparse convolution, as shown in Fig. 2, the position related to the input will be activated during the convolution process. Specifically, these positions are assigned value and combined with the to form a new . This process can be easily formulated as
where is a subset of , representing the position where the convolution needs to be performed. refers to the position related to which is restricted by kernel size.
Regular sparse convolution can effectively expand the receptive field like 2D convolution, which is beneficial for 3D sparse data. However, it also brings computation burden due to the generation of too many active locations, which can lead to a decrease in speed in subsequent convolution layers due to the large number of active points. In order to achieve a better trade-off between efficiency and receptive field, regular sparse convolution is often adopted in the down-sample layer in sparse CNNs.
In contrast, submanifold sparse convolution restricts an output location to be active if and only if the corresponding input location is active, which avoids the dilation in regular sparse convolution, although the receptive field is limited. With a simple modification on Eq. 2, , regular mode can be convert to submanifold mode.
2 Redundancy Analysis and Motivations
Despite the wide adoption of sparse CNNs in 3D object detection due to their generality and efficiency, spatial redundancy caused by sparse CNNs for this task is still under-explored. Considering the large number of task-independent background points in the 3D scene, a lot of redundant computation will be generated in each stage. In addition, due to the expansion of regular convolution, the background points will become denser, bringing more computational burden to the following stages. Therefore, in response to these problems, we have targeted the improvement of the sparse convolution operator. Details are shown in Sec. 4
Spatial Pruned Sparse Convolution
In this section, we propose a new efficient sparse convolution operator, named spatial pruned sparse convolution (SPS-Conv). Intending to effectively bring down the unnecessary computation costs caused by the 3D spatial redundancy, two variants of SPS-Conv are investigated, i.e., spatial pruned submanifold sparse convolution (SPSS-Conv) and spatial pruned regular sparse convolution (SPRS-Conv), and both of them are based on the same idea that the redundancy should be dynamically ameliorated in 3D space according to the specific contexts of different individuals. In the following, the magnitude-based sampling strategy is presented in Sec. 4.1, followed by the introduction of SPSS-Conv and SPRS-Conv in Sec. 4.2 and Sec. 4.3 respectively. Then, Sec. 4.4 manifests the generalization ability by showing how our method is applied to generic sparse CNNs.
In previous works , sampling module is often designed in a learnable manner, i.e., a convolution layer is trained to yield soft masks for sampling, while, at the time of inference, the soft masks are converted to hard ones where elements with lower responses will be simply discarded. Although these methods have been empirically shown effective in 2D tasks, the following issues still exist: 1) Additional supervision is required on the prediction branch, otherwise it is easy to fall into trivial solutions, i.e., the values in the soft mask are all close to 1. 2) The proportion of pruning is hard to be fixed, and it is instead automatically determined throughout the learning process, making it hard to satisfy the specific requirement with certain limited computation budget. 3) Two-stage fine-tuning is required to compensate for the performance loss of converting from soft mask to hard mask, consuming additional time and resources.
Recent literature has shown that the attention mechanism is able to reveal what the model really cares about. Following , the feature magnitude is adopted for yielding the model’s “attention” to highlight the informative elements. For the purpose of getting the magnitude-based attention mask, we first calculate the channel-wise absolute mean values on different voxels, and then the sigmoid function is applied to get the normalized output, which can be formulated as
where and denote the input feature and its channel. is the feature of the -th dimension. and refer to the spatial magnitude map and magnitude mask.
2 Spatial Pruned Submanifold Sparse Convolution
Considering the fact that the background regions often take the majority in 3D scenes, directly performing the submanifold convolution everywhere of the input volume inevitably involves substantial unnecessary calculations, leading to computational redundancy. Therefore, with a focus on alleviating this issue, we propose the spatial pruned submanifold sparse convolution (SPSS-Conv) that dynamically examines the regions of interest.
Specifically, with the guidance of the magnitude-based attention mask, we separate the elements into two disjoint subsets, i.e., the important set and the unimportant set , where . We note that various methods can be incorporated for accomplishing the division, such as using a fixed threshold or simply taking the elements with top-k scores. Then we re-weight the input feature map by multiplying it with the magnitude mask, applying submanifold convolution on important positions according to Eq. 1, and concatenate it with the features of unimportant positions as a residual, as shown in Fig. 3. To this end, the reconstructed output feature can be written as:
Since is sparser than , convolution enables task-beneficial locations to be highlighted. In contrast, those unimportant positions have relatively small activation values, thus the skip connection can not only reduce the computational overhead but also suppress the diffusion of redundant information.
3 Spatial Pruned Regular Sparse Convolution
The “dilation” effect of the regular sparse convolution inevitably causes a large number of adjacent positions to be activated, which brings a greater computational burden to subsequent layers. Recent work shows that the spatially dynamic sparsity in sparse convolution is essential for sophisticated 3D object detection. Motivated by this, SPRS-Conv is designed as a magnitude-based dynamic prediction down-sampling module, which can effectively suppress the amplification effect of spatial redundancy.
Similar to SPSS-Conv, is first divided into two disjoint subsets, i.e, and . Unlike the regular sparse convolution that dilates all possible neighborhood locations, leading to sub-optimal efficiency, we instead dynamically choose whether to expand for each position based on the magnitude mask yielded by Eq. 3. Further, we define , which is a set composed of all positions and the positions within their kernel range. The equation can be written as:
Considering the stride during down-sampling, we designed a stride mask to ignore those that should be skipped by the stride, which can be formulated as below:
where refers to the stride of the convolution on a specific dimension, % denotes mod operation. With , we can further modify Eq. 2 as:
As shown in Fig. 4, we operate at the specified positions according to Eq. 1 to get the output feature map. Then we down-sample it to the correct size according to the stride. Compared to the regular sparse convolution, our method dynamically prunes conditioned on its specific content. As shown in Fig. 1(b), SPRS-Conv greatly reduces the number of voxels in each stage, especially stage2 (about 50%), thus reducing the computational burden caused by background points
4 Spatial Pruned Convolution Network
Our method is model-agnostic thus it can be easily plugged into any existing sparse CNNs. A regular frame of sparse CNN is composed of a stem layer and four stages, each of which contains a down-sampling layer and two submanifold sparse convolution blocks. To demonstrate the effectiveness and generalization ability, we replace all regular sparse convolutions and submanifold sparse convolutions in sparse CNN, except those in the stem layer, with our proposed SPRS-Conv and SPSS-Conv, respectively. Detailed experimental results and analysis are presented in Sec. 5.
Experiments
3D object detection datasets. We evaluate our method on three challenging benchmarks KITTI , Waymo and nuScenes . KITTI dataset consists of 7,481 samples and 7,518 testing samples, where the training samples are generally divided into the train split (3, 712 samples) and the val split (3, 769 samples). Waymo dataset contains 1,000 sequences in total, with 798 for training and 202 for validation. As the Waymo dataset is really large-scale, we use data for training. Results in Waymo are evaluated in difficulty LEVEL_1 (L1) and LEVEL_2 (L2) objects. nuScenes dataset is a large-scale autonomous driving dataset, which contains 1,000 driving sequences in total. It is split into 700 scenes for training, 150 scenes for validation, and 150 scenes for testing. It is collected using a 32-beam synced LIDAR and 6 cameras with the complete 360o environment coverage.
Model configurations. i.e., VoxelR-CNN , PV-RCNN and SECOND for KITTI, CenterPoint for nuScenes and Waymo. We replace the submanifold convolution and regular convolution in sparse CNNs with SPSS-Conv and SPRS-Conv, respectively, other experimental hyperparameters flowing the default settings of the baseline methods .
Training. Unlike existing 2D approaches that design the sampling module in a learnable manner which needs a two-stage fine-tune. We train the model in an end-to-end manner. Our SPS-Conv is purely embedded into sparse CNNs without any modifications on network architectures or parameters (i.e. feature dimensions). Our modules contains one hyperparameter, i.e. pruning ratio and we choose to use top-k to divide the pruning part. For spatial pruned submanifold sparse convolution (SPSS-Conv) it refers to the proportion of positions in each stage that can be ignored for calculation, which can be symbolized as {, , , }. We set it as {0.3, 0.3, 0.3, 0.3 } for nuScenes and {0.5, 0.5, 0.5, 0.5} for Waymo and KITTI. As for spatial pruned regular sparse convolution (SPRS-Conv), pruning ratio is used for controlling the amount that needs to be inflated when downsampling. We symbolized it as {, , }, which are set as {0.5, 0.5, 0.5} in nuScenes and Waymo and {0.7, 0.5, 0.3} in KITTI.
2 Main Results
nuScenes. For the sake of fairness, we choose LIDAR-only methods for comparison, model ensembling and additional augmentation are not included during inference. Our method achieves competitive performance on both mAP and NDS on nuScenes split, in Tab. 1. Besides, We apply the method to the base model for a more detailed comparison, as shown in Tab. 2, our method can help the model to preserve the original performance while skipping redundant computation, specifically, SP-CenterPoint’s GFLOPs are only 54.5% of the base model.
KITTI. We make comparisons with recent state-of-the-art methods on KITTI split. As shown in Tab. 5, Our Spatial Pruned Voxel R-CNN achieves competitive results with other methods. To better demonstrate the effectiveness of our method, we validate our method on popular voxel-based detectors, Second, PVRCNN, and Voxel-RCNN. Tab. 5 shows that with our method, the GFLOPs of the sparse CNN are reduced by more than 50%, and the matching performance can still be obtained on KITTI split in in recall 11 positions.
Waymo. We present SPS-Conv based on the CenterPoint on the outdoor dataset Waymo in Tab. 3. The results demonstrate the generality of SPS-Conv on large-scale datasets with dense point clouds: SPS-conv still maintains the high performance while significantly reduces FLOPs.
3 Ablation Study
SPS-Conv pruning ratio. The pruning ratio (PR.) of the SPS-Conv is used to control the proportion of unimportant positions selected in each block, the higher the proportion, the fewer positions are involved in the calculation. We evaluate on nuScenes val split upon CenterPoint. As shown in Tab. 7 and Tab. 7, there is a relatively obvious and proportional decrease in GFLOPs, in contrast, the performance drop is not so severe, except when the pruning ratio increases to 0.9. This also reveals that 3D Scenes contain spatial redundancy, and when we selectively skip these redundant positions, it will not affect the performance.
Sampling and interpolation strategy in SPSS-Conv. In order to verify the effectiveness of our method, we conduct a combination of several sampling and interpolation methods, where the pruning ratio is limited to 0.5. Results are shown in Tab. 9, it is not difficult to find that the learnable method and magnitude method have comparable results which make us think that magnitude is another manifestation of network attention, so we select important regions according to this logic, which conforms to the inherent performance of the network. In addition, for the interpolation operation, compared to the average pooling operation, a simple skip connection can play a good role. A possible explanation is that these predicted unimportant regions are not task-essential, so even using the original values would not have much impact on the results.
Importance of position with high magnitude. To verify that the location chosen by this mechanism is task-friendly, we exchanged the calculation method of positions with high magnitude and other positions. As shown in Tab. 8, for SPSS-Conv, it can be seen that when we reverse the important position, there is a considerable performance drop in all categories. This phenomenon is more obvious in small objects, especially for the traffic cone and bicycle categories (even reaching about a 7 performance drop). Large object categories (such as vehicles) with more points are less sensitive to the sampling method but still have a notable performance drop. Compared with the experimental results of SPSS-Conv, the inversion experiment of SPRS has a more exaggerated performance loss. When inversion is not performed, SPRS will suppress the irrelevant features with low magnitude which reduced in the downsample part, showing that there is no obvious performance loss in the result. However, when we choose to suppress positions with large magnitudes, the spatial redundancy is amplified, and important features cannot be effectively expanded, and even discarded due to downsampling. After the above features are converted to Bird’s Eye View (BEV), since the number of foreground points from the input is weakened, the effective features extracted by BEVbackone are very limited, resulting in a great degree of performance degradation.
Magnitude-base pruning vs Random drop. Compared to magnitude-based pruning, we observe using random drop as an indicator will lead to a certain loss in performance (around 2%), shown in Tab. 8. This is caused by the randomness, part of the foreground is discarded, resulting performance degradation. However, the important part still has a high probability of being selected, which also guarantees performance to a certain extent. Despite of its degraded performance, the random drop method also has a certain degree of randomness. This is not desirable in practical applications as it may have the chance to lose some safety-critical areas which will cause problems in safety-critical applications.
4 Visualization
Visual analysis. To better understand points pruned by our magnitude criterion, we visualize point clouds before and after pruning. The comparison results are shown in Fig. 6(a), the original images and the pruned images are provided respectively. We observe that most of the foreground points are preserved. For the background areas, points that fall in vertical structures, such as light, poles, and trees, are also preserved as they tend to be hard negatives, and easily confused with foreground objects. These points require a deep neural network with a certain capability to process in order to recognize them as background. In contrast, background points in flat structures such as road points are largely removed because they are easily identifiable redundant points.
Why foreground points with high feature magnitude? We visualize magnitude in BEV for a clearer comparison, as shown in Fig. 6(b), It is not difficult to see that the feature magnitude of most foreground points is relatively large (vehicles). In contrast, the simple background points (ground) show in small magnitude, which are easier to be distinguished by the network. To gain more insights into why high feature magnitude corresponds to the above patterns, we conjecture that this is caused by the training objective in 3D object detection. When training a 3D object detection model, the focal loss is adopted as default in 3D object detection. When we look closer at the focal loss, it will incur a loss on positive samples and hard negatives while easy negatives are removed from the loss. Thus, this will generate gradients in the direction that can incur an update of features for areas with positive samples and hard negatives. This can eventually make a difference in their feature magnitudes in comparison with areas for easy negatives which are less frequently considered in the optimization objective.
Concluding Remarks
In this paper, we investigate the spatial redundancy in sparse CNNs and propose a new simple and efficient convolution operator SPSS-Conv with two variants, i.e., SPSS-ConV and SPRS-Conv. Extensive experiments have proved the following conclusions: 1) Magnitude can serve as a good indicator for sparse CNNs to dynamically identify the informative points in the diverse scenes. 2) The overwhelming background region in the 3D scene results in spatial redundancy, which can be pruned by our method without adversely affecting the performance. 3) Selectively inflating the elements near the region of interest during the down-sampling process not only saves computation but also keeps the necessary information intact. 4) The proposed SPSS-Conv can be effortlessly applied to generic sparse CNNs without specific structural constraints. We hope our work can provide new thoughts for inspiring future research in the community.
Acknowledgement
This work has been supported by Hong Kong Research Grant Council - Early Career Scheme (Grant No. 27209621), HKU Startup Fund, and HKU Seed Fund for Basic Research.