Focal Sparse Convolutional Networks for 3D Object Detection

Yukang Chen, Yanwei Li, Xiangyu Zhang, Jian Sun, Jiaya Jia

Introduction

A key challenge in 3D object detection is to learn effective representations from the unstructured and sparse 3D geometric data such as point clouds. In general, there are two ways for this job. The first is to process point clouds directly, based on PointNet++ networks. However, the neighbour sampling and grouping operations are time-consuming. This makes it improper for large-scale autonomous driving scenes that require real-time efficiency. The The second is to convert point clouds into voxelizations and apply 3D sparse convolutional neural networks (Sparse CNNs) for feature extraction . 3D Sparse CNNs resemble 2D CNNs in structures, including several feature stages and down-sampling operations. They typically consist of regular and submanifold sparse convolutions .

Although regular and submanifold sparse convolutions have been widely used, they have respective limitations. Regular sparse convolution dilates all sparse features. It inevitably burdens models with considerable computations. That is why backbone networks commonly limit its usage only in down-sampling layers . In addition, detectors aim to distinguish target objects from massive background features. But regular sparse convolution reduces the sparsity sharply and blurs feature distinctions.

On the other hand, submanifold sparse convolutions avoid the computation issue by restricting the output feature positions to the input. But it misses necessary information flow, especially for the spatially disconnected features. The above issues on regular and submanifold sparse convolutions limit Sparse CNNs to achieve high representation capability and efficiency. We illustrate the submanifold and regular sparse convolutional operations in Fig. 1.

These limitations originate from the conventional convolution pattern: all input features are treated equally in the convolution process. It is natural for 2D CNNs, and yet is improper for 3D sparse features. 2D convolution is designed for structured data. All pixels in the same layer typically share receptive field sizes. But 3D sparse data is with varying sparsity and importance in space. It is not optimal to handle non-uniform data with uniform treatment. In terms of sparsity, upon the distance to LIDAR sensors, objects present large sparsity variance. In terms of importance, the contribution of features varies with different locations for 3D object detection, e.g., foreground or background. Although 3D object detection is achieved , state-of-the-art methods still rely on RoI (region-of-interest) feature extraction. It corresponds to the idea that we should shoot arrows at the target in the feature extraction of 3D detectors.

In this paper, we propose a general format of sparse convolution by relaxing the conceptual difference between regular and submanifold ones. We introduce two new modules that improve the representation capacity of Sparse CNNs for 3D object detection. The first is focal sparse convolution (Focals Conv). It predicts cubic importance maps for the output pattern of convolutions. Features predicted as important ones are dilated into a deformable output shape, as shown in Fig 1. The importance is learned via an additional convolutional layer, dynamically conditioned on the input features. This module increases the ratio of valuable information among all features. The second is its multi-modal improved version of Focal sparse Convolution with Fusion (named as Focals Conv-F). Upon the LIDAR-only Focals Conv, we enhance importance prediction with RGB features fused, as image features typically contain rich appearance information and large receptive fields.

The proposed modules are novel in two aspects. First, Focals Conv presents a dynamic mechanism for learning spatial sparsity of features. It makes the learning process concentrated on the more valuable foreground data. With the down-sampling operations, valuable information increases in stages. Meanwhile, the large amount of background voxels are removed. Fig. 2 illustrates the learnable feature sparsity, including the common, crowded, and remote objects, where Focals Conv enriches the learned voxel features on the foreground without redundant voxels added in other areas. Second, both modules are lightweight. The importance prediction involves small overhead parameters and computation, as measured in Tab. 1. The RGB feature extraction of Focals Conv-F involves only several layers, instead of heady 2D detection or segmentation models .

The proposed modules of Focals Conv and Focals Conv-F can readily replace their original counterparts in sparse CNNs. To demonstrate the effectiveness, we build the backbone networks on existing 3D object detection frameworks . Our method enables non-trivial enhancement with small model complexity overhead on both the KITTI and nuScenes benchmarks. These results manifest that learnable sparsity with focal points is essential. Without bells and whistles, our approach outperforms state-of-the-art ones on the nuScenes test split .

Convolutional dynamic mechanism adapts the operations conditioned on input data, e.g., deformable convolutions and dynamic convolutions . The key difference is that our approach makes use of the intrinsic sparsity of data. It promotes feature learning to be concentrated on more valuable information. We deem the non-uniform property as a great benefit. We discuss the relations and differences to previous literature in Sec. 2.

Related Work

Dynamic mechanisms have been widely studied in CNNs, due to their advantages of high accuracy and easy adaption in scenarios. We discuss two kinds of related methods, i.e., kernel shape adaption , and input attention mask .

Kernel shape adaption. Kernel shape adaption methods adapt the effective receptive fields of networks. Deformable convolution predicts offsets for feature sampling. Its variant introduces an additional attention mask to modulate features. For 3D feature learning, KPConv learns local offsets for kernel points. MinkowskiNet generalizes sparse convolution to arbitrary kernel shape. Overall, these methods modify the input feature sampling process.

Deformable PV-RCNN applies offset prediction for feature sampling in 3D object detection. In contrast, focal sparse convolution improves the output feature spatial sparsity and makes it learned, helpful for 3D object detection.

Attention mask on input. Methods of seek spatial-wise sparsity for efficient inference. These methods receive dense images and prune unimportant pixels based on attention masks. These methods aim to sparsify dense data while we make use of intrinsic data sparsity. Although SBNet also utilizes the sparse property, it limits application to 2D BEV (bird-eye-views) images, and shares the static masks over all layers in the network. In contrast, our improved convolution is more adaptive and is applicable to related tasks, e.g., 3D instance segmentation .

2 3D Object Detection

LIDAR-only detectors. 3D object detection frameworks usually resemble 2D detectors, e.g., the R-CNN family and the SSD family . The main difference on 2D detectors lies in input encoders. VoxelNet encodes voxel features using PointNet and applies a RPN (region proposal network) . SECOND uses accelerated sparse convolutions and improves efficiency from VoxelNet . VoTr applies transformer architectures to voxels. Various detectors have been presented based on feature encoders. We validate the proposed approach on backbones of frameworks of on multiple datasets .

Completion-based detectors. Completion-based methods form another line of efforts in enriching foreground information. We focus on feature learning instead of point completion. PC-RGNN has a point completion module by a graph neural network. SIENet builds upon PCN for point completion in a two-stage framework. The completion process relies on the prior generated proposals. GSDN expands all features first through transposed convolutions and then by pruning. SPG designs a semantic point generation module for domain adaption 3D object detection. It is applied during data preprocessing, complicating the detection pipelines.

Multi modal fusion. Multi-modal fusion methods use more information than LIDAR-only ones. The KITTI benchmark had been dominated by LIDAR-only methods until PointPainting was proposed. It decorates raw point clouds with the corresponding image segmentation scores. PointAugmenting further replaces the segmentation model with an 2D object detection one . They are both decoration-based methods, which require image feature extraction on off-the-shelf 2D networks, before feeding into 3D detectors. Although promising results are achieved by these methods, the overall inference pipelines are complicated. Our multi-modal focal sparse convolution differs from the above methods in two aspects. First, we only require several jointly trained layers for image feature extraction, rather than the heavy segmentation or detection models. Second, we only strengthen the predicted important features, instead of the uniform decoration for all LIDAR features.

Focal Sparse Convolutional Networks

In this section, we first review the formulation of sparse convolution in Sec. 3.1. Then, the proposed focal sparse convolution and its multi-modal extension will be elaborated in Sec. 3.2 and Sec. 3.3. We finally introduce the resulting focal sparse convolutional networks in Sec. 3.4.

where kk enumerates all discrete locations in the kernel space Kd\mathnormal{K}^{d}. pˉk=p+k\bar{p}_{k}=p+k is the corresponding location around center pp, where kk is an offset distance from pp.

On this condition, the formulation becomes regular sparse convolution. It acts at all positions where any voxels exist in its kernel space. It does not skip any information gathering in the total spatial space.

This strategy involves two drawbacks. (i) It introduces considerable computation cost. The number of sparse features is doubled or even tripled, increasing burden for following layers. (ii) We empirically find that continuously increasing the number of sparse features may harm 3D object detection (Tab. 2). Crowded and unpromising candidate features may blur the valuable information. It degrades foreground features and further declines the feature discrimination capacity of 3D object detectors.

2 Focal Sparse Convolution

We factorize this process into three steps: (i) cubic importance prediction, (ii) important input selection, and (iii) dynamic output shape generation.

Cubic importance prediction. A cubic importance map IpI^{p} involves importance for candidate output features around the input feature at position pp. Each cubic importance map shares the same shape Kd\mathnormal{K}^{d} with the main processing convolution kernel weight, e.g.e.g., k3=3×3×3k^{3}=3\times 3\times 3 with the kernel size 3. It is predicted by an additional submanifold sparse convolution with a sigmoid function. The latter steps depend on the predicted cubic importance maps.

where I0pI^{p}_{0} is the center of the cubic importance map at position pp. And τ\tau is a pre-defined threshold (Tab. 3 and 6). Our formulation becomes the regular or submanifold sparse convolution when τ\tau is 0 or 1 respectively. We also find that using top-k ratio to select is an alternative of threshold.

We analyze the dynamic output shape in Tab. 2. For the remaining unimportant features, their output positions are fixed as input, i.e., submanifold. We found that directly removing them or using a fully dynamic manner without preserving them makes the training process unstable.

Supervision manners. In 3D object detection, we have a prior knowledge that foreground objects are more valuable information. Based on this prior, we apply focal loss as an objective loss function to supervise the importance prediction. We construct the objective targets for the centers of feature voxels inside 3D ground-truth boxes. We keep its loss weight as 1 for the generality of our modules.

Additional supervision comes from multiplying the predicted cubic importance maps to output features as attention. It makes the importance prediction branch differentiable naturally. It shares motivation with the kernel weight sparsification methods in the area of model compression. We empirically show that this attention manner benefits the performance for minor classes, e.g., Pedestrian and Cyclist on KITTI (investigated in Tab. 4).

3 Fusion Focal Sparse Convolution

We provide a multi-modal version of focal sparse convolution, as illustrated in Fig. 3 (via dashed lines). This extension is conceptually simple but effective. We extract RGB features from images and align LIDAR features to them. The extracted features are fused to input and important output sparse features in focal sparse convolution.

Feature extraction. The fusion module is lightweight. It contains a conv-bn-relu layer and a max-pooling layer. It down-samples the input image to 1/4 resolutions. It is followed by 3 conv-bn-relu layers with residual connection . The channel number is then reduced to be consistent with that of sparse features, with an MLP layer. This facilitates a simple summation of multi-modal features.

Feature alignment. A common issue during fusion is misalignment in the 3D-to-2D projection. Point cloud data is commonly processed by transformation and augmentation. Transformations include flip, re-scale, rotation, translation. The typical augmentation is ground-truth sampling , copying paste objects from other scenes. For these invertible transformations, we reverse the coordinates of sparse features with the recorded transformation parameters . For ground-truth sampling, we copy the corresponding 2D objects onto images. Rather than using an additional segmentation model or mask annotations , we directly crop objects in bounding boxes for simplification.

Fusion manners. The aligned RGB features are directly fused to sparse features in summation, as they share the same channel numbers. Although other fusion methods, e.g.e.g., concatenation or cross-attention, can be used, we choose the most concise summation for efficiency. The aligned RGB features are fused with sparse features twice in this module. It is first fused to input features for cubic importance prediction. Then we fuse RGB features only to important output sparse features, i.e., the first part in Eq. (5), instead of all of them (investigated in Tab. 10).

Overall, the multi-modal layers are lightweight in terms of parameters and fusion strtegies. They are jointly trained with detectors. It provides an efficient and economical solution for the fusion module in 3D object detection.

4 Focal Sparse Convolutional Networks

Both focal sparse convolution and its multi-modal extension can readily replace their counterparts in the backbone networks of 3D detectors. During training, we do not use any special initialization or learning rate settings for the introduced modules. The importance prediction branch is trained via back-propagation through the attention multiplication and objective loss function as introduced in Sec. 3.2.

The backbone networks in 3D object detectors typically consist of one stem layer and 4 stages. Each stage, except the first one, includes a regular sparse convolution with down-sampling and two submanifold blocks. In the first stage, there are one or two sparse convolutional layers. By default, each sparse convolution is followed by batch normalization and ReLU activation.

We validate focal sparse convolution on the backbone networks of existing 3D detectors . We directly apply focal sparse convolution at the last layer of certain stages. We analyze the stages for using our focal sparse convolution in experiments (ablated in Tab. 5 and 10).

Experiments

We conduct ablations and comparisons for Focals Conv and its multi-modal variant. More experiments, such as results on Waymo , are in the supplementary material.

KITTI. The KITTI dataset consists of 7,481 samples and 7,518 testing samples. The training samples are split into a train set with 3,717 samples and a val set with 3,769 samples. Models are commonly evaluated in terms of the mean Average Precision (mAP) metric. mAP is calculated with recall 40 positions (R40). We perform ablation studies with AP3D{}_{\textrm{3D}} (R40) on the val split. We conduct main comparisons with AP3D{}_{\textrm{3D}} (R40) on test split and AP3D{}_{\textrm{3D}} (R11) on the val split. For the optional multi-modal settings, RGB features are extracted from single front-view for fusion.

nuScenes. The nuScenes is a large-scale dataset, which contains 1,000 driving sequences in total. It is split into 700 scenes for training, 150 scenes for validation, and 150 scenes for testing. It is collected using a 32-beam synced LIDAR and 6 cameras with the complete 360o environment coverage. In evaluation, the main metrics are mAP and nuScenes detection score (NDS). In terms of multi-modal experiments, we use images of 6 views for fusion. For ablation study, models are trained on 14\frac{1}{4} training data and evaluated on the entire validation set, i.e., nuScenes 14\frac{1}{4} split.

Implementation details. In experiments, we validate our modules on state-of-the-art frameworks of PV-RCNN , Voxel R-CNN on KITTI , and CenterPoint on nuScenes . In LIDAR-only experiments, we apply Focals Conv in the first three stages of backbone networks. In multi-modal cases, we apply Focals Conv-F only in the first stage of the backbone network, for affordable memory and inference cost. We set the importance threshold τ\tau to 0.5. We keep other settings intact. More experimental details are provided in the supplementary material.

2 Ablation Studies

Improvements on KITTI. We first evaluate our methods over PV-RCNN in Tab. 1, as it is a high-performance, multi-class, and open-sourced framework. In Tab. 1, the 1st and 2nd lines show the reported results and results tested from the released model. We take the latter as the baseline. Focal S-Conv and Focals Conv-F achieve non-trivial improvement over this strong baseline.

Dynamic output shape. In Focals Conv, the output shape from every single voxel is dynamically determined by the predicted importance maps. We ablate this by fixing output shapes as regular dilation, without any other change. Tab. 2 shows that dilating all sparse features is harmful. It dramatically increases the number of unpromising voxel features.

Importance sampling. Focals Conv selects sparse features that need dilation with predicted importance. To ablate this module, we replace the importance selection (the important input selection step) with a random sample in Tab. 3 without other changes. It shows that large performance drop occurs without the guidance of importance. This validates that the importance prediction is necessary.

Supervision setting. The additional branch in Focals Conv is supervised by both attention multiplication and the objective loss. We ablate them in Tab. 4. Only using objective loss supervision is enough to ensure performance on Car. However, its performance on minor classes, Ped. and Cyc., is not optimal. Attention multiplication is beneficial to Ped. and Cyc. We assume that minor classes cannot get balanced supervision from the objective loss like the long-tailed distribution. In contrast, attention multiplication is object-agnostic, relaxing the imbalance to some degree.

Stages for using focal sparse convolution. Tab. 5 shows results of using Focals Conv in different numbers of stages. (1) Applying Focals Conv in the first stage, which already obtains clear improvement. The performance enhances as the used stage increases until all stages are involved. Since Focals Conv adjusts output sparsity, it is reasonable to be used in early stages that make effects on subsequent feature learning. The spatial feature space in the last stage is down-sampled to a very limited size, which might not be large enough for sparsity adaptation. Empirically, usage in the last layer of the first three stages is the best choice. It is thus used as the default setting in our experiments.

Importance threshold. We ablate the importance threshold τ\tau used in Focals Conv in Tab. 6. We run experiments with this value ranging from 0.1 to 0.9 and interval 0.2, without other change of settings. The accuracy AP3D (R40) on Car serves as the metric in this ablation. The performance is stable as the threshold value τ\tau varies.

Improvements over multi-modal baseline on nuScenes. We evaluate our multi-modal Focals Conv on the nuScenes 1/4 dataset. More improvement is presented in Tab. 9. We build a multi-modal CenterPoint baseline by fusing image features to the same fusion layer used in our methods, with the same fusion and feature extraction layers. This multi-modal CenterPoint enhances the LIDAR-only baseline from 56.1% to 59.0% mAP. Focals Conv-F improves to 61.7% mAP on this strong baseline.

Use stages and fusion scope for Focals Conv-F. We ablate the usage stages and fusion scope for Focals Conv-F in Tab. 10. Fusion scope is the scope of sparse features to fuse with RGB features at the output of Focals Conv-F. It shows that fusion in the early stages is beneficial, and becomes adverse in the last two stages. Imp. means only fusing onto important output features (judged by importance maps). When fusing in the first stage, it is better to fuse on important features, instead of all of them, making representation discriminative.

Model complexity and runtime. We report the model complexity and runtime comparisons in Tab. 1 and 9. The runtimes are evaluated on the same GPU machine. Focals Conv and its multi-modal variant only add a small overhead to model parameters and computation, on KITTI . This indicates that the performance improvement comes from the model capacity of sparsity learning, instead of increasing model sizes. On nuScenes , the overall runtime rises from 93 ms to 159 ms. But parameters are still limited. It is a common limitation in multi-view fusion methods. The multi-modal baseline also requires 145 ms. The reason is that there are 6-view images to process per frame.

3 Main Results

KITTI. We compare our Focals Conv modules upon Voxel R-CNN with previous state-of-the-art methods on both the KITTI test and val split. In Tab. 7, we compare with both LIDAR-only and multi-modal methods. The original Voxel R-CNN is comparable to PV-RCNN and is inferior to Pyramid-PV and VoTr-TSD . Focals Conv improves it to surpass these two new methods. Using Focals Conv-F, the multi-modal Voxel R-CNN achieves 82.28% AP3D on the KITTI test split. Tab. 8 shows comparisons on KITTI val split in AP3D in recall 11 positions. Focals Conv and Focals Conv-F enhance this leading result to 84.93% and 85.22% respectively in Car class.

nuScenes. On the nuScenes dataset, we evaluate our models on the test server and compare them with both LIDAR-only and multi-modal methods, as in Tab. 11. Focals Conv improves CenterPoint by a large margin to 63.8% mAP. Multi-modal methods present much better performance than LIDAR-only methods on the nuScenes dataset. CenterPoint v2⋆ includes PointPainting , Cascade R-CNN instance segmentation models pre-trained on nuImages, and five-model ensembling. As the testing augmentations are not unified or stated in previous methods, we provide two results of our final model. Focals Conv-F achieves 67.8% mAP and 71.8% mAP without any ensembling or testing augmentation. Focals Conv-F ‡ further achieves 70.1% mAP and 73.6% NDS with test-time augmentations . Both results outperform previous methods.

Conclusion and Discussion

This paper presents a focal sparse convolution and a multi-modal extension, which are simple and effective. They are end-to-end solutions for LIDAR-only and multi-modal 3D object detection. For the first time, we show that the learned sparsity with focal points is essential for 3D object detectors. Notably, focal and fusion sparse CNNs achieve leading performance on the large-scale nuScenes.

Limitations. In the multi-modal 3D detection that requires multiple views, e.g., 6 high-resolution images per frame in nuScenes , computation cost increases, although the image branch is already largely simplified.

Boarder Impacts. The proposed method replies on the sparsity of data distribution. It might reflect biases in data collection, including the ones of negative societal impacts.

Acknowledgements. This work is in part supported by The National Key Research and Development Program of China (No. 2017YFA0700800) and Beijing Academy of Artificial Intelligence (BAAI).

References

Appendix

Appendix A More Implementation Details

Our implementation is based on the open-sourced OpenPCDet , and the released code of CenterPoint .

KITTI. The 3D object detectors in this work convert point clouds into voxels as input data. On the KITTI dataset, the range of point clouds is clipped into [0, 70.4m] for X axis, [-40m,40m] for Y axis, and m for Z axis. The voxelization size for input is (0.05m, 0.05m, 0.1m).

nuScenes. On the nuScenes , the detection range is set to [-54m, 54m] for both X and Y axes, and [-5m, 3m] for the Z axis. The voxel size is set as (0.075m, 0.075m, 0.2m).

A.2 Data Augmentations

KITTI. On the KITTI dataset, data transformation and augmentations include random flipping, global scaling, global rotation, and ground-truth (GT) sampling . The random flipping is conducted along the X axis. The global scaling factor is sampled from 0.95 to 1.05. The global rotation is conducted around the Z axis. The rotation angle is sampled from -45o and 45o. The ground-truth sampling is to copy-paste some new objects from other scenes to the current training data, which enriches objects in the environments. For the multi-modal setting, we do not transform images with the corresponding operations, except ground-truth sampling. We copy-paste the corresponding image crops from other scenes onto the current training images.

nuScenes. On the nuScenes dataset, data augmentations includes random flipping, global scaling, global rotation, GT sampling , and an additional translation. The random flipping is conducted along both X and Y axes. The rotation angle is also randomly sampled in [-45o, 45o]. The global scaling factor is sampled in [0.9, 1.1]. The translation noise is conducted on all three axes, X, Y, and Z, with a factor independently sampled from 0 to 0.5. We also conduct the corresponding point-image GT sampling on the nuScenes. GT sampling is disabled in the last 4 epochs for performance enhancement.

A.3 Training Settings

KITTI. For model training on the KITTI dataset, i.e., PV-RCNN and Voxel R-CNN , we train the network for 80 epochs with the batch size 16. We adopt the Adam optimizer. The learning rate is set as 0.01 and decreases in the cosine annealing strategy. The weight decay is set as 0.01. The momentum is set as 0.9. The gradient norms of training parameters are clipped by 10.

nuScenes. For models trained on the nuScenes datasets, i.e., CenterPoint , we also train the network for 20 epochs with batch size 32. They are also trained with the Adam optimizer. The learning rate is initialized as 1e-3 and decreases in the cosine annealing strategy to 1e-4. The weight decay is set as 0.01. The gradient norms of training parameters are clipped by 35.

Appendix B Backbone Networks

We illustrate the structure of the backbone networks in Fig. S - 4. In this illustration, Reg block and Subm block mean the regular sparse convolutional block and the submanifold sparse convolutional block, respectively. The backbone networks are based on VoxelNet . It contains a stem layer and 4 stages. In the last three stages, Stage 1, 2, and 3, there a regular sparse convolutional block with stride as 2 for down-sampling. There are some detailed differences among different frameworks, as the following.

PV-RCNN and Voxel R-CNN. In the backbones of PV-RCNN and Voxel R-CNN , the channels for the stem and stages, {c0c_{0}, c1c_{1}, c2c_{2}, c3c_{3}, c4c_{4}}, are {16, 16, 32, 64, 64}. The numbers of Subm blocks in these stages, {n1n_{1}, n2n_{2}, n3n_{3}, n4n_{4}}, are {1, 2, 2, 2}. A Reg or Subm block is a conv-bn-relu layer, which includes a regular or submanifold convolution, a batch normalization layer , and a ReLU activation.

CenterPoint. In the backbone network of the CenterPoint detector, the backbone network is larger. The channels for the stem and stages, {c0c_{0}, c1c_{1}, c2c_{2}, c3c_{3}, c4c_{4}}, equal to {16, 16, 32, 64, 128}. The numbers of repeated Subm blocks in these stages, {n1n_{1}, n2n_{2}, n3n_{3}, n4n_{4}}, are {2, 2, 2, 2}. Compared to that in the PV-RCNN and Voxel R-CNN detectors, the Subm block is more complicated in this backbone network. Except the stem, it contains two sequential conv-bn-relu layers, with a residual connection.

B.2 Focal Sparse Convolution Usage

In our approach, the above architecture-level settings are directly inherited from the original PV-RCNN , Voxel R-CNN , and CenterPoint frameworks, without any adjustment, for a fair comparison. In the LIDAR-only task, we insert the Focals Conv in the last layer of Stage 1, 2, and 3. In the multi-modal task, we insert the Focals Conv - F only at the last layer of Stage 1. This relieves the efficiency and memory issues caused by the RGB feature extraction. Note that, in the CenterPoint detectors, it is also used in the last layer, not the total block. In other words, although there are two conv-bn-relu layers in each Subm block in CenterPoint , we only apply it as the last layer. For simplification, we do not double it as a block.

Appendix C Additional Experiments

We report the accuracy for 3D object detection and Bird’s Eye View (BEV) of Focals Conv-F upon Voxel R-CNN on the KITTI dataset in Tab. S - 12. The results are calculated by recall 40 positions with the IoU threshold of 0.7. It performs better than the strong Voxel R-CNN baseline on both APBEV{}_{\textrm{BEV}} and AP3D{}_{\textrm{3D}} in moderate and hard cases. We also provide the Prevision-Recall (PR) curves of Focals Conv-F on KITTI test split in Fig. S - 5.

C.2 Objective Loss Weight

The training of the focal sparse convolutional networks involves the objective loss function. We implement it as a focal loss as in Eq (9).

where i∈Ni\in N enumerates all sparse features in the current feature space. Following the original focal loss , we directly set γ=2\gamma=2 and it works well. For notational convenience, we define piˉ\bar{p_{i}} as follow

where pi∈p_{i}\in is the estimated probability for the class with label yi=1y_{i}=1. It is the estimation that whether the feature ii contributes any foreground objects.

We analyze the loss weight for this objective loss in Tab. S - 13. This ablation study is conducted upon the PV-RCNN detector on the KITTI datasets. The results on AP3D{}_{\textrm{3D}} with 40 recall positions are reported as the metric. We change the loss weight values from {0.1, 0.5, 1.0, 2.0}. It shows that too large or too small loss weight values degrade the results. Loss weights 0.5 and 1.0 present competitive performance. We remain the 1.0 loss weight as a default setting for simplification.

C.3 Improvements on the Waymo Open Dataset.

To show our generalization capacity, we conduct further experiments on Waymo dataset. We use 15\frac{1}{5} training data, following the default setting in the OpenPCdet codebase . As shown in Tab. S - 14, Focals Conv also brings non-trivial improvements on the Waymo dataset.

C.4 Improvements on the nuScenes val split.

Tab. S - 15 presents the improvements over CenterPoint on the nuScenes val split. The CenterPoint baseline in Tab. S - 15 is re-implemented in the same settings to Focals Conv and Focals Conv-F. It shows that both Focals Conv and Focals Conv-F bring non-trivial improvements. Notably, Focals Conv-F improves the plain CenterPoint by 4.9% mAP on the nuScenes val split. We further apply some tricks for performance enhancement, e.g., disabling ground-truth sampling in the last 4 epochs and double-flip testing .

C.5 Accuracy loss on some categories after fusion.

A surprising case is that the multi-modal fusion make the performance stay the same or worse on some popular categories, e.g., Car, Ped, Bar (from Focals Conv to Focals Conv - F in Tab. 11). The improvements over the baseline are consistent on all categories. To analyze this special case, we conduct ablations on augmentations on CenterPoint and the nuScenes 14\frac{1}{4} training set. We find ground-truth sampling (GT Sampl.) is the keypoint. As in Tab. A - 16, when GT Sampl. is used, the performance on some popular categories (e.g., Car, Bus, Ped, Bar) stays the same or worse. In contrast, when we disable GT Sampl. and apply all other transformations (flip, rotation, re-scaling, and translation), all categories are benefited from the fusion. We suppose that this is from the image-level copy-paste in GT Sampl. When other objects are pasted onto images, popular objects inevitably have more chance to be covered by the pasted, which degrades the performance on these categories.

C.6 Ablations on Voxel Size.

We ablate the effects of different voxel sizes upon Focals Conv-F on the nuScenes val split in Tab. S - 17. We change the voxel sizes in X and Y axes from 0.05m to 0.15m, with the interval 0.025m. The overall mAP achieves the best performance at the voxel size (0.075, 0.075, 0.2)m. However, the proper voxel sizes vary across different classes. This phenomenon deserves further analysis or a dynamic mechanism design in the future.

Appendix D Visualizations

We provide additional visual comparisons between the plain network and the focal sparse convolutional networks in Fig. S - 6. It shares the same settings to the Fig. 2 in the paper. These visualizations are based on the PV-RCNN detectors and on the KITTI dataset. In each visualization group, the top figure is the distribution of input voxels. The middle and the bottom figures are from the plain and the focal sparse convolutional networks, respectively. We project the coordinate centers of the output voxel features from the backbone networks onto the 2D image plane. The projection is based on the calibration matrices of KITTI .