LidarMultiNet: Towards a Unified Multi-Task Network for LiDAR Perception

Dongqiangzi Ye, Zixiang Zhou, Weijia Chen, Yufei Xie, Yu Wang, Panqu Wang, Hassan Foroosh

Introduction

LiDAR plays a major role in the field of autonomous driving. With the release of several large-scale multi-sensor datasets, (e.g. the Waymo Open Dataset (Sun et al. 2020) and the nuScenes (Caesar et al. 2020)) datasets, collected in real self-driving scenarios, LiDAR-based perception algorithms have significantly advanced in recent years. Thanks to the advancement of sparse convolution (Yan, Mao, and Li 2018; Choy, Gwak, and Savarese 2019), voxel-based LiDAR perception methods (Yin, Zhou, and Krähenbühl 2021) have become predominant on major 3D object detection and semantic segmentation benchmarks, and outperform their point-based, pillar-based, or projection-based counterparts (Fan et al. 2021; Lang et al. 2019; Qi et al. 2017) by a large margin in terms of both accuracy and efficiency. In voxel-based LiDAR perception networks, standard 3D sparse convolution is usually used in tandem with submanifold sparse convolution (Graham and van der Maaten 2017). Since standard 3D sparse convolution dilates the sparse features and increases the number of active sites, it is usually only applied as downsampling layer at each stage of the encoder followed by the submanifold sparse convolution layers. The submanifold sparse convolution maintains the number of active sites but it limits the information flow (Chen et al. 2022b) and the receptive field. However, a large receptive field is necessary to exploit the global contextual information, which is critical for 3D segmentation tasks.

In LiDAR-based perception, 3D object detection, semantic segmentation, and panoptic segmentation are usually implemented in distinct and specialized network architectures (Yin, Zhou, and Krähenbühl 2021; Zhang et al. 2020; Zhou, Zhang, and Foroosh 2021; Zhu et al. 2021b; Cheng et al. 2021), which are task-specific and difficult to adapt to other LiDAR perception tasks. Multi-task networks (Teichmann et al. 2018; Feng et al. 2021), unify closely-related tasks by sharing the weights and computation among them, and therefore expected to improve the performance of individual tasks while reducing the overall computational cost. However, so far prior LiDAR multi-task networks have been underperforming compared to their single-task counterparts and have been failing to demonstrate state-of-the-art performance (Feng et al. 2021). As a result, single-task networks are still predominant in major LiDAR perception benchmarks. In this paper, we bridge the gap between the performance of single LiDAR multi-task networks and multiple independent task-specific networks. Specifically, we propose to unify 3D semantic segmentation, 3D object detection, and panoptic segmentation in a versatile network that exploits the synergy between these tasks and achieves state-of-the-art performance, as shown in Figure 1.

Our main contributions are four-fold, and are summarized below:

We present a novel voxel-based LiDAR multi-task network that unifies three major LiDAR perception tasks and can be extended for new tasks with little increase in the computational cost by adding more task-specific heads.

We propose a Global Context Pooling (GCP) module to improve the global feature learning in the encoder-decoder network based on 3D sparse convolution.

We introduce a second-stage refinement module to refine the first-stage semantic segmentation of the foreground thing classes and produce accurate panoptic segmentation results.

We demonstrate start-of-the-art performance for LidarMultiNet on 5 major LiDAR benchmarks. Notably, LidarMultiNet reaches the official \nth1 place in the Waymo 3D semantic segmentation challenge 2022. LidarMultiNet reaches the highest mAPH L2 for a single model on the Waymo 3D object detection benchmark. On the nuScenes semantic segmentation and panoptic segmentation benchmarks, LidarMultiNet outperforms the previously published state-of-the-art methods. On the nuScenes 3D object detection benchmark, LidarMultiNet sets a new standard for state-of-the-art performance in LiDAR-only non-ensemble methods.

Related Work

LiDAR Detection and Segmentation One key challenge for LiDAR perception is how to efficiently encode the large-scale sparsely distributed point cloud into a uniform feature representation. The common practice is transforming the point cloud into a discretized 3D or 2D map through a 3D voxelization (Zhou and Tuzel 2018; Zhu et al. 2021b), Bird’s Eye View (BEV) projection (Yang, Luo, and Urtasun 2018; Lang et al. 2019; Zhang et al. 2020), or range-view projection (Wu et al. 2018; Sun et al. 2021). State-of-the-art LiDAR 3D object detectors (Yin, Zhou, and Krähenbühl 2021) typically project the 3D sparse tensor into a dense 2D BEV feature map and perform the detection on the BEV space. In contrast, LiDAR segmentation requires predicting the point-wise labels, hence a larger features map is needed to minimize the discretization error when projecting the voxel labels back to the points. Many methods (Tang et al. 2020; Xu et al. 2021; Ye et al. 2021b) also combine the point-level features with voxel features to retain the fine-grained features in a multi-view fusion manner.

In LiDAR-based 3D object detection, anchor-free detectors (Yin, Zhou, and Krähenbühl 2021) are predominant on major detection benchmarks and widely adopted for their efficiency. Our LidarMultiNet adopts the anchor-free 3D detection heads, which are attached to its 2D branch.

A second stage is often used in the detection framework (Shi et al. 2020; Yin, Zhou, and Krähenbühl 2021; Li, Wang, and Wang 2021; Sheng et al. 2021) to improve the detection accuracy through an RCNN-style network. It processes each object separately by extracting the features based on the initial bounding box prediction for refinement. LidarMultiNet adopts a second segmentation refinement stage based on the detection and segmentation results of the first stage.

LiDAR Panoptic Segmentation Recent LiDAR panoptic segmentation methods (Zhou, Zhang, and Foroosh 2021; Hong et al. 2021; Razani et al. 2021) usually derive from the well-studied segmentation networks (Zhang et al. 2020; Zhu et al. 2021b; Cheng et al. 2021) in a bottom-up fashion. This is largely due to the loss of height information in the detection networks, which makes them difficult to adjust the learned feature representation to the segmentation task. This results in two incompatible designs for the best segmentation (Xu et al. 2021) and detection (Yin, Zhou, and Krähenbühl 2021) methods. According to (Fong et al. 2022), end-to-end LiDAR panoptic segmentation methods still underperform compared to independently combined detection and segmentation models. In this work, our model can perform simultaneous 3D object detection and semantic segmentation and trains the tasks jointly in an end-to-end fashion.

Multi-Task Network Multi-task learning aims to unify multiple tasks into a single network and train them simultaneously in an end-to-end fashion. MultiNet (Teichmann et al. 2018) is a seminal work of image-based multi-task learning that unifies object detection and road understanding tasks in a single network. In LiDAR-based perception, LidarMTL (Feng et al. 2021) proposed a simple and efficient multi-task network based on 3D sparse convolution and deconvolutions for joint object detection and road understanding. In this work, we unify the major LiDAR-based perception tasks in a single, versatile, and strong network.

LidarMultiNet

The main architecture of LidarMultiNet is illustrated in Figure 2. A voxelization step converts the original unordered LiDAR points to a regular voxel grid. A Voxel Feature Encoder (VFE) consisting of a Multi-Layer Perceptron (MLP) and max pooling layers is applied to generate enhanced sparse voxel features, which serve as the input to the 3D sparse U-Net architecture. Lateral skip-connected features from the encoder are concatenated with the corresponding voxel features in the decoder. A Global Context Pooling (GCP) (Ye et al. 2022) module with a 2D multi-scale feature extractor bridges the last encoder stage and the first decoder stage. 3D segmentation head is attached to the decoder and outputs voxel-level predictions, which can be projected back to the point level through the de-voxelization step. Heads of BEV tasks, such as 3D object detection, are attached to the 2D BEV branch. Given the detection and segmentation results of the first stage, the second stage is applied to refine semantic segmentation and generate panoptic segmentation results.

The 3D encoder consists of 4 stages of 3D sparse convolutions with increasing channel width. Each stage starts with a sparse convolution layer followed by two submanifold sparse convolution blocks. The first sparse convolution layer has a stride of 2 except at the first stage, therefore the spatial resolution is downsampled by 8 times in the encoder. The 3D decoder also has 4 symmetrical stages of 3D sparse deconvolution blocks but with decreasing channel width except for the last stage. We use the same sparse convolution key indices between the encoder and decoder layers to keep the same sparsity of the 3D voxel feature map.

For the 3D object detection task, we adopt the detection head of the anchor-free 3D detector CenterPoint (Yin, Zhou, and Krähenbühl 2021) and attach it to the 2D multi-scale feature extractor. Besides the detection head, an additional BEV segmentation head also can be attached to the 2D branch of the network, providing coarse segmentation results and serving as an auxiliary loss during the training.

Global Context Pooling

3D sparse convolution drastically reduces the memory consumption of the 3D CNN for the LiDAR point cloud data, but it generally requires the layers of the same scale to retain the same sparsity in both encoder and decoder. This restricts the network to use only submanifold convolution (Graham, Engelcke, and Van Der Maaten 2018) in the same scale. However, submanifold convolution cannot broadcast features to isolated voxels through stacking multiple convolution layers. This limits the ability of CNN to learn long-range global information. Inspired by the Region Proposal Network (RPN) (Ren et al. 2015) in the 3D detection network, we design a Global Context Pooling (GCP) module to extract large-scale information through a dense BEV feature map. On the one hand, GCP can efficiently enlarge the receptive field of the network to learn global contextual information for the segmentation task. On the other hand, its 2D BEV dense feature can also be used for 3D object detection or other BEV tasks, by attaching task-specific heads with marginal additional computational cost.

Benefiting from GCP, our architecture could significantly enlarge the receptive field, which plays an important role in semantic segmentation. In addition, the BEV feature maps in GCP can be shared with other tasks (eg. object detection) simply by attaching additional heads with slight increase of computational cost. By utilizing the BEV-level training like object detection, GCP can enhance the segmentation performance furthermore.

Multi-task Training and Losses

During training, the BEV segmentation head is supervised with LBEV\mathcal{L}_{BEV}, a dense loss consisting of cross-entropy loss and Lovasz loss: LBEV=Lcebev+LLovaszbev\mathcal{L}_{BEV}=\mathcal{L}_{ce}^{bev}+\mathcal{L}_{Lovasz}^{bev}.

Our network is trained end-to-end for multiple tasks. Similar to (Feng et al. 2021), we define the weight of each component of the final loss based on the uncertainty (Kendall, Gal, and Cipolla 2018) as follows:

where σi\mathcal{\sigma}_{i} is the learned parameter representing the degree of uncertainty in taskitask_{i}. The more uncertain the taskitask_{i} is, the less Li\mathcal{L}_{i} contributes to Ltotal\mathcal{L}_{total}. The second part can be treated as a regularization term for σi\mathcal{\sigma}_{i} during training.

Instead of assigning an uncertainty-based weight to every single loss, we first group the losses belonging to the same task with fixed weights. The resulting three task-specific losses (i.e., LSEG\mathcal{L}_{SEG}, LDET\mathcal{L}_{DET}, LBEV\mathcal{L}_{BEV}) are then combined using weights defined based on the uncertainty:

Second-stage Refinement

Coarse panoptic segmentation result can be obtained directly by fusing the first-stage semantic segmentation and object detection results, i.e., assigning a unique ID to the points classified as one of the foreground thing classes within a 3D bounding box. However, the points within a detected bounding box can be misclassified as multiple classes due to the lack of spatial prior knowledge, as shown in Figure 5. In order to improve the spatial consistency for the thing classes, we propose a novel point-based approach as the second stage to refine the first-stage segmentation and provide accurate panoptic segmentation.

We merge the 2nd-stage predictions with the 1st-stage semantic scores to generate the final semantic segmentation predictions L^sem\hat{L}_{sem}. To refine segmentation score S2nd={rsi∣rsi∈(0,1)Kthing+1}i=1NS_{2nd}=\{{rs}_{i}|{rs}_{i}\in{(0,1)}^{{K}_{thing}+1}\}_{i=1}^{N}, we combine the point-wise mask scores with their corresponding box-wise class scores as follows:

where Kthing{K}_{thing} denotes the number of thing classes, ∅\emptyset denotes the rest stuff classes which would not be refined in the 2nd stage, Spoint={spi∣spi∈(0,1)}i=1N{S}_{point}=\{{sp}_{i}|{sp}_{i}\in(0,1)\}_{i=1}^{N} is the point-wise mask scores, and Sbox={sbi∣sbi∈(0,1)Kthing+1}i=1B{S}_{box}=\{{sb}_{i}|{sb}_{i}\in{(0,1)}^{{K}_{thing}+1}\}_{i=1}^{B} is the box classification scores. NN and BB denote the number of points and boxes.

In addition, the points not in any boxes can be considered as S2nd(∅)=1S_{2nd}(\emptyset)=1, which means their scores are the same as the 1st-stage scores. We then further combine the refined scores with the 1st-stage scores as follows:

where ϕ\phi denotes the index where points are not in any boxes, and S1st={sfi∣sfi∈(0,1)K}i=1N{S}_{1st}=\{{sf}_{i}|{sf}_{i}\in{(0,1)}^{K}\}_{i=1}^{N} is the 1st stage scores.

The scores SfinalS_{final} are used to generate the semantic segmentation results L^sem\hat{L}_{sem} through finding the class with the maximum score. It is intuitive to infer the final panoptic segmentation results through the 1st-stage boxes and the final semantic segmentation results SboxS_{box} and L^sem\hat{L}_{sem}. First, we extract points for a box where points and the box have the same semantic category. Then the extracted points will be assigned a unique index as the instance id for the panoptic segmentation.

Experiments

In this section, we perform extensive tests of the proposed LidarMultiNet on five major benchmarks of the large-scale Waymo Open Dataset (Sun et al. 2020) (3D Object Detection and 3D Semantic Segmentation) and nuScenes dataset (Caesar et al. 2020; Fong et al. 2022) (Detection, LiDAR Segmentation, and Panoptic Segmentation).

Waymo Open Dataset (WOD) contains 1150 sequences in total, split into 798 in the training set, 202 in the validation set, and 150 in the test set. Each sequence contains about 200 frames of LiDAR point cloud captured at 10 FPS with multiple LiDAR sensors. Object bounding box annotations are provided in each frame while the 3D semantic segmentation labels are provided only for sampled frames. WOD uses Average Precision Weighted by Heading (APH) as the main evaluation metric for the detection task. There are two levels of difficulty, LEVEL_2 (L2) is assigned to examples where either the annotators label as hard or if the example has less than 5 LiDAR points, while LEVEL_1 (L1) is assigned to the rest of the examples. Both L1 and L2 examples participate in the computation of the primary metric mAPH L2.

For the semantic segmentation task, we use the v1.3.2 dataset, which contains 23,691 and 5,976 frames with semantic segmentation labels in the training set and validation set, respectively. There are a total of 2,982 frames in the final test set. WOD has semantic labels for a total of 23 classes, including an undefined class. Intersection Over Union (IOU) metric is used as the evaluation metric.

NuScenes contains 1000 scenes with 20 seconds duration each, split into 700 in the training set, 150 in the validation set, and 150 in the test set. The sensor suite contains a 32-beam LiDAR with 20Hz capture frequency. For the object detection task, the key samples are annotated at 2Hz with ground truth labels for 10 foreground object classes (thing). For the semantic segmentation and panoptic segmentation tasks, every point in the keyframe is annotated using 6 more background classes (stuff) in addition to the 10 thing classes. NuScenes uses mean Average Precision (mAP) and NuScenes Detection Score (NDS) metrics for the detection task, mIoU and Panoptic Quality (PQ) (Kirillov et al. 2019) metrics for the semantic and panoptic segmentation. Note that nuScenes panoptic task ignores the points that are included in more than one bounding box. As a result, the mIoU evaluation is typically higher than using full semantic labels.

Implementation Details

On the Waymo Open Dataset, the point cloud range is set to [−75.2m,75.2m][-75.2m,75.2m] for xx axis and yy axis, and [−2m,4m][-2m,4m] for the zz axis, and the voxel size is set to (0.1m,0.1m,0.15m)(0.1m,0.1m,0.15m). Following (Yin, Zhou, and Krähenbühl 2021), we transform the past two LiDAR frames using the vehicle’s pose information and merge them with the current LiDAR frame to produce a denser point cloud and append a timestamp feature to each LiDAR point. Points of past LiDAR frames participate in the voxel feature computation but do not contribute to the loss calculation.

On the nuScenes dataset, the point cloud range is set to [−54m,54m][-54m,54m] for xx axis and yy axis, and [−5m,3m][-5m,3m] for the zz axis, and the voxel size is set to (0.075m,0.075m,0.2m)(0.075m,0.075m,0.2m). Following the common practice (Caesar et al. 2020; Yin, Zhou, and Krähenbühl 2021) on nuScenes, we transform and concatenate points from the past 9 frames with the current point cloud to generate a denser point cloud. Following (Yin, Zhou, and Krähenbühl 2021), we apply separate detection heads in the detection branch for different categories.

During training, we employ data augmentation which includes standard random flipping, and global scaling, rotation and translation. We also adopt the ground-truth sampling (Yan, Mao, and Li 2018) with the fade strategy (Wang et al. 2021). We train the models using AdamW (Loshchilov and Hutter 2017) optimizer with one-cycle learning rate policy (Gugger 2018), with a max learning rate of 3e-3, a weight decay of 0.01, and a momentum ranging from 0.85 to 0.95. We use a batch size of 2 on each of the 8 A100 GPUs. For the one-stage model, we train the models from scratch for 20 epochs. For the two-stage model, we freeze the 1st stage and finetune the 2nd stage for 6 epochs.

Waymo Open Dataset Results

We tested the performance of LidarMultiNet on the WOD 3D Semantic Segmentation Challenge. Since the semantic segmentation challenge only considers semantic segmentation accuracy, our model is trained with focus on semantic segmentation, while object detection and BEV segmentation both serve as the auxiliary tasks. Since there is no runtime constraint, most participants employed the Test-Time Augmentation (TTA) and model ensemble to further improve the performance of their methods. For details regarding the TTA and ensemble, please refer to the supplementary material. Table 1 is the final WOD semantic segmentation leaderboard and shows that our LidarMultiNet (Ye et al. 2022) achieves a mIoU of 71.13 and ranks the \nth1 place on the leaderboardhttps://waymo.com/open/challenges/2022/3d-semantic-segmentation, accessed on August 06, 2022., and also has the best IoU for 15 out of the total 22 classes. Note that our LidarMultiNet uses only LiDAR point cloud as input, while some other entries on the leaderboard (e.g. SegNet3DV2) use both LiDAR points and camera images and therefore require running additional 2D CNNs to extract image features.

For a better reference, we also test the result of LidarMultiNet that is trained with both detection and segmentation as the main tasks. LidarMultiNet reaches a mIoU of 69.69 on WOD 3D segmentation test set without TTA and model ensemble.

Ablation Study on the 3D Semantic Segmentation Validation Set

We ablate each component of the LidarMultiNet and the results on the 3D semantic segmentation validation set are shown in Table 2. Our baseline network reaches a mIoU of 69.90 on the validation set. On top of this baseline, multi-frame input (i.e., including the past two frames) brings a 0.59 mIoU improvement. The GCP further improves the mIoU by 0.94. The auxiliary losses, (i.e., BEV segmentation and 3D object detection) result in a total improvement of 0.63 mIoU, and the 2nd-stage improves the mIoU by 0.34, forming our best single model on the WOD validation set. TTA and ensemble further improve the mIoU to 73.05 and 73.78, respectively.

Evaluation on the 3D Object Detection Benchmark

To demonstrate that LidarMultiNet can outperform single-task models on both detection and segmentation tasks, we tested it on the WOD 3D object detection benchmark and compared with state-of-the-art 3D object detection methods. The model is trained with both detection and segmentation as the main tasks, and its detection head is trained for detecting three classes, (i.e. vehicle, pedestrian, and cyclist). The model is trained for 20 epochs with the fade strategy, i.e., ground-truth sampling for the object detection task is disabled for the last 5 epochs. The results on WOD test and validation set are shown in Table 3 and Table 4. Our LidarMultiNet method reaches the highest mAPH L2 of 76.35 on the test set for a single model without TTA and outperforms the state-of-art 3D object detectors, including the multi-modal detectors which also leverage camera information. LidarMultiNet also outperforms other multi-frame fusion methods that require more past frames. Moreover, the same LidarMultiNet model reaches a mIoU of 71.93 on the WOD semantic segmentation validation set. In comparison, the other detectors and segmentation methods on the WOD benchmarks are all single-task models dedicated for either object detection or semantic segmentation.

Effect of the Joint Multi-task Training

Table 5 shows an ablation study comparing the first-stage result of the jointly-trained model with independently-trained models. The segmentation-only model removes the detection head, while keeping only the 3D segmentation head and the BEV segmentation head. The detection-only model keeps only the detection head. Compared to the single-task models, the jointly-trained model performs better on both segmentation and detection. In addition, by sharing part of the network among the tasks, the jointly-trained model is also more efficient than directly combining the independent single-task models.

NuScenes Benchmarks

On the nuScenes detection, semantic segmentation, and panoptic segmentation benchmarks, we compare LidarMultiNet with state-of-the-art LiDAR-based methods. The test set and validation set results of all three benchmarks are summarized in Table 6 and Table 7, respectively. As shown in the tables, a single-model LidarMultiNet without TTA outperforms the previous state-of-the-art methods on each task. Combining the independently trained state-of-the-art single-task models (i.e., DRINet++, LargeKernel3D, and Panoptic-PHNet) reaches 80.4 mIOU, 70.5 NDS, and 80.2 PQ on the test set. In comparison, one single LidarMultiNet model without TTA outperforms their combined performance by 1.0% in mIOU, 1.1% in NDS, and 1.3% in PQ.

In summary, to the best of our knowledge, LidarMultiNet is the first time that a single LiDAR multi-task model surpasses previous single-task state-of-the-art methods for the three major LiDAR-based perception tasks.

Effect of the Second-Stage Refinement

An ablation study on the effect of the proposed 2nd stage is shown in Table 8. With the 1st-stage detection and semantic predictions, LidarMultiNet already can get high panoptic segmentation results by directly fusing these two results together. The proposed 2nd stage further improves both semantic segmentation and panoptic segmentation results.

Top-down LiDAR Panoptic Segmentation

CNN-based top-down panoptic segmentation methods have shown competitive performance compared to the bottom-up methods in the image domain. However, most previous LiDAR panoptic segmentation methods (Zhou, Zhang, and Foroosh 2021; Li et al. 2022a) adopt the bottom-up design due to the need of an accurate semantic prediction. And the cumbersome network structures with multi-view or point-level features fusion make them difficult to perform well in the object detection task. On the other hand, thanks to the GCP module and joint training design, LidarMultiNet can reach top performance on both object detection and semantic segmentation tasks. Even without a dedicated panoptic head, LidarMultiNet already outperforms the previous state-of-the-art bottom-up method.

Conclusion

We present the LidarMultiNet , which reached the official \nth1 place in the Waymo Open Dataset 3D semantic segmentation challenge 2022. LidarMultiNet is the first multi-task network to achieve state-of-the-art performance on all five major large-scale LiDAR perception benchmarks. We hope our LidarMultiNet can inspire future works in the unification of all LiDAR perception tasks in a single, versatile, and strong multi-task network.

Appendix A Appendix

Hyperparameters of LidarMultiNet are summarized in Table 9. Size of the input features of VFE is set to 16. The 3D encoder in our network has 4 stages of 3D sparse convolutions with increasing channel width 32, 64, 128, 256. The downsampling factor of the encoder is 8. The 3D decoder has 4 symmetrical stages of 3D sparse deconvolution blocks with decreasing channel widths 128, 64, 32, 32. 2D depth and width of the GCP module are set to 6, 6 and 128, 256, respectively.

Test-Time Augmentation and Ensemble

In order to further improve the performance on the WOD semantic segmentation benchmark, we apply Test-Time Augmentation (TTA). Specifically, we use flipping with respect to xzxz-plane and yzyz-plane, [0.950.95, 1.051.05] for global scaling, and [±22.5°\pm 22.5\degree, ±45°\pm 45\degree, ±135°\pm 135\degree, ±157.5°\pm 157.5\degree, 180°180\degree] for yaw rotation. Besides, we found pitch and roll rotation were helpful in the segmentation task, and we use ±8°\pm 8\degree for pitch rotation and ±5°\pm 5\degree for roll rotation. In addition, we also apply ±0.2m\pm 0.2m translation along zz-axis for augmentation.

Besides the best single model, we also explored the network design space and designed multiple variants for model ensemble. For example, more models are trained with smaller voxel size (0.05m,0.05m,0.15m)(0.05m,0.05m,0.15m), smaller downsample factor (4×4\times), different channel width (64), without the 2nd stage, or with more past frames (4 sweeps). For our submission to the leaderboard, a total of 7 models are ensembled to generate the segmentation result on the test set.

Runtime and Model Size

We tested the runtime and model size of LidarMultiNet on Nvidia A100 GPU. Compared with single-task models (nuScenes detection: 107ms, segmentation: 112ms, summation: 219ms), our multi-task network shows notable efficiency (nuScenes the 1st stage runtime: 126ms, model size: 135M, the 2nd stage runtime: 9ms, model size: 4M, WOD the 1st stage runtime: 145ms, model size: 131M, the 2nd stage runtime: 18ms, model size: 4M)

Visualizations

Some example qualitative results of LidarMultiNet on the validation sets of Waymo Open Dataset and nuScenes dataset are visualized in Figure 6 and Figure 7, respectively.

References