TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers
Xuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang, Yilun Chen, Hongbo Fu, Chiew-Lan Tai
Introduction
As one of the fundamental tasks in self-driving, 3D object detection aims to localize a set of objects in 3D space and recognize their categories. Thanks to the accurate depth information provided by LiDAR, early works such as VoxelNet Zhou2018VoxelNetEL and PointPillar Lang2019PointPillarsFE achieve reasonably good results using only point clouds as input. However, these LiDAR-only methods are generally surpassed by the methods using both LiDAR and camera data on large-scale datasets with sparser point clouds, such as nuScenes Caesar2020nuScenesAM and Waymo Sun2020ScalabilityIP. LiDAR-only methods are surely insufficient for robust 3D detection due to the sparsity of point clouds. For example, small or distant objects are difficult to detect in LiDAR modality. In contrast, such objects are still clearly visible and distinguishable in high-resolution images. The complementary roles of point clouds and images motivate researchers to design detectors utilizing the best of the two worlds, i.e., multi-modal detectors.
Existing LiDAR-camera fusion methods roughly fall into three categories: result-level, proposal-level, and point-level. The result-level methods, including FPointNet Qi2018FrustumPF and RoarNet Shin2019RoarNetAR, use off-the-shelf 2D detectors to seed 3D proposals, followed by a PointNet Qi2017PointNetDL for object localization. The proposal-level fusion methods, including MV3D Chen2017Multiview3O and AVOD Ku2018Joint3P, perform fusion at the region proposal level by applying RoIPool Ren2015FasterRT in each modality for shared proposals. These coarse-grained fusion methods show unsatisfactory results since rectangular regions of interest (RoI) usually contain lots of background noise. Recently, a majority of approaches have tried to do point-level fusion and achieved promising results. They first find a hard association between LiDAR points and image pixels based on calibration matrices, and then augment LiDAR features with the segmentation scores Vora2020PointPaintingSF; Xu2021FusionPaintingMF or CNN features Sindagi2019MVXNetMV; Meyer2019SensorFF; Huang2020EPNetEP; Wang2021PointAugmentingCA; Zhang2020MultiModalityCA of the associated pixels through point-wise concatenation. Similarly, Liang2019MultiTaskMF; Liang2018DeepCF; Xie2020PIRCNNAE; Yoo20203DCVFGJ first project a point cloud onto the bird’s eye view (BEV) plane and then fuse the image features with the BEV pixels.
Despite the impressive improvements, these point-level fusion methods suffer from two major problems, as shown in Fig. 1. First, they simply fuse the LiDAR features and image features through element-wise addition or concatenation, and thus their performance degrades seriously with low-quality image features, e.g., images in bad illumination conditions. Second, finding the hard association between sparse LiDAR points and dense image pixels not only wastes many image features with rich semantic information, but also heavily relies on high-quality calibration between two sensors, which is usually hard to acquire due to the inherent spatial-temporal misalignment Zhao2021LIFSegLA.
To address the shortcomings of the previous fusion approaches, we introduce an effective and robust multi-modal detection framework in this paper. Our key idea is to reposition the focus of the fusion process, from hard-association to soft-association, leading to the robustness against degenerated image quality and sensor misalignment.
Specifically, we design a sequential fusion method that uses two transformer decoder layers as the detection head. To our best knowledge, we are the first to use transformer for LiDAR-camera 3D detection. Our first decoder layer leverages a sparse set of object queries to produce initial bounding boxes from LiDAR features. Unlike input-independent object queries in 2D carion2020endtoend; Sun2020SparseRE, we make the object queries input-dependent and category-aware so that the queries are enriched with better position and category information. Next, the second transformer decoder layer adaptively fuses object queries with useful image features associated by spatial and contextual relationships. We leverage a locality inductive bias by spatially constraining the cross attention around the initial bounding boxes to help the network better visit the related positions. Our fusion module not only provides rich semantic information to object queries, but also is more robust to inferior image conditions since the association between LiDAR points and image pixels are established in a soft and adaptive way. Finally, to handle objects that are difficult to detect in point clouds, we introduce an image-guided query initialization module to involve image guidance on the query initialization stage. Overall, the corporation of these components significantly improves the effectiveness and robustness of our LiDAR-camera 3D detector. To summarize, our contributions are fourfold:
Our studies investigate the inherent difficulties of LiDAR-camera fusion and reveal a crucial aspect to robust fusion, namely, the soft-association mechanism.
We propose a novel transformer-based LiDAR-camera fusion model for 3D detection, which performs fine-grained fusion in an attentive manner and shows superior robustness against degenerated image quality and sensor misalignment.
We introduce several simple yet effective adjustments for object queries to boost the quality of initial bounding box predictions for image fusion. An image-guided query initialization module is also designed to handle objects that are hard to detect in point clouds.
We achieve the state-of-the-art 3D detection performance on nuScenes and competitive results on Waymo. We also extend our model to the 3D tracking task and achieve the 1st place in the leaderboard of the nuScenes tracking challenge.
Related Work
LiDAR-only 3D Detection aims to predict 3D bounding boxes of objects in given point clouds Yang2018PIXORR3; Qi2019DeepHV; Qi2020ImVoteNetB3; Vora2020PointPaintingSF; Zhou2019EndtoEndMF; Chen2020EveryVC; Shi2021FromPT; Zhu2019ClassbalancedGA; Zhu2020SSNSS; Chen2020ObjectAH. Due to the unordered, irregular nature of point clouds, many 3D detectors first project them onto a regular grid such as 3D voxels Zhou2018VoxelNetEL; Yan2018SECONDSE, pillars Lang2019PointPillarsFE or range images Sun2021RSNRS; Fan2021RangeDetID. After that, standard 2D or 3D convolutions are used to compute the features in the BEV plane, where objects are naturally separated, with their physical sizes preserved. Other works Shi2019PointRCNN3O; Yang20203DSSDP3; Yang2019STDS3; Shi2020PVRCNNPF directly operate on raw point clouds without quantization. The mainstream of 3D detection head is based on anchor boxes Lang2019PointPillarsFE; Zhou2018VoxelNetEL following the 2D counterparts, while Yin2020Centerbased3O; Wang2020PillarbasedOD adopt a center-based representation for 3D objects, largely simplifying the 3D detection pipeline. Despite the popularity of adopting the transformer architecture as a detection head in 2D carion2020endtoend, 3D detection models for outdoor scenarios mostly utilize the transformer for feature extraction Pan20203DOD; Mao2021VoxelTF; Sheng2021Improving3O. However, the attention operation in each transformer layer requires a computation complexity of for points, requiring a carefully designed memory reduction operation when handling LiDAR point clouds with millions of points per frame. In contrast, our model retains an efficient convolution backbone for feature extraction and leverages a transformer decoder with a small set of object queries as the detection head, making the computation cost manageable. The concurrent works Liu2021GroupFree3O; Liu2021SuppressandRefineFF; misra20213detr adopt transformer as a detection head but focus on indoor scenarios and extending these methods to outdoor scenes is non-trivial.
LiDAR-Camera 3D Detection has gained increasing attention due to the complementary roles of point clouds and images. Early works Qi2018FrustumPF; Shin2019RoarNetAR; Chen2017Multiview3O adopt result-level or proposal-level fusion, where the fusion granularity is too coarse to release the full potential of two modalities. Since PointPainting Vora2020PointPaintingSF was proposed, the point-level fusion methods Sindagi2019MVXNetMV; Huang2020EPNetEP; Wang2021PointAugmentingCA have shown great advantages and promising results. However, such methods are easily affected by the sensor misalignment due to the hard association between points and pixels established by calibration matrices. Moreover, the simple point-wise concatenation ignores the quality of real data and contextual relationships between two modalities, and thus leads to degraded performance when the image features are defective. In our work, we explore a more robust and effective fusion mechanism to mitigate these limitations during LiDAR-camera fusion.
Methodology
In this section, we present the proposed method TransFusion for LiDAR-camera 3D object detection. As shown in Fig. 2, given a LiDAR BEV feature map and an image feature map from convolutional backbones, our transformer-based detection head first decodes object queries into initial bounding box predictions using the LiDAR information, and then performs LiDAR-camera fusion by attentively fusing object queries with useful image features. Below we will first provide the preliminary knowledge about a transformer architecture for detection and then present the detail of TransFusion.
Transformer Vaswani2017AttentionIA has been widely used for 2D object detection Zhu2021DeformableDD; Sun2020SparseRE; Gao2021FastCO; Yao2021EfficientDI since DETR carion2020endtoend was proposed. DETR uses a CNN backbone to extract image features and a transformer architecture to convert a small set of learned embeddings (called object queries) into a set of predictions. The follow-up works Zhu2021DeformableDD; Sun2020SparseRE; Yao2021EfficientDI further equip the object queries with positional information Slightly different concepts might be introduced, e.g., reference points in Deformable-DETR Zhu2021DeformableDD and proposal boxes in Sparse-RCNN Sun2020SparseRE.. The final predictions of boxes are the relative offsets w.r.t. the query positions to reduce optimization difficulty. We refer readers to the original papers carion2020endtoend; Zhu2021DeformableDD for more details. In our work, each object query contains a query position providing the localization of the object and a query feature encoding instance information, such as the box’s size, orientation, etc.
2 Query Initialization
Input-dependent. The query positions in the seminal works carion2020endtoend; Zhu2021DeformableDD; Sun2020SparseRE are randomly generated or learned as network parameters, regardless of the input data. Such input-independent query positions will take extra stages (decoder layers) for their models carion2020endtoend; Zhu2021DeformableDD to learn the moving process towards the real object centers. Recently, it has been observed in 2D object detection Yao2021EfficientDI that with a better initialization of object queries, the gap between 1-layer structure and 6-layer structure could be bridged. Inspired by this observation, we propose an input-dependent initialization strategy based on a center heatmap to achieve competitive performance using only one decoder layer.
3 Transformer Decoder and FFN
The decoder layer follows the design of DETR misra20213detr and the detailed architecture is provided in the supplementary Sec. A. The cross attention between object queries and the feature maps (either from point clouds or images) aggregates relevant context onto the object candidates, while the self attention between object queries reasons pairwise relations between different object candidates. The query positions are embedded into -dimensional positional encoding with a Multilayer Perceptron (MLP), and element-wisely summed with the query features. This enables the network to reason about both context and position jointly.
The object queries containing rich instance information are then independently decoded into boxes and class labels by a feed-forward network (FFN). Following CenterPoint Yin2020Centerbased3O, our FFN predicts the center offset from the query position as , bounding box height as , size as , yaw angle as and the velocity (if available) as . We also predict a per-class probability for semantic classes. Each attribute is computed by a separate two-layer convolution. By decoding each object query into prediction in parallel, we get a set of predictions as output, where is the predicted bounding box for the -th query. Following misra20213detr, we adopt the auxiliary decoding mechanism, which adds FFN and supervision after each decoder layer. Hence, we can have initial bounding box predictions from the first decoder layer. We leverage such initial predictions in the LiDAR-camera fusion module to constrain the cross attention, as explained in the next section.
4 LiDAR-Camera Fusion
SMCA for Image Feature Fusion. Multi-head attention is a popular mechanism to perform information exchange and build a soft association between two sets of inputs, and it has been widely used for the feature matching task Sarlin2020SuperGlueLF; Sun2021LoFTRDL. To mitigate the sensitivity towards sensor calibration and inferior image features brought by the hard-association strategy, we leverage the cross-attention mechanism to build the soft association between LiDAR and images, enabling the network to adaptively determine where and what information should be taken from the images.
Specifically, we first identify the specific image in which the object queries are located using previous predictions as well as the calibration matrices, and then perform cross attention between the object queries and the corresponding image feature map. However, as the LiDAR features and image features are from completely different domains, the object queries might attend to visual regions unrelated to the bounding box to be predicted, leading to a long training time for the network to accurately identify the proper regions on images. Inspired by Gao2021FastCO, we design a spatially modulated cross attention (SMCA) module, which weighs the cross attention by a 2D circular Gaussian mask around the projected 2D center of each query. The 2D Gaussian weight mask is generated in a similar way as CenterNet Zhou2019ObjectsAP, where is the spatial indices of the weight mask , is the 2D center computed by projecting the query prediction onto the image plane, is the radius of the minimum circumscribed circle of the projected corners of the 3D bounding box, and is the hyper-parameter to modulate the bandwidth of the Gaussian distribution. Then this weight map is element-wisely multiplied with the cross-attention map among all the attention heads. In this way, each object query only attends to the related region around the projected 2D box, so that the network can learn where to select image features based on the input LiDAR features better and faster. The visualization of the attention map is shown in Fig. 3. The network typically tends to focus on the foreground pixels close to the object center and ignore the irrelevant pixels, providing valuable semantic information for object classification and bounding box regression. After SMCA, we use another FFN to produce the final bound box predictions using the object queries containing both LiDAR and image information.
5 Label Assignment and Losses
Following DETR misra20213detr, we find the bipartite matching between the predictions and ground truth objects through the Hungarian algorithm Kuhn1955TheHM, where the matching cost is defined by a weighted sum of classification, regression, and IoU cost:
where is the binary cross entropy loss, is the L1 loss between the predicted BEV centers and the ground-truth centers (both normalized in $L_{iou}\lambda_{1},\lambda_{2},\lambda_{3}$ are the coefficients of the individual cost terms. We provide sensitivity analysis of these terms in the supplementary Sec. C. Since the number of predictions is usually larger than that of GT boxes, the unmatched predictions are considered as negative samples. Given all matched pairs, we compute a focal loss Lin2017FocalLF for the classification branch. The bounding box regression is supervised by an L1 loss for only positive pairs. For the heatmap prediction, we adopt a penalty-reduced focal loss following CenterPoint Yin2020Centerbased3O. The total loss is the weighted sum of losses for each component. We adopt the same label assignment strategy and loss formulation for both decoder layers.
6 Image-Guided Query Initialization
Since our object queries are currently selected using only LiDAR features, it potentially leads to sub-optimality in terms of the detection recall. Empirically, our model already achieves high recall and shows superior performance over the baselines (Sec. 5). Nevertheless, to further leverage the ability of high-resolution images in detecting small objects and make our algorithm more robust against sparse LiDAR point clouds, we propose an image-guided query initialization strategy, which selects object queries leveraging both the LiDAR and camera information.
Specifically, we generate a LiDAR-camera BEV feature map by projecting the image features onto the BEV plane through cross attention with LiDAR BEV features . Inspired by Roddick2020PredictingSM, we use the multiview image features collapsed along the height axis as the key-value sequence of the attention mechanism, as shown in Fig. 4. The collapsing operation is based on the observation that the relation between BEV locations and image columns can be established easily using camera geometry, and usually there is at most one object along each image column. Therefore, collapsing along the height axis can significantly reduce the computation without losing critical information. Although some fine-grained image features might be lost during this process, it already meets our need as only a hint on potential object positions is required. Afterward, similar to Sec. 3.2, we use to predict the heatmap, which is averaged with the LiDAR-only heatmap as the final heatmap . Using to select and initialize the object queries, our model is able to detect objects that are difficult to detect in LiDAR point clouds.
Note that proposing a novel method to project the image features onto the BEV plane is beyond the scope of this paper. We believe that our method could benefit from more research progress Roddick2019OrthographicFT; Roddick2020PredictingSM; Philion2020LiftSS in this direction.
Implementation Details
Training. We implement our network in PyTorch paszke2017automatic using the open-sourced MMDetection3D mmdet3d2020. For nuScenes, we use the DLA34Yu2018DeepLA of the pretrained CenterNet as our 2D backbone and keep its weights frozen during training, following Wang2021PointAugmentingCA. We set the image size to , which performs comparably with full resolution (). VoxelNet Zhou2018VoxelNetEL; Yan2018SECONDSE is chosen as our 3D backbone. Our training consists of two stages: 1) We first train the 3D backbone with the first decoder layer and FFN for 20 epochs, which only needs the LiDAR point clouds as input and produces the initial 3D bounding box predictions. We adopt the same data augmentation and training schedules as prior LiDAR-only works Yin2020Centerbased3O; Zhu2019ClassbalancedGA. Note that we also find the copy-and-paste augmentation strategy Yan2018SECONDSE benefits the convergence but could disturb the real data distribution, so we disable this augmentation for the last 5 epochs following Wang2021PointAugmentingCA (they called a fade strategy). 2) We then train the LiDAR-camera fusion and the image-guided query initialization module for another 6 epochs. We find that this two-step training scheme performs better than joint training, since we can adopt more flexible augmentations for the first training stage. See supplementary Sec. B for the detailed hyper-parameters and settings on Waymo.
Testing. During inference, the final score is computed as the geometric average of the heatmap score and the classification score . We use all the outputs as our final predictions without Non-maximum Suppression (NMS) (see the effect of NMS in supplementary Sec. D). It is noteworthy that previous point-level fusion methods such as PointAugmenting Wang2021PointAugmentingCA rely on two different models for camera FOV and LiDAR-only regions if the cameras are not 360-degree cameras, because only points in the camera FOV could fetch the corresponding image features. In contrast, we use a single model to deal with both camera FOV and LiDAR-only regions, since object queries located outside camera FOV will directly ignore the fusion stage and the initial predictions from the first decoder layer will be a safeguard.
Experiments
In this section, we first make comparisons with the state-of-the-art methods on nuScenes and Waymo. Then we conduct extensive ablation studies to demonstrate the importance of each key component of TransFusion. Moreover, we design two experiments to show the robustness of our TransFusion against inferior image conditions. Besides TransFusion, we also include a model variant, which is based on the first training stage, i.e., producing the initial bounding box predictions using only point clouds. We denote it as TransFusion-L and believe that it can serve as a strong baseline for LiDAR-only detection. We provide the qualitative results in supplementary Sec. I.
nuScenes Dataset. The nuScenes dataset is a large-scale autonomous-driving dataset for 3D detection and tracking, consisting of 700, 150, and 150 scenes for training, validation, and testing, respectively. Each frame contains one point cloud and six calibrated images covering the 360-degree horizontal FOV. For 3D detection, the main metrics are mean Average Precision (mAP) Everingham2009ThePV and nuScenes detection score (NDS). The mAP is defined by the BEV center distance instead of the 3D IoU, and the final mAP is computed by averaging over distance thresholds of across ten classes. NDS is a consolidated metric of mAP and other attribute metrics, including translation, scale, orientation, velocity, and other box attributes. Following CenterPoint Yin2020Centerbased3O, we set the voxel size to .
Waymo Open Dataset. This dataset consists of 798 scenes for training and 202 scenes for validation. The official metrics are mAP and mAPH (mAP weighted by heading accuracy). The mAP and mAPH are defined based on the 3D IoU threshold of 0.7 for vehicles and 0.5 for pedestrians and cyclists. These metrics are further broken down into two difficulty levels: LEVEL1 for boxes with more than five LiDAR points and LEVEL2 for boxes with at least one LiDAR point. Unlike the 360-degree cameras in nuScenes, the cameras in Waymo only cover around 250 degrees horizontally. The voxel size is set to .
nuScenes Results. We submitted our detection results to the nuScenes evaluation server. Without any test time augmentation or model ensemble, our TransFusion outperforms all competing non-ensembled methods on the nuScenes leaderboard at the time of submission. As shown in Table 2, our TransFusion-L already outperforms the state-of-the-art LiDAR-only methods by a significant margin (+5.2% mAP, +2.9% NDS) and even surpasses some multi-modality methods. We ascribe this performance gain to the relation modeling power of the transformer decoder as well as the proposed query initialization strategies, which are ablated in Sec. 5.3. Once enabling the proposed fusion components, our TransFusion receives remarkable performance boost (+3.4% mAP, +1.5% NDS) and outperforms all the previous methods, including FusionPainting Xu2021FusionPaintingMF, which uses extra data to train their segmentation sub-networks. Moreover, thanks to our soft-association mechanism, TransFusion is robust to inferior image conditions including degenerated image quality and sensor misalignment, as shown in the next section.
Waymo Results. We report the performance of our model over all three classes on Waymo validation set in Table 2. Our fusion strategy improves the mAPH of pedestrian and cyclist classes by 0.3 and 1.5x, respectively. We suspect two reasons for the relatively small improvement brought by the image components. First, the semantic information of images might have less impact on the coarse-grained categorization of Waymo. Second, the initial bounding boxes from the first decoder layer are already with accurate locations since the point clouds in Waymo are denser than those in nuScenes (see more discussions in supplementary Sec. H). Note that CenterPoint achieves a better performance with a multi-frame input and a second-stage refinement module. Such components are orthogonal to our method and we leave a more powerful TransFusion for Waymo as the future work. PointAugmenting achieves better performance than ours but relies on CenterPoint to get the predictions outside camera FOV for a full-region detection, making their system less flexible.
Extend to Tracking. To further demonstrate the generalization capability, we evaluate our model in a 3D multi-object tracking (MOT) task by performing tracking-by-detection with the same tracking algorithms adopted by CenterPoint. We refer readers to the original paper Yin2020Centerbased3O for details. As shown in Table 3, our model significantly outperforms CenterPoint and sets the new state-of-the-art results on the leaderboard of nuScenes tracking.
2 Robustness against Inferior Image Conditions
We design three experiments to demonstrate the robustness of our proposed fusion module. Since the nuScenes test set only allows at most three submissions, all the experiments are conducted on the validation set. For fast iteration, we reduce the first stage training to 12 epochs and remove the fade strategy. All the other parameters are the same as the main experiments. To avoid overstatement, we additionally build two baseline LiDAR-camera detectors by equipping our TransFusion-L with two representative fusion methods on nuScenes: fusing LiDAR and image features by point-wise concatenation (denoted as CC) and the fusion strategy of PointAugmenting (denoted as PA).
Nighttime. We first split the validation set into daytime and nighttime based on scene descriptions provided by nuScenes and show the performance gain under different situations in Table 4. Our method brings a much larger performance gain during nighttime, where the worse lighting negatively affects the hard-association based fusion strategies CC and PA.
Degenerated Image Quality. In Table 5, we randomly drop several images for each frame by setting the image features of such images to zero during inference. Since both CC and PA fuse LiDAR and image features in a tightly-coupled way, their performance drops significantly when some images are not available during inference. In contrast, our TransFusion is able to maintain a high mAP under all cases. When all the six images are not available, CC and PA suffer from and mAP degradation, respectively, while TransFusion still keeps the mAP at a competitive level of . This advantage comes from the sequential design and the attentive fusion strategy, which first generates initial predictions based on LiDAR data and then only gathers useful information from image features adaptively. Moreover, we could even directly disable the fusion module if the camera malfunctioning is known, such that the whole system could still work seamlessly in a LiDAR-only mode.
Sensor Misalignment. We evaluate different fusion methods under a setting where LiDAR and images are not well-calibrated following RoarNet Shin2019RoarNetAR. Specifically, we randomly add a translation offset to the transformation matrix from camera to LiDAR sensor. As shown in Fig. 5, TransFusion achieves better robustness against the calibration error compared with other fusion methods. When two sensors are misaligned by , the mAP of our model only drops by 0.49%, while the mAP of PA and CC degrades by 2.33% and 2.85%, respectively. In our method, the calibration matrix is only used for projecting the object queries onto images, and the fusion module is not strict with the projected locations since the attention mechanism could adaptively find the relevant image features around based on the context information. The insensitivity towards sensor calibration also enables the possibility to pipelining the 2D and 3D backbones such that the LiDAR features are fused with the features from the previous images Vora2020PointPaintingSF.
3 Ablation Studies
We conduct ablation studies on the nuScenes validation set to study the effectiveness of the proposed components.
Query Initialization. In Table 6, we study how the query initialization strategy affects the performance of the initial bounding box prediction. a) the first row is TransFusion-L. b) when the category-embedding is removed, NDS drops to 63.9%. d)-f) shows the performance of the models trained without the input-dependent strategy. Specifically, we make the query positions as a set of learnable parameters () to capture the statistics of potential object locations in the dataset. The model under this setting only achieves 33.8% NDS. Increasing the number of decoder layers or the number of training epochs boosts the performance, but TransFusion-L still outperforms the model in (f) by 9.0% NDS. a), c): In contrast, with the proposed query initialization strategy, our TransFusion-L does not require more decoder layers.
Fusion Components. To study how the image information benefits the detection results, we ablate the proposed fusion components by removing the feature fusion module (denoted as w/o Fusion) and the image-guided query initialization (denoted as w/o Guide). As shown in Table 7, the image feature fusion and image-guided query initialization bring 4.8% and 1.6% mAP gain, respectively. The former provides more distinctive instance features, which are particularly critical for classification on nuScenes, where some categories are challenging to distinguish, such as trailer and construction vehicle. The latter affects less, since TransFusion-L already has enough recall. We believe the latter will be more useful when point clouds are sparser. Compared with other fusion methods, our fusion strategy brings a larger performance gain with a modestly increasing number of parameters and latency. To better understand where the improvements are from, we show the mAP breakdown on different subsets based on the range in Table 8. Our fusion method gives larger performance boost for distant regions where 3D objects are difficult to detect or classify in LiDAR modality.
Conclusion
We have designed an effective and robust transformer-based LiDAR-camera 3D detection framework with a soft-association mechanism to adaptively determine where and what information should be taken from images. Our TransFusion sets the new state-of-the-art results on the nuScenes detection and tracking leaderboards, and shows competitive results on Waymo detection benchmark. The extensive ablative experiments demonstrate the robustness of our method against inferior image conditions. We hope that our work will inspire further investigation of LiDAR-camera fusion for driving-scene perception, and the application of a soft-association based fusion strategy to other tasks, such as 3D segmentation.
Acknowledgements. This work is supported by Hong Kong RGC (GRF 16206819, 16203518, T22-603/15N), Guangzhou Okay Information Technology with the project GZETDZ18EG05, and City University of Hong Kong (No. 7005729).
References
Supplementary Material
The supplementary document is organized as follows:
Sec. A depicts the detailed network architectures of our transformer decoder layers.
Sec. B provides the implementation details of TransFusion and the training settings on nuScenes and Waymo.
Sec. C reports our sensitivity analysis of the matching cost during label assignment.
Sec. D presents the effect of NMS on TransFusion and CenterPoint.
Sec. E provides the results of using PointPillars as our 3D backbone.
Sec. F discusses the effect of 2D backbone in TransFusion.
Sec. G shows the results with different number of object queries.
Sec. H discusses the performance gain of image information on Waymo.
Sec. I provides visualization results on the nuScenes and Waymo datasets.
A Network Architectures
The detailed architectures of the respective transformer decoder layers for initial bounding box prediction and LiDAR-camera fusion are shown in Fig. 6. Following Liu2021GroupFree3O, we adopt the common practice of transformer except that we use the learned positional encoding instead of the fixed sine positional encoding Vaswani2017AttentionIA. For the image-guided query initialization module, we use the LiDAR BEV features as query sequence and collapsed image features as key-value sequence, and only perform cross attention to save the computation cost. Our model can benefit from the efficient attention mechanisms in recent works such as Zhu2021DeformableDD.
B Implementation Details
Our implementation is based on the open-sourced codebase MMDetection3D mmdet3d2020, which provides many popular 3D detection methods, including PointPillar, VoxelNet, and CenterPoint. For our 3D backbone, we set its hyper-parameters according to CenterPoint-Voxel’s official implementation. For the transformer-decoder-based detection head, the hidden dimension is set to 256 and dropout is set to . We use and queries for nuScenes and Waymo since the max numbers of objects in one frame are 142 and 185, respectively. Since our object queries are non-parametric, we are able to modify the number of queries during inference. We provide the ablations on the number of object queries in Sec. G. When selecting object queries from the heatmap, we pick local maximum pixels whose values are greater than or equal to their 8-connected neighbors. To avoid mistakenly suppress nearby instances for small objects, we do not check the local maximum for pedestrian and traffic cone on nuScenes and for pedestrian and cyclist on Waymo. Following PointAugmenting Wang2021PointAugmentingCA, we adopt DLA34 of CenterNet pre-trained on monocular 3D detection task as our 2D backbone for nuScenes. Since there is no public available 2D backbone pre-trained on Waymo dataset, we train a Faster-RCNN Ren2015FasterRT using the 2D labels provided by Waymo and use its ResNet50 and FPN as our 2D backbone. We freeze the weight of image backbone during training and set the image resolution to half of the full resolution for both nuScenes and Waymo to speed up the training process.
nuScenes. Following the common practice, we transform previous ten LiDAR sweeps into the current frame to produce a denser point cloud input for both training and inference. The detection range is set to for X and Y axes, and for Z axis. The maximum numbers of non-empty voxels for training and inference are set to 120,000 and 160,000, respectively. In terms of the data augmentation strategy, we adopt random flipping along both X and Y axes, global scaling with a random factor from , global rotation between , as well as the copy-and-paste augmentation Yan2018SECONDSE. We follow CBGS Zhu2019ClassbalancedGA to perform class-balanced sampling and optimize the network using the AdamW optimizer with one-cycle learning rate policy, with max learning rate , weight decay , and momentum to . We train the 3D backbone with the first decoder layer and FFN for 20 epochs, and the LiDAR-camera fusion components for 6 epochs with batch size of 16 using 8 Tesla V100 GPUs. We use gradient clipping at an norm of 0.1 to stabilize the training process. The weighting coefficients of heatmap loss, classification loss, and regression loss are , and , respectively. The coefficients of matching cost are , respectively. The sensitivity analysis of the matching cost coefficients is provided in the next section.
Waymo. For Waymo, we only use a single sweep as input and set the detection range to for X and Y axes, and for Z axis. The maximum number of non-empty voxels is set to 150,000. We adopt the same training strategies as nuScenes except: 1) The first stage training last for 36 epochs with batch size of 16 under a max learning rate of . 2) The weighting coefficient of regression loss is changed to , following CenterPoint. 3) The matching cost coefficients are set to , respectively.
C Label Assignment Strategy
Following DETR, we perform label assignment by finding the bipartite matching between predictions and ground -truth objects through the Hungarian algorithm. In Table 4, we study the effect of the coefficient of each matching cost term on the detection performance of TransFusion-L. Since the matching results are only decided by the relative values of individual matching cost terms, we keep and try different values for and . As shown in Table 4, we find the network suffers from a convergence issue when the coefficient of the classification cost is too large, and the detection performance is not sensitive to the coefficient within a reasonable range. Since the weighting coefficients of the matching cost need some tuning, we additionally propose a heuristic label assignment strategy (denoted as Heuristic) to avoid hyper-parameter tuning. The Heuristic assignment strategy follows the simple rules: each GT box will only be assigned to the predicted box with the same category and the smallest center distance. If a conflict appears, the predicted box will be matched to the closer GT box. In this way, we also find the one-to-one matching between prediction and GT but with some GT boxes unused. We find Heuristic works quite well for uncrowded scenes but for objects in a crowded scene, such as pedestrian or traffic cone in nuScenes, it is unable to prevent duplicate predictions, which will be further explained in Sec. D.
D Effect of NMS
Recently, many works carion2020endtoend; Sun2020SparseRE; Zhu2021DeformableDD in 2D detection have focused on removing the last non-differentiable component, Non-Maximum Suppression (NMS), in the detection pipeline. OneNet Sun2021WhatMF systematically compares the end-to-end detectors with non-end-to-end detectors, and claims that the one-to-one positive sample assignment as well as classification cost in the matching cost are the two key factors in producing a large score gap between duplicate prediction and building an end-to-end detection system without NMS. We refer readers to the original paper Sun2021WhatMF for more details. Following DETR’s label assignment strategy, our method naturally satisfies these two requirements and do not need NMS. As show in Table 10, our method still maintains a high mAP without NMS while CenterPoint drops about 12% mAP. This advantage eliminates the hand-designed NMS post-processing step and makes TransFusion more practical and handy for deployment to new scenarios in the real applications. Besides, since the Heuristic strategy mentioned in Sec. C does not have classification cost involved in the assignment stage, this strategy is unable to produce a large score gap between duplicate prediction. This is why it does not perform as well as the baseline model on Pedestrian and Traffic cone.
E Pillar-based 3D Backbone
To demonstrate our framework’s compatibility with other 3D backbones, we use PointPillars as our 3D backbone to produce the BEV features, while keeping all the other settings the same as the main experiments. The voxel size is set to . As shown in Table 11, our model outperforms CenterPoint by a remarkable margin under the same pillar-based backbone, showing its great generalization ability.
F Discussions of the 2D Network
Current multi-modality detection models usually employ CNN features from 2D networks pre-trained on different tasks (i.e., segmentation or detection) and with different resolution (i.e., different levels from ResNet or DLA). There is no existing work analyzing what kind of image features are most useful for a 3D detection model, and using improper image features might prevent the release of the potential for a multi-modality detection system. We believe that the sequential design of our method enables a flexible and off-the-shelf experiment base to explore the effects of different image features. Therefore, we explore this question by fixing the 3D backbone with the first decoder layer and performing the second stage training with different image features.
From Table 4, we find image features of the 2D instance segmentation model bring the largest performance boost compared with that of detection models. In terms of different levels of the feature pyramid, the feature map of level 0 (stride 4) brings a slightly larger performance gain. We suspect the image features at that level contain more fine-grained information which is important to distinguish small or distant objects. Image features from level 1 (stride 8) and level 2 (stride 16) can bring a similar gain with a smaller resolution of feature maps, while image features from level 3 (stride 32) yields a drop of 1.2% mAP in comparison with the level-0 counterpart due to the row resolution.
G Adapt Queries at Test Time
Unlike DETR, our object queries are non-parametric and input-dependent. These two characteristics allow us to use different numbers of queries during inference. It could be useful when we have some prior knowledge about a scene, such as its crowdedness. In Table 13, we provide the performance evaluated under different object queries for the same model trained under queries. Note that we use to get all the numbers in the main text for its better performance-efficiency trade-off and use for online submission for a slightly better performance.
H Dicussions on Waymo
Our TransFusion brings smaller performance gain over TransFusion-L on Waymo compared with that on nuScenes. We speculate that this is mainly due to the following two reasons:
As shown in Table 2, compared with TransFusion-L, TransFusion brings the largest performance increase for bicycle (+8.7%), motorcycle (+5.4%), and construction vehicle (+4.9%) in terms of mAP on nuScenes. Due to the geometrical ambiguity, objects from the above three categories are difficult to distinguish using LiDAR information only, and thus the semantic information of images is extraordinarily important for more accurate classification. However, the categorization of Waymo is rather coarse-gained (i.e., vehicle, pedestrian, cyclist), which hides the improvement brought by the image information to some extent.
The LiDAR point clouds in Waymo are much denser than those in nuScenes (see Sec. I for visualizations). Thus the bounding box predictions of TransFusion-L are already with accurate localization, which reduces the room for further improvement by image fusion.
I Qualitative Results
We first compare the detection results of TransFusion and TransFusion-L on the nuScenes dataset in Fig. 7. The image information improves the performance of the LiDAR-only baseline through reducing the False Positive and False Negative. More visualization results on Waymo and nuScenes datasets are shown in Fig. 8 and Fig. 9, respectively.