M$^2$BEV: Multi-Camera Joint 3D Detection and Segmentation with Unified Birds-Eye View Representation

Enze Xie, Zhiding Yu, Daquan Zhou, Jonah Philion, Anima Anandkumar, Sanja Fidler, Ping Luo, Jose M. Alvarez

Introduction

The ability of perceiving objects and the environment in a unified framework is a core requirement for robotic systems including autonomous vehicles (AV). Because for these applications, the performance of downstream tasks such as localization, mapping and planning highly rely on the quality of the perception of different tasks. The perception system of an autonomous agent needs to have several separate components: (1) 3D perception - understanding the dynamics and scene layout in 3D is an informative world representation for localization and planning. (2) Holistic understanding - the ability to jointly perceive both the objects and the environment needed by diverse downstream tasks. (3) Multi-sensor - the need to comprehensively sense the surroundings from multiple sensors with different views and/or different modalities for added redundancy and reliability. In this paper, we are interested in designing the perception system for autonomous vehicles (AV).

For designing perception systems with the above considerations, many scene understanding methods and benchmarks have been proposed. There are two most important tasks in AV perception: 3D object detection and BEV segmentation. 3D object detection is one of the popular tasks, with the canonical input to the detector being LiDAR point clouds . In cases where LiDAR is not available, multi-camera 3D object detection presents an alternative . The goal of multi-camera 3D object detection is to predict 3D bounding boxes in a BEV (ego vehicle) coordinate system given only monocular camera inputs. Another important task is BEV segmentation, the goal of BEV segmentation is to perform semantic segmentation of the environment, e.g., drivable area and lane boundaries, in the BEV frame. Unlike detection, segmentation allows dense prediction of “stuff” classes belonging to the static environment, a necessary step for map construction. In this work, our central motivation is to provide a single unified framework for joint 3D object detection and BEV segmentation, under a multi-view camera-only perception setting.

It is worth mentioning that existing camera-based methods are not suitable for 360∘ multi-task AV perception without significant changes, and we are the first address at this problem with a unified framework. We illustrate this in more details with three mainstream camera-based methods: (1) Monocular 3D object detection methods, e.g. CenterNet and FCOS3D , predict 3D bounding boxes within each view separately. They require additional post-processing steps to fuse the predictions across different views and remove redundant bounding boxes. These steps are typically not robust nor differentiable, and therefore are not amenable to end-to-end joint reasoning with downstream planning tasks. (2) Pseudo LiDAR based methods, e.g. pseudo-LiDAR . These methods can reconstruct the 3D voxel with predicted depth, but are sensitive to errors in depth estimation and often require additional depth annotations and training supervisions. (3) Transformer based method. Recently, DETR3D uses a transformer framework that 3D object queries are projected to multi-view 2D images and interact with the image features in a top-down manner. Although DETR3D supports multi-view 3D detection, it does not support BEV segmentation and multi-tasking for only considering object queries without dense BEV representations.

Our approach. We argue that a unified BEV feature representation matters for 360∘ multi-task AV perception. In this paper, we obtain the BEV representation based on an efficient feature transformation, where BEV representations are obtained by reconstructing multi-view 2D image features to 3D voxels along the rays. As shown in Fig. 2, a unified BEV representation allows us to easily support multiple tasks such as detection and segmentation with minimal additional computational cost. This makes our work different from naively stacking separate task-specific networks in a sequential manner. Our contributions can be summarized as follows:

We propose a unified framework to transform multi-camera images to a Bird’s-Eye View (BEV) representation for multi-task AV perception, including 3D object detection and BEV segmentation. To the best of our knowledge, this is the first work target at predicting these two challenging tasks in a single framework.

We propose several novel designs such as efficient BEV encoder, dynamic box assignment, and BEV centerness. These designs help GPU memory-efficiency and significantly improve the performance on both tasks.

We show that large-scale pre-training with 2D annotation (e.g. nuImage) and 2D auxiliary supervision can significantly improve the performance of 3D tasks and benefits label efficiency. As a result, our method achieves state-of-the-art performance on both 3D object detection and BEV segmentation on nuScenes, showing that BEV representation is promising for next-generation AV perception.

Specifically, To make the framework usable in real-world scenarios with limited computational budget, we propose several empirical designs to significantly improve the accuracy and the GPU memory efficiency. The first one is an efficient BEV encoder, which uses a “Spatial to Channel” (S2C) operator to transform a 4D voxel tensor to a 3D BEV tensor, thus avoiding the usage of memory expensive 3D convolutions. The second one is the dynamic box assignment dedicated for 3D detection tasks. It uses a learning-to-match strategy to assign ground-truth 3D boxes with anchors. The third one is BEV centerness specially designed for BEV segmentation. It is motivated by the fact that longer distance area in BEV havs less pixels in image. We thus re-weight pixels according to the distance to the ego-car and assign larger weights to farther samples. The last one is the 2D detection pre-training and auxiliary supervision on the 2D image encoder. It can significantly speed up the training convergence and improve the performance of 3D tasks.

Related Work

Monocular 3D Detection. Monocular 3D object detection is similar to 2D object detection in image space but more challenging than 2D detection since estimating 3D information from 2D images is ill-posed. Early approaches predict 2D boxes first and add another sub-network to regress to 3D boxes , or score 3D cuboids placed on the ground against the 2D proposals . Another line of work predicts pseudo dense depth and transforms an RGB image into other representations such as OFT and Pseudo-Lidar . Several recent works directly predict the depth of objects based on 2D detectors. SS3D predicts 2D boxes, keypoints, and distance, and proposes a 3D IoU to facilitate training. FCOS3D adopts an advanced 2D anchor free detector FCOS , and adds 3D distance and 3D box prediction in the detection branch. PGD models the geometric relations across different objects to facilitate the depth estimation task.

Multi-view 3D Detection. Multiple datasets have been released for AV research recently that provide data from a full multi-camera sensor rig. Monocular 3D detectors can of course be extended to the multi-camera setting by independently processing each view and fusing the results from all the views in the post-processing stage. But hard-coding a post-processing is suboptimal and adds burdensome hyperparameters. Since post-processing is typically not differentiable, this approach cannot afford end-to-end training with downstream tasks such as planning. ImVoxelNet is a multi-view 3D detector that projects 2D image features to 3D voxels. Detection is then based on the obtained voxel representation. The projection of ImVoxelNet is the same as Atlas , a framework that focuses on 3D reconstruction. DETR3D extends DETR from 2D to 3D by generating object queries in 3D and back-projecting to the 2D image space to aggregate 2D image features. Similar to DETR, DETR3D also uses set-to-set loss for end-to-end detection without NMS. However, it is not obvious that how to extend DETR3D to bird’s-eye view segmentation since it does not have dense BEV feature representation.

BEV Segmentation. VPN proposes a simple view transformation module to transform a feature map from the perspective view to the bird’s-eye view with two-layer MLPs and performs indoor layout segmentation. PON proposes a transformer architecture to convert an image representation to a BEV representation for autonomous driving segmentation. LSS lifts 2D features into a 3D BEV representation by estimating an implicit depth distribution and performs BEV segmentation and planning. NEAT proposes to use neural attention fields to predict a map in bird’s-eye view scene coordinates. Concurrent work Panoptic-BEV uses a two-branch transformer to project the front view image into BEV and performs panoptic segmentation in the BEV plane. However, it only considers a single view and does not estimate the height of objects.

Method

In this section, we present the detailed design of M2BEV, our newly proposed multi-view pipeline for joint 3D object detection and segmentation with a unified BEV representation. In Sec. 3.1, we will start with detailed descriptions of 5 core components. In Sec. 3.2, we illustrate in details on how the novel unified BEV representation make it possible for the two originally disjoint tasks. In Sec. 3.3, we introduce some important designs based on our baseline method, which significantly improve the results of both tasks to very competitive performance, as shown in Fig. 4 a,b. In Sec. 3.4, we give detailed illustrations on the loss function of M2BEV.

Overview. Our framework takes NN RGB images from multi-view cameras as input and corresponding extrinsic/intrinsic parameters. The outputs are 3D bounding boxes of objects and segmentation of maps. The multi-view images are first fed to the image encoder and output 2D features. These 2D multi-view features are then projected to 3D space to construct the 3D voxel. The 3D voxel is then fed to an efficient BEV encoder to obtain the BEV feature. Finally, the task-specific head, e.g. 3D detection or BEV segmentation head, is added on the BEV feature. The details of each part in the pipeline are described as below. We also introduce more detailed implementation in appendix.

Part1: 2D Image Encoder. Given NN images ∈RH×W×3\in R^{H\times W\times 3}, we run the forward-pass with a shared CNN backbone for all images, e.g. ResNet, and a feature pyramid network (FPN) to create 4-level features F1,F2,F3,F4F_{1},F_{2},F_{3},F_{4} with shape H2i+1×W2i+1×C\frac{H}{2^{i+1}}\times\frac{W}{2^{i+1}}\times C for each image. We then upsample these features to H4×W4\frac{H}{4}\times\frac{W}{4}, then concatenate them and add one 1×\times1 conv to fuse them to form a tensor FF. The result is a set of features given as input (e.g. 6 views for nuScenes). These multi-view features are projected from 2D image coordinate to 3D ego-car coordinate in the next step.

Part2: 2D→\rightarrow3D Projection. The 2D→\rightarrow3D projection is the key module that make the multi-task training possible in our work. The multi-view features F∈RN×H4×H4×CF\in R^{N\times\frac{H}{4}\times\frac{H}{4}\times C} are combined and projected to 3D space to obtain a voxel V∈RX×Y×Z×CV\in R^{X\times Y\times Z\times C}, as shown in Fig. 4c. In Sec. 3.2, we explain precisely how we implement such 2D→\rightarrow3D projection. The voxel feature contains image features with all the views thus it is a unified feature representation. In the next step, the voxel feature is fed to 3D BEV encoder to obtain the BEV feature (reduce ZZ dimension).

Part3: 3D BEV Encoder. Given input 4D tensor voxel V∈RX×Y×Z×CV\in R^{X\times Y\times Z\times C}, we need to use BEV encoder to reduce ZZ dimension and output BEV feature B∈R12X×12Y×CB\in R^{\frac{1}{2}X\times\frac{1}{2}Y\times C}. The intuitive idea is to use several 3D convolutions with stride=2 in ZZ dimension but this way is very slow and inefficiency. Here we propose an efficient BEV encoder by transform 4D tensor to 3D with a novel “Spatial to Channel” (S2C) operator. Details see Sec. 3.3.

Part4: 3D Detection Head. With a unified BEV feature BB, it is possible for us to leverage the popular head designed for lidar-based 3D detection. Specifically, we directly adopt the detection head from PointPillars , which generates dense 3D anchors in BEV and then predicts the category, box size, and direction of each object. The PointPillars’s detection head is super simple and efficient, which only contains three parallel 1×\times1 convolutions. Different from PointPillars, we propose a dynamic box assignment to assign anchors with ground-truth, because we find it is more friendly for camera-based setting. We provide details in Sec. 3.3.

Part5: BEV Segmentation Head. Benefit from the powerful BEV representation, our segmentation head is also very simple. We directly add four 3×33\times 3 convolutions on the BEV feature and use one 1×11\times 1 convolution to get the final prediction∈RH×W×N\in R^{H\times W\times N}, where NN is the number of categories. Here N=2N=2 because we follow LSS to generate map ground-truth with two categories related to the environment: drivable area and lane boundary. We also propose a BEV centerness strategy to re-weight the loss for each pixel with a different physical distance, as described in Sec. 3.3.

2 Efficient 2D→→\rightarrow3D Projection

where DD is the depth of pixel Pi,jP_{i,j}. If DD is unknown, each pixel in PP is mapped to a set of points in the camera ray in 3D space.

Our Approach. Here, we assume the depth distribution along the ray is uniform, which means that all voxels along a camera ray are filled with the same features corresponding to a single pixel in PP in 2D space. This uniform assumption increases the computational and memory efficiency by reducing the number of learned parameters.

Comparison with LSS . The most related work to M2BEV is LSS , which implicitly predicts a non-uniform depth distribution and lifts 2D features from H×WH\times W to H×W×DH\times W\times D, typically with D≥50D\geq 50, where DD is the size of the categorical depth distribution. This step is very memory-expensive, prohibiting LSS from using a larger network or high-resolution images as input; We evaluate LSS and find the GPU memory of LSS is 3×\times higher than ours. Moreover, the image size of popular AV datasets, e.g. nuScenes, is 1600×9001600\times 900, and advanced monocular 3D detectors typically use ResNet-101 as a backbone. However, LSS only uses a small backbone (EfficientNet-B0) and the input image size is only 128×384128\times 384. In contrast, we do not implicitly estimate the depth when lifting 2D features to 3D as in LSS. As a result, our projection is more efficient and does not need learned parameters, which allows us to use a larger backbone (ResNet-101) and higher resolution input (1600×9001600\times 900).

3 Improvement Designs

Efficient BEV Encoder. Given a 4D tensor voxel V∈RX×Y×Z×CV\in R^{X\times Y\times Z\times C} input, we first propose a “Spatial to Channel (S2C)” operation to transform VV from 4D tensor X×Y×Z×CX\times Y\times Z\times C to 3D tensor X×Y×(ZC)X\times Y\times(ZC) via “torch.reshape\tt torch.reshape” operation. Then we use several 2D convolutions to reduce the channel dimension.

Remark: Comparison to 3D convolutions with stride 2 on ZZ dimension. We observe that 3D convolution is more “expensive” by being slower and more memory consuming than S2C+2D convolutions. It is impossible to build a heavy BEV encoder with 3D convolutions. However, S2C allows us to easily use and stack more 2D convolutions.

Dynamic Box Assignment. Many LiDAR-based works such as PointPillars assign 3D anchors for ground-truth boxes using a fixed intersection-over-union (IoU) threshold. However, we argue that this hand-crafted assignment is suboptimal for our problem because our BEV feature does not consider the depth in LiDAR, thus the BEV representation may encode less-accurate geometric information. Insipred by FreeAnchor that use learning-to-match assignment in 2D detection, we extend this assignment to 3D detection. The main difference is that the original FreeAnchor assigns 2D anchors and ground truth (GT) boxes in the image coordinate frame, while we assigns 3D anchors in the BEV coordinate.

During training, we first predict the class ajclsa_{j}^{cls} and the location ajloca_{j}^{loc} for each anchor aj∈Aa_{j}\in A and select a bag of anchors for each ground-truth box based on IoU. We use a weighted sum of classification score and localization accuracy to distinguish the positive anchors; the intuition behind this practice is that, an ideal positive anchor should have high confidence in both classification and localization. The rest of the anchors with low classification scores or large localization errors are set as negative samples. Kindly refer to for more details.

BEV Centerness. The concept of “centerness” is commonly used in 2D detectors to re-weight positive samples. Here we extend the concept of “centerness” in a non-trivial distance-aware manner, from 2D image coordinate to 3D BEV coordinate. This process is illustrated in Fig. 5a in more details. The motivation is that area in BEV space farther away from the ego car correspond to fewer pixels in the images. So an intuitive idea is make the network pay more attention on the farther area. Specifically, BEV centerness is defined as below:

where (xi,yi)(x_{i},y_{i}) is one point in the BEV frame, ranging from -50m to +50m, and (xc,yc)(x_{c},y_{c}) is the center point corresponding to the location of the ego vehicle. We use sqrt\rm sqrt here to slow down the increase of the centerness. The BEV centerness ranges from 1 to 2 and is used as a loss weight in Eq. 5. Thus, errors in predictions for samples far away from the center are punished more. We show that BEV centerness improves BEV segmentation in different ranges in Fig. 5a. The distance is farther, the IoU improvement is higher.

2D Detection Pre-training. We empirically found that pre-training the model on large-scale 2D detection dataset, e.g. nuImage dataset, could significantly improve the 3D accuracy as detailed in Fig. 5b. The nuImage dataset contains 93000 images categorized into 10 different classes with instance-level annotation. More specifically, we pre-train a Cascade Mask R-CNN on nuImage for 20 epoches. The box mAP after the pre-training are 52.5 and 56.4 with ResNet-50 and ResNeXt-101 as backbone respectively. When training on nuScenes, the pretrained weights from the 2D dertector’s backbone are then used to initialize the 2D encoder in M2BEV pipeline. The rest of the layers are randomly initialized.

2D Auxiliary Supervision. As detailed in Fig. 3, after obtaining the image features, we add a 2D detection head on the features at different scales and calculating the losses with 2D GT bboxes generated from the 3D boxes in ego-car coordinate. The 2D detection head is implemented in the same way as proposed in FCOS . It is worth noting that the auxiliary head is only used during the training phases and will be removed during the inference phase. As a result, it does not introduce additional computation cost in inference. Fig. 5c illustrates how 2D GT boxes are generated from 3D annotations. The 3D GT boxes from the ego-car coordinates are back-projected to the 2D image space with the camera intrinsic parameters. In this way, 2D box GTs can be obtained without additional efforts.

Remark: By using 2D detection as pre-training and auxiliary supervision, the image features are more aware of objects thus boosting up 3D accuracy.

4 Training Losses

Our final loss is a combination of the 3D detection loss Ldet\mathcal{L}_{det}, BEV segmentation loss Lseg\mathcal{L}_{seg}, and 2D auxiliary detection loss Ldet2d\mathcal{L}_{det_{2d}}:

where Ldet3d\mathcal{L}_{det_{3d}} is same as the loss function introduced in PointPillars :

where NposN_{pos} is the number of positive samples, and we set βcls=1.0\beta_{cls}=1.0, βloc=0.8\beta_{loc}=0.8 and βdir=0.8\beta_{dir}=0.8. We empirically find that for camera-based methods, larger βdir\beta_{dir} is better. The classification loss is Focal Loss, the direction loss is binary cross-entropy loss. 3D Boxes (with 2D velocity) are defined by (x,y,z,w,h,l,θ,vx,vy)(x,y,z,w,h,l,\theta,v_{x},v_{y}) and we use Smooth-L1 loss for each item with loss weight [1,1,1,1,1,1,1,0.2,0.2][1,1,1,1,1,1,1,0.2,0.2].

For Lseg3d\mathcal{L}_{seg_{3d}}, we use a combination of Dice loss Ldice\mathcal{L}_{dice} and binary cross entropy loss Lbce\mathcal{L}_{bce}, as shown in below:

where βdice=1\beta_{dice}=1 and βbce=1\beta_{bce}=1.

For Ldet2d\mathcal{L}_{det_{2d}}, the loss is same as FCOS , as shown below:

Experiments

Dataset. We evaluate M2BEV on the nuScenes dataset . nuScenes contains 1000 video sequences collected in Boston and Singapore with 700/150/150 scenes for training/validation/testing. Each sample consists of a LiDAR scan and images from 6 cameras: front_left\tt front\_left, front\tt front, front_right\tt front\_right, back_left\tt back\_left, back\tt back, back_right\tt back\_right. nuScenes includes 10 categories for 3D bounding boxes.

Evaluation metrics. For the detection task, we use the standard evaluation metrics of Average Precision (mAP), and nuScenes detection score (NDS) . mAP defines a match by considering the 2D center distance on the ground plane rather than IoU-based affinities. NDS is a weighted sum of several metrics related to the intuitive notion on what detections are important for safe driving. For bird’s-eye view segmentation, we follow LSS and use IoU scores as the metric.

Network architecture. We use ResNet block structure with deformable convolution as our image backbone. ResNet-50 is used in the ablation study, and ResNeXt-101 is used for the final result when comparing to the state-of-the-art methods. The architecture for the detection head is the same as PointPillars: three parallel 1×\times1 convolutions to predict classes, boxes, and directions. The segmentation head has four 3×\times3 convolutions followed by one 1×\times1 convolution. The voxel size is 400×400×12400\times 400\times 12 where each (Δx,Δy,Δz)(\Delta x,\Delta y,\Delta z) bin in the voxel is (0.25m, 0.25m, 0.5m).

Training and inference. AdamW is used with learning rate 1e−31e^{-3} and weight decay 1e−21e^{-2}. We train for 12 epochs in all experiments and we use “polylr” to gradually decrease the learning rate. The batch size is 1 sample per GPU and each sample has 6 images. We do not use any augmentation, e.g. flipping, during training or testing, and keep the input resolution fixed at 1600×9001600\times 900. The model is trained on 3 DGX nodes. Each node has 8 Tesla-V100 GPUs. The total training time is about 6 hours with the ResNet-50 as backbone and 23 hours with ResNeXt-101 backbone.

2 Comparison with state of the art

3D object detection. We evaluate our model on the official nuScenes detection benchmark. The result is shown in Tab. 1. On the validation set, our method outperforms PGD , previous best method, by a large margin with 4.8% mAP and 4.2% NDS, demonstrating the importance of using a unified BEV representation. On the test set, our method outperforms baselines with only camera data and achieves more than 1% mAP than DD3D and DETR3D . Note that DD3D and DETR3D use external 3D depth data to pre-train the network, which is fundamentally different from other methods. We also compare M2BEV with detection only and joint training, and we find that joint training slightly hurts the detection results.

BEV segmentation. In Tab. 4.1, we compare M2BEV with other BEV segmentation methods . M2BEV achieves significantly higher IoU than LSS in IoU of drivable area (+3.0%) and lane boundary (+18.1%), which indicates that depth estimation is not necessary in BEV segmentation. We also compare M2BEV with segmentation only and joint training, and we get similar conclusion in above that multi-task learning slight hurt the performance of both tasks. In Fig. 7, we also visualize both detection and segmentation results in images and BEV space.

3 Ablation Studies

3D detection. As shown in Tab. 4(a), the naive baseline with the original fixed IoU anchor matching in PointPillars only obtains 19.7% mAP and 27.8% NDS. There are several observations: 1) When replaced with dynamic anchor matching strategy, the mAP and NDS are improved by 7.8% and 4.8%. This result indicates that rule-based anchor assignment is not optimal in our task. 2) 2D nuImage detection pre-training largely improves mAP by 5.8% and NDS by 6.0%. 3) the Spatial-to-Channel (S2C) operation in BEV encoder helps improve the detection result by more than 1%. 4) 2D auxiliary supervision slightly improves both the mAP and NDS without inference cost. In the next paragraph, we give a deeper discussion on how 2D detection pre-training improves our 3D task.

2D detection pre-training. We pre-train a Cascande Mask R-CNN on COCO and nuImage datasets. COCO is a generic 2D detection dataset and nuImage is an autonomous driving 2D detection dataset. We verify that there is no overlap between nuImage training set and nuScenes val/test set.

In Tab. 4(a), we can observe that 2D detection pre-training is beneficial over ImageNet pre-training. First, with coco pre-training, mAP and NDS directly improve about 3%. However, coco dataset still has a considerable domain gap with nuScenes, while the domain shift between nuImage and nuScenes is small. When using nuImage pre-training, the mAP and NDS can be further improved by 2.7% and 3.5%. We demonstrate that nuImage pre-train also largely improves other 3D detectors, such as FCOS3D, in Tab. 4(f).

From the left figure in Fig. 6, we see that with nuImage pre-training, training converges faster and performance largely increases. From the right figure, we train a model with (1) only 50% nuScenes data (with nuImage pre-training) and (2) 100% nuScenes data (with ImageNet pre-training). We find that the former has similar mAP and NDS compared with the latter but uses only 50% 3D annotations.

Remark: The experiment provides a new insight leveraging large-scale 2D box annotation to boost 3D detection performance. 2D box annotation is much cheaper than 3D and much easier to obtain. This observation suggests that 2D labels can be leveraged to reduce the need for 3D annotation.

BEV segmentation. As shown in Tab. 4(b), our naive baseline achieves 61.3% and 22.4% IoU for drivable area and lanes, respectively. When adding nuImage pre-training, the segmentation IoU significantly improves by 6.2% and 6.0%. Then the Spatial-to-Channel operator further largely boosts up the baseline by 10.6% and 12.2%. Finally, BEV centerness also helps to improve the results by increasing the performance on far-away objects. As shown in Tab. 4(e), the “S2C” operation with 2D convolutions in efficient BEV encoder can stack more layers to refine BEV feature, while being more efficiency than naive stacking fewer 3D convolution layers.

Multi-task joint training. In the above paragraphs, we ablate on different tasks individually. In Tab. 4(c), we demonstrate performance when we jointly train both two tasks. Interestingly, we observe that 3D object detection and BEV segmentation do not help to improve each other, and joint training slightly hurts the performance of each task. We observe that the location distribution of objects and maps do not have strong correlation, e.g. many cars are not in the drivable area. Other works also point out that not all tasks benefit from joint training. We leave this challenge for future work.

Remark: Although multi-tasking 3D detection and BEV segmentation causes slight drop in performance, the advantages of multi-task inference in autonomous driving still outweigh this issue. A shared network can support many tasks with little extra computational cost. In addition, the small performance drop is marginal compared to the competitive performance of the proposed framework.

Backbones. Tab. 4(d) shows results of M2 BEV with different backbones. It can be seen that better features extracted by deeper and advanced design networks improve the performance as expected.

Runtime efficiency. In Tab. 4.1, we compare the runtime efficiency of M2BEV with FCOS3D , DETR3D and LSS .

Single task. M2BEV is more efficient than monocular detector FCOS3D and multi-view detector DETR3D, because FCOS3D needs to merge results individually from different cameras in the post-processing step, while DETR3D uses a Transformer decoder, which is complex and inefficient due to a cross-attention module.

Multi-task. In M2BEV, both tasks share most of the features, the two heads have few parameters, and the inference speed for a single task or multiple tasks is nearly the same. However, simply combining “FCOS3D+LSS” is very low-efficiency, and M2BEV is 4×\times faster than “FCOS3D+LSS”.

Robustness to calibration error. We also evaluate how the performance of our method changes with test time camera extrinsic noise. Specifically, we increase the extrinsic noise level from 1e−31e^{-3} to 2e−12e^{-1} during the test time. As shown in Fig. 8, when the extrinsic noise is small, e.g. <1e−2<1e^{-2}, M2BEV is relatively robust to calibration errors and the performance drop is only about 1%. However, further increasing the extrinsic noise level to a relative large number, e.g. >5e−2>5e^{-2}, would considerably degrade the performance of both 3D detection and BEV segmentation. We consider this an important problem for autonomous driving in future research.

4 Limitations

The proposed M2BEV framework is not perfect, when road conditions are complicated, there are failure cases in both 3D detection and BEV segmentation as partly shown in Fig. 7. Although our method is competitive in camera-based methods, there is still significant room to improve compared with LiDAR-based methods. Test-time camera extrinsic noise is also an inevitable issue in real-world scenarios. As studied in Sec. 4.3 and shown in Fig. 8, when severe calibration errors are present, some degradation in the prediction quality can be observed with the proposed framework.

Conclusion

3D object detection and map segmentation are the two most important tasks for multi-camera AV perception. This work proposes a framework for performing both tasks in one network. The key idea is to project multi-view features from the image plane to the BEV space to create a unified BEV representation. Detection and segmentation branches then operate on the BEV representation. We additionally demonstrate pre-training on cheap 2D data can improve the label efficiency for 3D tasks.

We believe our framework is not limited to only 3D detection and BEV segmentation. In the future, we are interested in tasks involving temporal information, such as 3D object tracking, motion prediction, and trajectory forecasting.

Appendix 0.A Additional Implementation Details

Codebase. Our work uses mmDetection3Dhttps://github.com/open-mmlab/mmdetection3d as the codebase. On the ResNet-50 backbone, we add the deformable convolution (DCN) in stage 3 and 4 features following . On the ResNeXt-101 backbone, DCN is added from stage 2 to stage 4 features. “SyncBN” is used in both backbone and the pyramid features.

Training details. Training is conducted on 3 DGX nodes with 24 GPUs, where we use warm-up to avoid the training collapse. The warmup iteration is set to 1000, and the starting learning rate is 1e-6. In the warmup stage, the learning rate gradually increases to 1e-3. We also use mix-precision training to decrease the GPU cost and speed up training.

Detection head. We use 3D rotate non-maximum suppression (NMS) to remove redundant boxes. The thresholds of NMS and the score of each box are set to 0.2 and 0.05. We set the maximum number of objects in one frame to be 500. For anchor generation, an anchor set with 4 sizes and 2 rotations are generated on each point of the feature map. The anchor sizes are: [0.86,2.59,1]\mathtt{[}0.86,2.59,1], [0.57,1.73,1]\mathtt{[}0.57,1.73,1], [1,1,1]\mathtt{[}1,1,1], [0.4,0.4,1]\mathtt{[}0.4,0.4,1] whereas the rotations are: [0∘,90∘][0^{\circ},90^{\circ}].

Data pre-processing. The images from the 6 views are loaded with image normalization. The “mean” and “std” are set to [123.675,116.28,103.53]\mathtt{[}123.675,116.28,103.53] and [58.395,57.12,57.375]\mathtt{[}58.395,57.12,57.375] respectively following the common setting. Note that we do not include other data augmentations such as color jitter, and random rescaling.

Appendix 0.B More Visualizations

We additionally visualize 6 groups of ground-truth and predicted results in Fig. 9, Fig. 10 and Fig. 11.

Fig. 9 is a night driving scene. The result indicates that M2BEV learns to see in the dark. For example, a very tiny car in a far distance can be successfully detected by MBEV{}^{B}EV, while it is missed to be annotated as ground-truth by human annotators.

For multi-camera 3D detection, an obvious challenge is that extra post-processing and care are needed for objects appearing between different cameras. In Fig. 10, M2BEV correctly localizes buses that appear between two different cameras, which shows the advantage from having a unified BEV representation.

Fig. 11 shows under a very crowded scene, M2BEV can still detect most of the objects and segment maps although part of these objects and maps are heavily occluded.

References