BEVStereo: Enhancing Depth Estimation in Multi-view 3D Object Detection with Dynamic Temporal Stereo

Yinhao Li, Han Bao, Zheng Ge, Jinrong Yang, Jianjian Sun, Zeming Li

Introduction

Due to the stability and inexpensive cost of vision sensors, camera-based 3D object detection has received extensive concern. Specially, the multi-view schemes (Wang et al. 2022b; Huang et al. 2021; Liu et al. 2022a; Li et al. 2022b; Huang and Huang 2022; Liu et al. 2022b; Li et al. 2022a) show significantly promising, and have made lots of breakthroughs. However, there is still a substantial performance gap compared with LiDAR-based approaches (Lang et al. 2019; Yan, Mao, and Li 2018; Yin, Zhou, and Krahenbuhl 2021), since it exposes a notoriously ill-posed issue for perceiving depth.

Contemporary multi-view detectors (Huang et al. 2021; Huang and Huang 2022; Li et al. 2022a) predict a discrete depth distribution for each point of the field of view (FOV), which enables to project features from image representation to BEV map. The unified BEV map is the key to learning harmonious results since the overlap regions of adjacent views represent more complete to directly forecast results. Such sweetness is hard to be enjoyed by the monocular-based detector (Wang et al. 2021b), as a post-processing strategy is needed to remove repetitive and low-quality 3D boxes in overlap areas.

The above paradigm is based on an important preconceived assumption, i.e., the perceived depth distribution in FOV needs accurate enough. However, most of them perceive depth by only feeding into single-frame images, which is actually an ill-posed solution (Huang et al. 2021; Huang and Huang 2022; Li et al. 2022a). Several studies (Yao et al. 2018; Xue et al. 2019; Bae, Budvytis, and Cipolla 2022) point out that predicting depth needs multi-view stereo condition, which requires images from different views to construct cost volume. Fortunately, the automatic driving scenario is often processed in a continuous time sequence, enabling us to leverage temporal views for constructing multi-view stereo.

To carry out the traditional temporal stereo technology like (Yao et al. 2018) is non-trivial in automatic driving scenarios, which manifests in two aspects:

Large memory cost. When we replace the depth module in BEVDepth with a basic temporal stereo method (Yao et al. 2018), the memory cost grows to 3.5 times that of BEVDepth despite bringing a 1.6 percent promotion on NDS, making it a tremendous burden to apply it to a detection task;

Failing to reason the depth of moving objects and static ego vehicle cases. Temporal stereo approaches are unable to handle several situations (Wang, Pang, and Lin 2022) like a static ego vehicle and moving objects since the parallax angle tends to 0 if ego vehicle is static and the stereo is unable to match if the object is moving. However, after statistics in nuScenes scene, over 10% of the frames’ ego vehicles are static, while approximately 25% of the objects are moving. Therefore, these two shortcomings limit its application to autonomous driving scenarios.

MVS methods (Wang, Pang, and Lin 2022; Wang et al. 2022a) expose that the majority of the computational memory cost is associated with constructing cost volume due to its calculation procedure of dense similarity. It naturally motivates us to construct a sparse cost volume for cutting computational memory. To this end, we propose a dynamic mechanism to sample a small number of reference candidate features for building cost volume instead of all ones along the depth axis. It is implemented by predicting two modeling parameters, i.e., depth center μ\mu and depth range σ\sigma. This far, it can significantly reduce computational memory. Going into one step, we introduce a parameter evolution method for μ\mu and σ\sigma, which is carried out by applying the EM algorithm to update the modeling parameters μ\mu and σ\sigma. With the evolution technique, it is possible to continuously improve reference candidate features that are more important for cost volume while adjusting to situations including moving objects and stationary ego vehicles. This insight is similar to MaGNet (Bae, Budvytis, and Cipolla 2022), but it not only fails to deal with complex outdoor situations but introduces redundantly learnable parameters to update μ\mu and σ\sigma. Finally, we also introduce an advanced variant of Circle NMS (Yin, Zhou, and Krahenbuhl 2021), which takes objects’ size into account for better removing duplicate 3D boxes.

We instantiate our proposed methods to advanced BEVDepth (Li et al. 2022a), namely BEVStereo. By conducting comprehensive experiments on nuScence benchmark (Caesar et al. 2020), it shows significant improvements in the 3D object detection task. In conclusion, the contributions of this work are as three-fold as follows:

We point out that the MVS technology is a promising method for tackling the ill-posed issue of depth perception in camera-based 3D object detection task. But it exposes two fatal flaws in the automatic driving scenarios, i.e., either large memory cost issue or moving objects and static ego vehicles.

We introduce a dynamic temporal stereo technique, which can save extreme memory cost to construct cost volume. Moreover, a parameter evolution algorithm is proposed to tackle moving and noisy features of objects.

BEVStereo improves mAP and NDS by 1.7% and 1.7% on nuScenes dataset, while achieving the new SOTA performance on the camera-only track. Extensive experiments verify that our approach can effectively be adapted to moving objects and static ego vehicles.

Related Work

Many approaches have made their effort on predicting objects directly from single images. For the purpose of 3D object detection, Cai et al. (Cai et al. 2020) calculates the depth of the objects by integrating the height of the objects in the image with the height of the objects in the real world. Based on FCOS (Tian et al. 2019), FCOS3D (Wang et al. 2021b) extends it to 3D object detection by changing the classification branch and regression branch which predicts 2D and 3D attributes at the same time. M3D-RPN (Brazil and Liu 2019) treats mono-view 3D object detection task as a stand-alone 3D region proposal network, narrowing the gap between LiDAR-based approaches and camera-based methods. D4LCN (Ding et al. 2020) replaces 2D depth map with pseudo LiDAR representation to better present 3D structure. DFM (Wang, Pang, and Lin 2022) integrates temporal stereo to mono-view 3D object recognition, improving the quality of depth estimation while minimizing the negative effects of difficult situations that temporal stereo is unable to handle.

Multi-view 3D Object Detection

Current multi-view 3D object detectors can be divided into two schemas: LSS-based (Philion and Fidler 2020) schema and transformer-based schema.

BEVDet (Huang et al. 2021) is the first study that combines LSS and LiDAR detection head which uses LSS to extract BEV feature and uses LiDAR detection head to propose 3D bounding boxes. By introducing previous frames, BEVDet4D (Huang and Huang 2022) acquires the ability of velocity prediction. To reduce memory usage, M2BEV (Xie et al. 2022) decreases the learnable parameters and achieves high efficiency on both inference speed and memory usage. BEVDepth (Li et al. 2022a) uses LiDAR to generate depth GT for supervision and encodes camera intrinsic and extrinsic parameters to enhance the model’s ability of depth perception.

DETR3D (Wang et al. 2022b) extends DETR (Carion et al. 2020) into 3D space, using transformer to generate 3D bounding boxes. Based on DETR, PETR (Liu et al. 2022a) and PETRV2 (Liu et al. 2022b) adds position embedding onto it. BEVFormer (Li et al. 2022b) uses deformable transformer to extract features from images and uses cross attention to link the feature between frames for velocity prediction.

Depth Estimation

Based on the number of images used for depth estimation, depth estimation methods can be divided into single-view depth estimation and multi-view depth estimation.

Although predicting depth from a single image is obviously ill-posed, it is still possible to estimate some of the depth of the objects by using the context as a signal. Therefore, many approaches (Bhat, Alhashim, and Wonka 2021; Eigen and Fergus 2015; Eigen, Puhrsch, and Fergus 2014a; Fu et al. 2018) use CNN method to predict depth.

For the task of multi-view depth estimation, Constructing cost volume is an effective way to predict depth (Zhu et al. 2021; Wei et al. 2021, 2022). MVSNet (Yao et al. 2018) is the first research that uses cost volume for depth estimation. RMVSNet (Yao et al. 2019) reduces memory cost by introducing GRU module. MVSCRF (Xue et al. 2019) adds CRF module onto MVSNet. PointMVSNet (Chen et al. 2019) uses point algorithm to optimize the regression of depth estimation. Cascade MVSNet (Gu et al. 2020) uses cascade structure, making it able to use large depth range and a small amount of depth intervals. Fast-MVSNet (Yu and Gao 2020) uses sparse cost volume and Gauss-Newton layer to speed up MVSNet. Wang et al. (Wang et al. 2021a) use adaptive patchmatch and multi-scale fusion to achieve good performance while mataining high efficiency. Bae et al. (Bae, Budvytis, and Cipolla 2022) introduce MaGNet to better fuse single-view depth estimation and multi-view depth estimation.

Method

BEVStereo is a stereo-based multi-view 3D object detector. By applying our temporal stereo technique, it is able to handle complex outdoor scenarios while maintaining memory efficiency. We also propose a size-aware circle NMS approach to improve the proposal suppression process.

LSS-based (Philion and Fidler 2020) multi-view 3D object detectors currently include four components: an image encoder to extract the image features, a depth module to generate depth and context, then outer product them to get point features, a view transformer to convert the feature from camera view to the BEV view, and a 3D detection head to propose the final 3D bounding boxes.

Temporal stereo methods to predict depth

MVS-based (Yao et al. 2018) methods predict depth by constructing cost volume. For every pixel on the reference feature, they initially put forth a number of candidates along the depth axis. They next convert these candidates from reference to source using a homography warping operation in order to retrieve the relevant source feature and create the cost volume. After cost volume is constructed. For the purpose of predicting the confidence of each depth candidate, 3D convolution is performed to regularize the cost volume.

Dynamic Temporal Stereo

Based on BEVDepth (Li et al. 2022a), BEVStereo changes the way of generating depth prediction. Instead of predicting depth from a single image, BEVStereo predicts both depth from single feature (mono depth) and depth from temporal stereo (stereo depth). For mono depth, we directly predict depth prediction, which is the same as BEVDepth. For stereo depth, we firstly predict depth center (μ\mu) and depth range (σ\sigma), then μ\mu and σ\sigma are used to generate depth distribution. Additionally, Weight Net is used to create a weight map that will be applied on stereo depth. Mono depth and weighted stereo depth are combined to get the final depth. Our framework overview is illustrated in Fig. 1.

Our Depth Module simultaneously predicts mono depth, μ\mu, σ\sigma and context. After iterating μ\mu and σ\sigma by our EM method, they are used to generate the stereo depth. The process of iterating μ\mu and σ\sigma is illustrated in Fig. 2.

We choose to estimate μ\mu and σ\sigma, which stand for the depth center and depth range of the cost volume. Compared to other stereo-based methods of splitting bins along the depth dimension (Yao et al. 2018; Wang et al. 2022a), our method can dynamically choose the search area while also lowering the number of candidates. After estimating μ\mu and σ\sigma of the reference frame, we can dynamically select candidates for each pixel based on the depth center and range of cost volume and obtain the depth of these candidates. These candidates are used for homography warping operation to fetch the feature from source frame, as illustrated in Equ. 1, where PP denotes the coordinate of the point, DD denotes the depth of the candidate, srcsrc denotes source frame, refref denotes reference frame, Mref2srcM_{ref2src} denotes the transformation matrix from the reference frame to source frame and KK denotes the intrinsic matrix. The reference feature and the warpped source feature are used to construct cost volume. Similarity Net is followed to predict the confidence score of all candidates.

Inspired by the EM algorithm, We attempt to make the expectation of μ\mu closer to the depth gt during the iteration process. Since we compute each point’s confidence after sampling a number of points close to μ\mu, it is only natural that we use this knowledge to further our objectives. As a result, we update μ\mu using the weight sum method, which causes μ\mu to become the expectation of the sample points for each iteration. The update rule is illustrated in Eq. 2, where DiD_{i} denotes the depth of the iith candidate and PiP_{i} denotes the probability of the iith candidate. When facing cases like static ego vehicle and moving objects, all candidates share the same low probability since it is hard to find the best match point on the source feature, μ\mu is able to maintain its value by using the weight sum technique. For other scenarios, the value of μ\mu will approach the true depth value in the process of iteration. Surprisingly, we discover that when μ\mu and mono depth are trained together, the quality of initial μ\mu is also enhanced under the direction of mono depth. Therefore, in all kinds of scenarios, our dynamic temporal stereo approach can improve depth prediction. As μ\mu is being updated in the process of iteration, it is also critical to find the suitable σ\sigma to set the searching range. In accordance with existing information, the searching range should be reduced when the confidence of μ\mu is high and expanded when it is low, we update σ\sigma following Equ. 3 where PμP_{\mu} denotes the confidence of μ\mu. Without introducing any learnable parameters, the search range is optimized during iteration.

To prevent the scenario where the projected μ\mu is far from the depth gt, making it difficult to optimize μ\mu during iteration. we divide the depth into different ranges and use our iteration technique in each split range. After the iteration process is finished, the depth map is generated following Equ. 4 where P denotes the computed depth confidence and D denotes the depth of the split bins along the depth axis for each pixel.

Weight Net

Even while the temporal stereo is capable of accurately predicting depth, there are still some areas where it is unreliable because some reference feature points do not correlate to positions on source feature. Therefore, we introduce Weight Net to better combine mono depth and stereo depth. To do this, we apply the same homography warping operation to fetch the mono depth of the source frame, using μ\mu as the depth. A similarity net is then applied to the warped mono depth from the source frame and the mono depth from the reference frame to construct the weight map.

Size-aware Circle NMS

The distance between the centers of two bounding boxes is used by circle NMS (Yin, Zhou, and Krahenbuhl 2021) function as a criterion for suppression. Circle NMS achieves excellent efficiency and good performance by bypassing the difficult process of computing rotated IoU of bouding boxes. However, ignoring the size of boxes will result in two drawbacks as illustrated in Fig. 3: 1) No matter how closely the boxes overlap, the NMS algorithm yields the same output as long as the box centers are fixed. 2) When boxes are placed differently, boxes with 0 IoU may be removed while boxes with high IoU are kept.

We propose size-aware circle NMS, which avoids computing rotated IoU while taking into consideration the size of the boxes. We separate the distance of two bounding boxes’ centers into x axis and y axis. We use xthrex_{thre} and ythrey_{thre} as threholds of x axis and y axis, which are computed following Equ. 5 and Equ. 6, where θ\theta denotes the orientation, ww denotes the hyper parameter of scale factor, dxd_{x} denotes the length of the box and dyd_{y} denotes the width of the box. The box will be suppressed when the distance in x axis is smaller than xthrex_{thre} and distance in y axis is smaller than ythrey_{thre}. By applying size-aware circle NMS, the blue box with a lower score will be suppressed in scenarios like the left portion of Fig 3 because it has a greater xthrex_{thre} and ythrey_{thre}. The blue box will be suppressed in scenarios like the right portion of Fig. 3 because the distances in the x and y axes are more likely to be smaller than xthrex_{thre} and ythrey_{thre} in the mean time.

Experiment

In this section, we first describe the experimental settings that we employ before going into the specifics of our implementation strategy. Experiments involving heavy ablation are carried out to confirm the efficacy and validity of BEVStereo.

We decide to run our experiments on the nuScenes (Caesar et al. 2020) dataset. For training, we use LiDAR and image data, but we only use image data for inference. In the case of image data, the key frame image and the furthest sweep connected to it are used, whereas in the case of LiDAR data, only the key frame data is used. We assess the results of our method using detection and depth metrics. Memory usage is also used to assess the effectiveness of our method. To be more specific, we report the mean Average Precision (mAP), nuScenes Detection Score (NDS), mean Average Translation Error (mATE), mean Average Scale Error (mASE), mean Average Orientation Error (mAOE), mean Average Velocity Error (mAVE), and mean Average Attribute Error (mAAE). We follow the established evaluation procedures for the depth estimation task (Eigen, Puhrsch, and Fergus 2014b), reporting scale invariant logarithmic error (SILog), mean absolute relative error (Abs Rel), mean squared relative error (Sq Rel), mean log10 error (log10), and root mean squared error (RMSE) to assess our approach.

Implementation details

We implement BEVStereo based on BEVDepth (Li et al. 2022a). The feature map we employ for building the cost volume has a downsampling rate of 4 while the depth feature’s final form remains unchanged. The MVS (Yao et al. 2018) approach is applied to replace the depth module in BEVDepth with the same input resolution and output resolution in order to fairly demonstrate the effectiveness of our method. The learning rate is set to 2e-4, the EMA technique is also used, and AdamW (Loshchilov and Hutter 2017) is used as the optimizer. During training, we use both image and BEV data augmentation.

Analysis

We perform numerous experiments to examine the mechanism of BEVStereo in order to better understand how it works. We choose BEVDepth (Li et al. 2022a) as baseline, we also implement MVSNet (Yao et al. 2018) on BEVDepth as a comparison to show the distinct benefit that BEVStereo provides, detection results and recall results are used for comparison.

We keep track of memory usage and detection results to demonstrate how effectively we use our memory. We also monitor the same matrics for the MVS-based (Yao et al. 2018) approach for fair comparison.

As illustrated in Tab. 6, BEVStereo increases the metrics on mAP, mATE, and NDS considerably at the expense of adding little memory consumption. When compared to using MVS (Yao et al. 2018) on BEVDepth (Li et al. 2022a), BEVStereo considerably reduces memory usage while boosting performance.

Performance analysis

To begin with, we demonstrate the performance comparison under the nuScenes (Caesar et al. 2020) evaluation metrics. As shown in Tab. 1, Our BEVStereo outperforms BEVDepth on mAP, mATE and NDS. Tab. 2 shows that the accuracy of depth estimation is improved by introducing our design.

We assess the performance of BEVStereo under challenging conditions such as moving objects, and static ego vehicles in order to show how well it adapts to complicated outdoor environments. Tab. 3 demonstrates that BEVStereo still has the ability to improve performance even while MVS approach fails when dealing with moving objects. The static objects, which make up the majority of MVS schema’s contribution, are also used to evaluate our method. As shown in Tab. 4, BEVStereo’s ability of perceiving static objects is even higher than BEVDepth with MVS. We choose frames whose ego vehicle has a low velocity for evaluation since MVS cannot handle situations when this occurs. As can be seen in Tab. 5, BEVStereo still improves performance even when MVS fails in these conditions. It is important to note that BEVStereo still produces the similar results when faced with circumstances like moving objects and static ego vehicles if μ\mu is not updated during the inference step. This demonstrates that our schema is capable of guiding the Depth Module to produce better μ\mu and maintaining the initial prediction of μ\mu in the face of these eventualities.

Ablation Study

We conduct various experiments during the inference stage by modifying the number of iterations just to verify the function of iterating μ\mu and σ\sigma. As illustrated in Tab. 7, the detection results improve as the number of iterations grows.

Weight Net

We run the experiment under identical conditions without Weight Net to assess its validity. Weight Net promotes the detection results, as shown in Tab. 1.

Size-aware Circle NMS

We compare BEVStereo with the size-aware circle NMS to BEVStereo with the conventional circle NMS as our baseline. They are subjected to class-aware and class-agnostic procedures in order to test the validity of size-aware circle NMS.

As shown in Tab. 8, our size-aware circle NMS improves on the matrices of mAP, mATE, and NDS when using class-aware NMS. The traditional distance-based circle NMS has completely lost its capacity to suppress under class-agnostic circumstance, while our size-aware circle NMS continues to function well.

Efficient Voxel Pooling v2

In the previous version of Efficient Voxel Pooling (Li et al. 2022a), threads within the same warp access memory discontinuously, leading to more memory transactions, which results in poor performance. We enhance Efficient Voxel Pooling by improving the way threads are mapped, as illustrated in Fig. 4. For each block, we employ 32 and 4 threads on the x and y axes. First, 128 point coordinates are loaded into shared memory by all the threads in one block. Then, one point feature at a time is processed by each warp. According to the point coordinates, the point feature is atomically accumulated to the matching BEV feature. The 128 point features are processed round robin by four warps in a block till they are finished. In this manner, performance-limiting memory transactions from the L2 cache and global memory are diminished.

We compare the latency of Efficient Voxel Pooling v1 and Efficient Voxel Pooling v2 using various resolutions. Efficient Voxel Pooling v2 is able to reduce the latency up to 40%.

Visualization

As illustrated in Fig. 5, we can find that BEVStereo has the ability to promote the accuracy of depth estimation on both moving and static objects. We also visualize the detection results, as shown in Fig. 6 which also demonstrates the performance promotion brought by BEVStereo.

Benchmark Result

We compare BEVStereo with other state-of-the-art methods like CenterPoint (Yin, Zhou, and Krahenbuhl 2021), FCOS3D (Wang et al. 2021b), DETR3D (Wang et al. 2022b), BEVDet (Huang et al. 2021), PETR (Liu et al. 2022a), BEVDet4D (Huang and Huang 2022) and BEVFormer (Li et al. 2022b). We evaluate our BEVStereo on the nuScenes test and val set. As shown in Tab. 9 and Tab. 10, BEVStereo achieves the highest score of camera-based methods on both mAP and NDS.

Conclusion

In this paper, a novel multi-view 3D object detector is proposed, namely BEVStereo. BEVStereo improves performance without significantly increasing memory usage by applying dynamic temporal stereo technique to create temporal stereo. Some complex scenarios that other stereo-based approaches cannot handle can be resolved by our method. In addition, we propose size-aware circle NMS, which takes the size of boxes into account while avoiding the laborious computation of rotated IoU. Under both class-aware and class-agnostic circumstances, our size-aware circle NMS performs satisfactorily. Last but not least, we present Efficient Voxel Pooling v2, which speeds up voxel pooling by improving the efficiency of memory accesses.

Acknowledgements

Throughout the process of developing BEVStereo, I have received a great deal of guidance and assistance. I would like to thank Haotian Zhang, Yuefeng Wu and Tai Wang for their wonderful collaboration and patient support.

References