BEVDet4D: Exploit Temporal Cues in Multi-camera 3D Object Detection
Junjie Huang, Guan Huang
Introduction
Recently, autonomous driving draws great attention in both the research and the industry community. The vision-based perception tasks in this scene include 3D object detection, BEV semantic segmentation, motion prediction, and so on. Most of them can be partly solved in the spatial-only 3D working space with a single frame of data. However, with respect to the time-relevant targets like velocity, current vision-based paradigms with merely a single frame of data perform far poorer than those with sensors like LiDAR or radar. For example, the velocity error of the recently leading method BEVDet in the vision-based 3D object detection is 3 times that of the LiDAR-based method CenterPoint and 2 times that of the radar-based method CenterFusion . To close this gap, we propose a novel paradigm dubbed BEVDet4D in this paper and pioneer the exploitation of vision-based autonomous driving in the spatial-temporal 4D space.
As illustrated in Fig. 2, BEVDet4D makes the first attempt at accessing the rich information in the temporal domain. It simply extends the naive BEVDet by retaining the intermediate BEV features in the previous frames. Then it fuses the retained feature with the corresponding one in the current frame just by a spatial alignment operation and a concatenation operation. Other than that, we kept most other details of the framework unchanged. In this way, we place just a negligible extra computational budget on the inference process while enabling the paradigm to access the temporal cues by querying and comparing the two candidate features. Though simple in constructing the framework of BEVDet4D, it is nontrivial to build its robust performance. The spatial alignment operation and the learning targets should be carefully designed to cooperate with the elegant framework so that the velocity prediction task can be simplified and superior generalization performance can be achieved with BEVDet4D.
We conduct comprehensive experiments on the challenge benchmark nuScenes to verify the feasibility of BEVDet4D and study its characteristics. Fig. 1 illustrates the trade-off between inference speed and performance of different paradigms. Without bells and whistles, the BEVDet4D-Tiny configuration reduces the velocity error by 62.9% from 0.909 mAVE to 0.337 mAVE. Besides, the proposed paradigm also has significant improvement in the other indicators like detection score (+2.6% mAP), orientation error (-12.0% mAOE), and attribute error (-25.1% mAAE). As a result, BEVDet4D-Tiny exceeds the baseline by +8.4% on the composite indicator NDS. The high-performance configuration dubbed BEVDet4D-Base scores high as 42.1% mAP and 54.5% NDS, which has surpassed all published results in vision-based 3D object detection . Last but not least, BEVDet4D achieves the aforementioned superiority just at a negligible cost in inference latency, which is meaningful in the scenario of autonomous driving.
Related Works
Vision-based 3D object detection is a promising perception task in autonomous driving. In the last few years, fueled by the KITTI benchmark monocular 3D object detection has witness a rapid development . However, the limited data and the single view disable it in developing more complicated tasks. Recently, some large-scale benchmarks have been proposed with sufficient data and surrounding views, offering new perspectives toward the paradigm development in the field of 3D object detection. Based on these benchmarks, some multi-camera 3D object detection paradigms have been developed with competitive performance. For example, inspired by the success of FCOS in 2D detection, FCOS3D treats the 3D object detection problem as a 2D object detection problem and conducts perception just in image view. Benefitting from the strong spatial correlation of the targets’ attribute with the image appearance, it works well in predicting this but is relatively poor in perceiving the targets’ translation, velocity, and orientation. PGD further develops the FCOS3D paradigm by searching and resolving the outstanding shortcoming (i.e. the prediction of the targets’ depth). This offers a remarkable accuracy improvement on the baseline but at the cost of more computational budget and additional inference latency. Following DETR , DETR3D proposes to detect 3D objects in an attention pattern, which has similar accuracy as FCOS3D. Although DETR3D requires just half the computational budget, the complex calculation pipeline slows down its inference speed to the same level as FCOS3D. PETR further develops the performance of this paradigm by introducing the 3D coordinate generation and position encoding. Besides, they also exploit the strong data augmentation strategies just as BEVDet . Another concurrent work dubbed Graph-DETR3D also extends the DETR3D from two expects. Analogous to the second stage in CenterPoint , Graph-DETR3D samples multiple points in the 3D space instead of a single point when generating the features of the object queries. Another modification is making the multi-scale training become feasible for DETR3D paradigm by dynamically adjusting the depth target according to the scaling factor. As a novel paradigm, BEVDet makes the first attempt at applying a strong data augmentation strategy in vision-based 3D object detection. As BEVDet explicitly encodes features in the BEV space, it is scalable in multiple aspects including multi-tasks learning, multi-sensors fusion, and temporal fusion. BEVDet4D is the temporal extension of BEVDet .
So far, few works have exploited the temporal cues in vision-based 3D object detection. Thus, the existing paradigms perform relatively poorly in predicting the time-relevant targets like velocity than the LiDAR-based or radar-based methods. To the best of our knowledge, is the only one pioneer in this perspective. However, they predict the results based on a single frame and exploit the 3D Kalman filter to update the results for the temporal consistency of results between image sequences. The temporal cues are exploited in the post-processing phase instead of the end-to-end learning framework. Differently, we make the first attempt in exploiting the temporal cues in the end-to-end learning framework BEVDet4D, which is elegant, powerful, and still scalable. BEVFormer is a concurrent work of BEVDet4D. Analogous to those in the VID literature, they mainly focus on the feature fusion in the spatial-temporal 4D working space with the attention mechanism . The comparable velocity precision of BEVFormer is achieved by fusing features from multiple adjacent frames (i.e. 4 frames in total), which is analogous to most LiDAR-based methods with points from multiple sweeps. This is fundamentally different from the proposed BEVDet4D, which uses merely two adjacent frames and achieved a higher velocity precision in a more elegant pattern.
2 Object Detection in Video
Video object detection mainly fueled by the ImageNet VID dataset is analogous to the well-known tasks of common object detection which performs and evaluates the object detection task in the image-view space. The difference is that detecting objects in video can access the temporal cues for improving detection accuracy. The methods in this area access the temporal cues mainly according to two kinds of mediums: the predicting results or the intermediate features. The former is analogous to in vision-based 3D object detection, who optimizes the prediction results in a tracking pattern. The latter reutilizes the features from the previous frame based on some special architectures like LSTM for feature distillation , attention mechanism for feature querying , and optical flow for feature alignment . Specific for the scene of autonomous driving, BEVDet4D is analogous to the flow-based methods in mechanism but accesses the spatial correlation according to the ego-motion and conducts feature aggregation in the 3D space. Besides, BEVDet4D mainly focuses on the prediction of the velocity targets which is not in the scope of the common video object detection literature.
Methodology
As illustrated in Fig. 2, the overall framework of BEVDet4D is built upon the BEVDet baseline which is consists of four kinds of modules: an image-view encoder, a view transformer, a BEV encoder, and a task-specific head. All implementation details of these modules are kept unchanged. To exploit the temporal cues, BEVDet4D extends the baseline by retaining the BEV features generated by the view transformer in the previous frame. Then the retained feature is merged with the one in the current frame. Before that, an alignment operation is conducted to simplify the learning targets which will be detailed in the following subsection. We apply a simple concatenation operation to merge the features for verifying the BEVDet4D paradigm. More complicated fusing strategies have not been exploited in this paper.
Besides, the feature generated by the view transformer is sparse, which is too coarse for the subsequential modules to exploit the temporal cues. Therefore, an extra BEV encoder is applied to adjust the candidate features before the temporal fusion. In practice, the extra BEV encoder consists of two naive residual units , whose channel number is set the same as the input feature.
2 Simplify the Velocity Learning Task
Following nuScense , we denote the global coordinate system as , the ego coordinate system as , and the targets coordinate system as . As illustrated in Fig. 3, we construct a virtual scene with a moving ego vehicle and two target vehicles. One of the targets is static (i.e., painted green) in the global coordinate system, while the other one is moving (i.e., painted blue). The objects in two adjacent frames (i.e., frame and frame ) are distinguished with different transparentness. The position of the objects is formulated as . denotes the coordinate system where the position is defined in. denotes the time when the position is recorded. We use to denote the transformation from the source coordinate system into the target coordinate system.
Instead of directly predicting the velocity of the targets, we tend to predict the translation of the targets in the two adjacent frames. In this way, the learning task can be simplified as the time factor is removed and the positional shifting can be measured just according to the difference between the two BEV features. Besides, we tend to learn the position shifting that is irrelevant to the ego-motion. In this way, the learning task can also be simplified as the ego-motion will make the distribution of the targets’ positional shifting more complicated.
For example, due to the ego-motion, a static object (i.e., the green box in Fig. 3) in the global coordinate system will be changed into a moving object in the ego coordinate system. More specifically, the receptive field of the BEV features is symmetrically defined around the ego. Considering the two features generated by the view transformer in the two adjacent frames, their receptive fields in the global coordinate system are diverse due to the ego-motion. Given a static object, its position in the global coordinate system is denoted as and in the two adjacent frames. The positional shifting in the two features should be formulated as:
According to Eq. 1, if we directly concatenate the two features, the learning target (i.e., the positional shifting of the target in the two features) of the following modules is relevant to the ego motion (i.e., ). To avoid this, we shift the target in the adjacent frame by to remove the fraction of ego-motion.
According to Eq. 2, the learning target is set as the object’s motion in the current frame’s ego coordinate system, which is irrelevant to the ego-motion.
In practice, the alignment operation in Eq. 2 is achieved by feature alignment. Given the candidate features of the previous frame and the current frame , the aligned feature can be obtained by:
Alone with Eq. 3, bilinear interpolation is applied as may not be a valid position in the sparse feature of . The interpolation is a sub-optimal method that will lead to precision degeneration. The magnitude of the precision degeneration is negatively correlated with the resolution of the BEV features. A more precise method is to adjust the coordinates of the point cloud generated by the lifting operation in the view transformer . However, it is deprecated in this paper as it will destroy the precondition of the acceleration method proposed in the naive BEVDet . The magnitude of the precision degeneration will be quantitatively estimated in the ablation study Section. 4.3.2.
Experiment
We conduct comprehensive experiments on a large-scale dataset, nuScenes . nuScenes dataset includes 1000 scenes with images from 6 cameras with surrounding views, points from 5 Radars and 1 LiDAR. It is the up-to-date popular benchmark for 3D object detection and BEV semantic segmentation . The scenes are officially split into 700/150/150 scenes for training/validation/testing. There are up to 1.4M annotated 3D bounding boxes for 10 classes: car, truck, bus, trailer, construction vehicle, pedestrian, motorcycle, bicycle, barrier, and traffic cone. Following CenterPoint , we define the region of interest (ROI) within 51.2 meters in the ground plane with a resolution of 0.8 meters by default.
For 3D object detection, we report the official predefined metrics: mean Average Precision (mAP), Average Translation Error (ATE), Average Scale Error (ASE), Average Orientation Error (AOE), Average Velocity Error (AVE), Average Attribute Error (AAE), and NuScenes Detection Score (NDS). The mAP is analogous to that in 2D object detection for measuring the precision and recall, but defined based on the match by 2D center distance on the ground plane instead of the Intersection over Union (IOU) . NDS is the composite of the other indicators for comprehensively judging the detection capacity. The remaining metrics are designed for calculating the positive results’ precision on the corresponding aspects (e.g., translation, scale, orientation, velocity, and attribute).
Following BEVDet , models are trained with AdamW optimizer, in which gradient clip is exploited with learning rate 2e-4, a total batch size of 64 on 8 NVIDIA GeForce RTX 3090 GPUs. Sublinear memory cost is used for GPU memory management. We apply a cyclic policy , which linearly increases the learning rate from 2e-4 to 1e-3 in the first 40% schedule and linearly decreases the learning rate from 1e-3 to 0 in the remainder epochs. By default, the total schedule is terminated within 20 epochs.
We keep all data processing settings the same as BEVDet . Specifically, we use to denote the width and height of the input image. By default in the training process, the source images with 1600900 resolution are processed by random flipping, random scaling with a range of , random rotating with a range of , and finally cropping to a size of . The cropping is conducted randomly in the horizon direction but is fixed in the vertical direction (i.e., , where and are the upper bound and the lower bound of the target region.) In the BEV space, the input feature and 3D object detection targets are augmented by random flipping, random rotating with a range of , and random scaling with a range of . Following CenterPoint , all models are trained with CBGS . In testing time, the input images are scaled by a factor of and cropped to resolution with a region defined as .
We conduct all experiments based on MMDetection3D . The inference speed is the average upon 6019 validation samples . For monocular paradigms like FCOS3D and PGD , the inference speeds are divided by a factor of 6 (i.e. the number of images in a single sample ), as they take each image as an independent sample. By default, the inference acceleration method proposed in BEVDet is applied.
2 Benchmark Results
We comprehensively compare the proposed BEVDet4D with the baseline method BEVDet and other paradigms like FCOS3D , PGD , DETR3D , PETR Graph-DETR3D and BEVFormer . Their numbers of parameters, computational budget, inference speed, and accuracy on the nuScenes val set are all listed in Tab. 1. Some state-of-the-art methods with other sensors are also listed for comparison like LiDAR-based method CenterPoint and radar-based method CenterFusion .
The high-speed version dubbed BEVDet4D-Tiny scores 47.6% NDS on nuScenes val set, which exceeds the baseline (i.e. BEVDet-Tiny with 39.2% NDS) by a large margin of +8.4% NDS. The improvement in the composite indicator NDS mainly derives from the reduction of the orientation error, the velocity error, and the attribute error. Specifically, benefitting from the well-designed BEVDet4D paradigm, the velocity error is significantly decreased by -62.9% from BEVDet-Tiny 0.909 mAVE to BEVDet4D-Tiny 0.337 mAVE. For the first time, the precision of the velocity prediction in the camera-based methods notably exceeds the CenterFusion 0.540 mAVE, who relies on the multi-sensor fusion with camera and radar for high precision in this aspect. Besides, at a similar inference speed, velocity precision of BEVDet4D-Tiny is also comparable with the state-of-the-art LiDAR-based method PointPillar (i.e. 17.9 FPS and 0.323 mAVE) implemented in . With respect to the orientation prediction, the proposed method also reduces the error in this aspect by -12.0% from BEVDet-Tiny 0.523 mAOE to BEVDet4D-Tiny 0.460 mAOE. This is because the orientation and velocity of the targets are strong-coupled. Analogously, the attribute error is reduced by -25.1% from BEVDet-Tiny 0.247 mAAE to BEVDet4D-Tiny 0.185 mAAE.
While upgrading the paradigm to BEVDet4D-Base analogous to BEVDet-Base , the promotion on the baseline is slightly narrowed to +7.3% on the composite indicator NDS from BEVDet-Base 47.2% NDS to BEVDet4D-Base 54.5% NDS. This surpasses the concurrent work of BEVFormer by +2.8% NDS (i.e. 54.5% NDS v.s. 51.7% NDS), while running faster than it in test time (i.e. 1.9 FPS v.s. 1.7 FPS). With test time augmentation, we further push the performance boundary to 55.2% NDS. It is worth noting that, thanks to the few framework adjustments, BEVDet4D achieves the aforementioned performance improvement at the cost of negligible extra inference latency.
2.2 nuScenes test set
For the nuScenes test set, we train the BEVDet4D-Base configuration on the train and val sets. A single model with test time augmentation is adopted. As listed in Tab. 2, BEVDet4D ranks first on the nuScenes vision-based 3D object detection leader board with a score of 56.9% NDS, substantially surpassing the previous leading method BEVDet by +8.7% NDS. It also exceeds the concurrent work of BEVFormer by +3.4% NDS and significantly exceeds those relied on additional data for pre-training like DD3D , DETR3D , and PETR . Besides the composite indicator, BEVDet4D has leading performance in most other indicators like mAP, mATE, mAOE, mAVE and mATE. With respect to the ability of generalization, the previous leading method BEVDet has merely +0.5% performance growth from val set 47.7% NDS to test set 48.2% NDS. However, with the same configuration, the performance boosting of BEVDet4D is +1.7% NDS from val set 55.2% NDS to test set 56.9% NDS. This indicates that exploiting temporal cues in BEVDet4D can also help improve the models’ generalization performance.
3 Ablation Studies
In this subsection, we empirically show how the robust performance of BEVDet4D is built. BEVDet-Tiny without acceleration is adopted as a baseline. In other words, the spatial alignment operation in BEVDet4D is conducted within the view transformer by adjusting the pseudo point cloud . The results of different configurations are listed in Tab. 3. Some key factors are discussed one by one in the following.
Directly concatenate the current frame feature with the previous one in configuration Tab. 3 (A), the overall performance drops from 39.2% NDS to 37.6% NDS by -1.6%. This modification degrades the models’ performance, especially on the translation and the velocity aspects. We conjecture that, due to the ego-motion, the positional shift of the same static object between the two candidate features will confuse the following modules’ judgment on the object position. With respect to the moving object, it is more complicated for the modules to judge out the velocity target defined in the current frame’s ego coordinate system from the positional shift between the two candidate features which is described in Eq. 1. As to this end, the module needs to remove the fraction of ego-motion from this positional shift and consider the time factor.
By conducting translation-only align operation in configuration Tab. 3 (B), we enable the modules to utilize the position-aligned candidate features to make better perceptions of the static targets. Besides, the velocity predicting task is simplified by removing the fraction of the ego-motion. As result, the translation error is reduced by -5.4% to 0.672, which has surpassed the baseline configuration with a translation error of 0.691. Moreover, the velocity error is also reduced by -23.2% from configuration Tab. 3 (A) 1.544 mAVE to configuration Tab. 3 (B) 1.186 mAVE. However, this velocity error is still larger than that of the baseline configuration. We conjecture that the distribution of the positional shift is far from that of the velocity due to the inconsistent time duration between the two adjacent frames.
Further removing the time factor in configuration Tab. 3 (C), we let the module directly predict the targets’ positional shift in two candidate features. This modification successfully simplifies the learning targets and makes the trained module more robust on the validation set. The velocity error is thus further reduced by a large margin of -59.6% to 0.479 mAVE, which is just 52.7% of the naive BEDVet .
In configuration Tab. 3 (D), we apply an extra BEV encoder before concatenating the two candidate features. This slightly enlarges the computational budget by 2.8%. The change of inference speed is negligible. However, this modification offers comprehensive improvement on the baseline (i.e., Tab. 3 (C)). The overall performance is improved by +0.9% NDS from 44.0% to 44.9%. By adjusting the loss weight of velocity prediction in the training process, configuration Tab. 3 (E) reduces the velocity error to 0.435.
By considering the rotation variance of the ego pose in the align operation, configuration Tab. 3 (F) further reduces the velocity error by 13.6% from 0.435 (i.e., Tab. 3 (E)) to 0.376. This indicates that a precise align operation can help increase the precision of velocity prediction.
To search for the optimal test time interval between the current frame and the reference one, we use the unlabeled camera sweeps with 12Hz instead of the annotated camera frames (2Hz) in configuration Tab. 3 (G). The time interval between two camera sweeps is denoted as . We select three different time intervals in each training configuration and judge the adjusting direction by comparing them in test time. In this way, we can avoid the training disturbance in searching for this hyper-parameter. According to Fig. 4, The optimal interval is around 15T which is set as the test time interval by default in this paper. During the training process, we conduct data augmentation by randomly sampling time intervals within . As a result, configuration Tab. 3 (G) further reduces the velocity error by 12.8% from 0.376 to 0.328.
3.2 Precision Degeneration of the Interpolation
We use configuration Tab. 3 (C) to exploit the precision degeneration of the interpolation operation. Several ablation configurations are constructed in Tab. 4 to study the factors like the BEV resolution and the interpolation operation. When a low BEV resolution of 0.8m0.8m is applied, we observed a slight drop in velocity precision from configuration Tab. 4 (A) 0.479 mAVE to (B) 0.499 mAVE. This indicates that aligning the feature map after the view transformation with interpolation operation will introduce systematic error. However, the precondition of acceleration in BEVDet can be maintained in configuration Tab. 4 (B). Benefitting from the acceleration method, the inference speed can be scaled up to 15.6 FPS, which is twice that of the configuration Tab. 4 (A).
When a high BEV resolution of 0.4m0.4m is applied, the performance difference between aligning within the view transformation and aligning after view transformation with interpolation operation is negligible (i.e. Tab. 4 (C) with 45.2 NDS v.s. Tab. 4 (D) with 45.3 NDS). High BEV resolution can help reduce the precision degeneration caused by the interpolation operation. Besides, from the perspective of inference acceleration, conducting aligning operations within the view transformation is deprecated.
3.3 The Position of the Temporal Fusion
It is not trivial to select the position of the temporal fusion in the BEVDet4D framework. We compare some typical positions in Tab. 5 to study this problem. Among all configurations, conducting temporal fusion after the extra BEV encoder in configuration Tab. 5 (B) is the most applicable one with the lowest velocity error of 0.429 mAVE. When bringing forward the temporal fusion in configuration Tab. 5 (A), the velocity error is increased by +11.9% to 0.480 mAVE. This indicates that the BEV feature generated by the view transformer is too coarse to be directly applied. An extra BEV encoder before temporal fusion can help alleviate this problem. When we postpone the temporal fusion to the back of the BEV encoder in configuration Tab. 5 (C), the overall performance degenerates to 39.4% NDS which is close to the baseline BEVDet with 39.2% NDS. More precisely, the feature from the previous frame helps slightly reduce the velocity error from 0.909 mAVE to 0.838 but increases the translation error from 0.691 mATE to 0.720 mATE. This indicates that the BEV encoder plays an important role in effectuating the proposed BEVDet4D paradigm by resisting the positional misleading from the previous frame feature and estimating the velocity according to the difference between the two candidate features.
Conclusion
We pioneer the exploitation of vision-based autonomous driving in the spatial-temporal 4D space by proposing BEVDet4D to lift the scalable BEVDet from spatial-only 3D working space into spatial-temporal 4D working space. BEVDet4D retains the elegance of BEVDet while substantially pushing the performance in multi-camera 3D object detection, particularly in the velocity prediction aspect. Future works will focus on the design of framework and paradigm for actively mining the temporal cues.