FB-BEV: BEV Representation from Forward-Backward View Transformations
Zhiqi Li, Zhiding Yu, Wenhai Wang, Anima Anandkumar, Tong Lu, Jose M. Alvarez
Introduction
The most intuitive method for projecting camera features onto the BEV plane involves estimating the depth value of each pixel in the image and using the camera calibration parameters to determine the corresponding position of each pixel in 3D space , as shown in Figure 1 (left). We refer to this process as forward projection, where the 2D pixels take the initiative in projection and the 3D space passively accepts features from the images. The accuracy of the predicted depth for each pixel is critical to achieving high-quality BEV features. However, accurately estimating the depth value of each pixel is challenging . To address this challenge, Lift-Splat-Shoot (LSS) pioneered the use of depth distribution to model the uncertainty of each pixel’s depth . One limitation of LSS is that it generates discrete and sparse BEV representation . As shown in Figure 2, the density of BEV features decreases with distance. When using the default settings of LSS on the nuScenes dataset, only 50% of the grids can receive valid image features through projection.
The motivation behind backward projection is opposite to that of forward projection. For the backward projection paradigm, the points in 3D space take the initiative . For instance, BEVFormer sets the coordinates of the 3D space to be filled in advance and then projects these 3D points back onto the 2D image , as shown in Figure 1 (middle). As a result, each predefined 3D space position can obtain its corresponding image features. The BEV representation obtained by this method is denser than that of LSS, with each BEV grid filled with the corresponding image features.
The drawbacks of backward projection are also apparent, as shown in Figure 3. Although yielding a denser BEV representation, it comes at the cost of establishing numerous false correspondences between 3D and 2D space due to occlusion and depth mismatch . The absence of depth information during the projection process is the main cause. Without depth as a reference, each 3D coordinate on the ray is equally related to the same 2D coordinate, equivalent to having a uniform depth distribution for this pixel in forward projection. As a result, the distance prediction of the objects along the longitudinal direction become ambiguous. Backward projection thus tends to be inferior to forward projection in depth utilization. Recently, the advantage of forward projection has been further highlighted since more accurate depth distribution obtained from depth supervision is shown to improve 3D perception .
Considering the pros and cons discussed above, we propose forward-backward view transformation to address the limitations of existing VTMs, as shown in Figure 1 (right). To address the issue of sparse BEV representations in forward projection, we leverage backward projection to refine the sparse region from forward projection. Meanwhile, backward projection is prone to false-positive features due to the lack of depth guidance. We thus propose a depth-aware backward projection design to suppress false-positive features by measuring the quality of each projection relationship through depth consistency. The depth consistency is determined by the distance of depth distributions between a 3D point and its corresponding 2D projection point. Using this depth-aware method, unmatched projections are given lower weights, which reduces the interference caused by false-positive BEV features. In addition, for the objection detection task, we only care about foreground objects, so we densify only the foreground regions of the BEV plane while using backward projection. This not only reduces the computational burden but also avoids the introduction of false-positive features in the background areas. With the sparse regions refined for forward projection and false-positives features reduced for backward projection, our forward-backward projection not only solves the defects of existing projection methods but also realizes the effective ensemble of existing projection methods. Our contributions can be summarized as follows:
We propose a forward-backward projection strategy that generates dense BEV features with strong representation ability through bidirectional projection. Our approach addresses the limitations of existing projection methods, which result in either sparse BEV features or false-positive features caused by inaccurate projection.
To address the pitfalls of existing forward projection methods for producing sparse BEV representations, we employ backward projection to refine the blank grid that not be activated by forward projection. This makes the model more suitable for large-scale BEVs.
We propose a novel depth-aware backward projection method that overcomes the limitations of existing methods in effectively utilizing depth information. Our approach integrates depth consistency into the projection process to establish a more accurate mapping relationship between the 3D and 2D spaces.
Our FB-BEV model has been extensively evaluated on the nuScenes dataset. The results demonstrate that it outperforms other methods for camera-based 3D object detection and achieves the state-of-the-art 62.4% NDS on the nuScenes test set.
Related work
We introduce related BEV perception works according to the VTMs they use.
The Lift-Splat-Shoot (LSS) method is the archetypal technique of this category. LSS utilizes a depth distribution to model depth uncertainty and project multi-view features into the same Bird’s Eye View (BEV) space. Subsequent methods have largely adhered to this paradigm. For instance, BEVDet applies this forward projection approach to the field of multi-view 3D detection. CaDDN and BEVDepth proposes the use of LiDAR point clouds to generate depth ground truth for supervising the depth prediction module. BEVDepth demonstrates that an accurate depth prediction module can significantly enhance model performance. Similarly, BEVstereo further underscores the importance of precise depth estimation to model performance. Furthermore, BEVFusion extends this paradigm to the multi-modality perception domain and improves the projection efficiency of the LSS paradigm. The most notable disadvantage of VTM in LSS is low efficiency. Subsequent research has made significant progress in improving efficiency through engineering implementation . In response to the sparseness of BEV features, MatrixVT mainly focuses on improving the calculation efficiency in the process of BEV generation, rather than densely stressing BEV features.
2 Backward Projection Methods
OFT is among the first methods to adapt the backward projection paradigm. This paradigm does not involve complex accumulation in 3D space , which is the least efficient step in forward projection. Subsequent works such as ImVoxelNet and M2BEV extend this paradigm from monocular to multi-view perception where 3D space is divided into voxels. DETR3D does not introduce dense BEV features, but performs end-to-end learning of object queries in 3D and projects object centers back to image space. BEVFormer aggregates features at different heights on the BEV space without introducing voxelized representation, therefore reducing the resource consumption. BEVFormer also introduces deformable sampling points and temporal features, promoting further development of camera-based perception. For the perception heads, BEVFormer adopts Deformable DETR and Panoptic SegFormer . BEVFormerV2 further exploits the potential of backward projection by adapting the modern image backbone via perspective supervision. PolarFormer and PolarDETR adopt polar coordinates rather than Cartesian coordinates to conduct the projection process. Methods project 3D anchors onto 4D features rather than 3D features. PersFormer uses Inverse Perspective Mapping (IPM) to guide the projection point on the image space. However, existing methods seldom consider introducing depth in the projection process or even consider getting rid of the dependence on depth as an advantage . We argue that it increases ambiguity in the projection process without depth to measure the quality of the projection.
In addition to different view projection paradigms, researchers have also explored using longer temporal information to enhance the spatial perception capacity. .
3 Projection-Free Methods
In addition to the above two paradigms, some methods can generate BEV representations without relying on projections. PETR and PETRv2 implicitly learn the view transformation through global attention and use camera parameters to encode position features. CFT uses view-aware attention to adaptively learn the BEV features required for each view, and even get rid of the dependence on camera calibration parameters. BEVSegFormer automatically learns the correspondence between 3D and 2D space without relying on the projection process.
Method
To address the limitations of existing view transformation modules, we propose a novel Forward-Backward View Transformation method named FB-BEV. FB-BEV employs a two-pronged approach. Firstly, a VTM based on forward projection will generate an initial sparse BEV representation. To obtain a denser BEV representation while minimizing the computational burden, a foreground region proposal network is employed to select the foreground BEV grids. Subsequently, another VTM utilizes these foreground grids as BEV queries and refines them by projecting them back onto the images with a depth-aware mechanism.
As illustrated in Figure 4, FB-BEV mainly consists of three key modules: a view transformation module with forward projection denoted as F-VTM, a foreground region proposal network denoted as FRPN, and a view transformation module with depth-aware backward projection denoted as B-VTM. In addition, we have a depth net to predict the depth distributions, and the distributions will be utilized in both VTMs. F-VTM generates a complete BEV representation from the multi-view features by projecting each pixel feature into the 3D space based on the corresponding depth distribution. FRPN is a lightweight binarized mask predictor used to select the regions where the foreground object is located. B-VTM is only responsible for optimizing BEV grids located in the foreground region generated by FRPN.
2 Forward Projection
Our forward projection module F-VTM follows the paradigm of LSS . Lift and Splat are two fundamental steps in modern forward projection techniques used for view transformation. The Lift step projects each pixel in the 2D image onto the 3D voxel space based on its corresponding depth distribution. The Splat step aggregates the feature values of pixels within each voxel by sum pooling. For specific implementation, our F-VTM is based on BEVDet and BEVDepth , which represent the current state-of-the-art design of forward projection. We denote the BEV features from F-VTM as .
3 Foreground Region Proposal Network
During the inference phase, with BEV feature from F-VTM as input and the predicted binary mask , we filter out unnecessary BEV grids with a mask logit lower than threshold . Thus we obtain a set of discrete BEV grids , where is the location of each foreground BEV grid. Each BEV grid can be seen as a BEV query that requires further refinement. To maintain feature consistency in the foreground area, we have selected BEV grids that contain both blank and non-blank grids.
4 Depth-Aware Backward Projection
The depth-aware backward projection module serves a dual purpose. Firstly, it effectively fills the BEV with arbitrary resolution and can choose only to generate BEV features of specified regions, thereby compensating for the sparse features generated by forward features. Secondly, when combined with a forward projection method, they provide a more comprehensive BEV representation. In this section, we first introduce the depth consistency used to improve the quality of backward projection in 3.4.1 and then introduce our detailed implementation in 3.4.2
Forward projection alleviates this problem by predicting different weights for different depths. Specifically, for each point , it predict a weight for each discrete depth , and , is a set of discrete depths, is the initial depth and is the depth interval. Thus, while considering two discrete depth and on point , it falls onto the 3D points and based on Equation 1. The forward projection method leverages predicted depth weights and to generate distinguishing features.
where and . The depth consistency introduced in this paper serves a similar role as the depth weight in forward projection. It is worth mentioning that we obtain the depth distribution of point via bilinear projection.
Forward projection employs discrete depth values to generate corresponding discrete 3D projection points in 3D space. While sacrificing continuity in depth, the accuracy of BEV features by forward projection is also affected. For our depth-aware backward projection, we guarantee the ability to densely fill 3D space at arbitrary resolutions, while leveraging depth consistency to guarantee projection quality.
4.2 Implementation
Depth consistency is a general mechanism that can be plugged into any existing backward projection method. In this paper, our depth-aware backward projection is based on the spatial cross-attention in BEVFormer . The projection process of the original Spatial Cross-Attention (SCA) of BEVFormer can be formulated as:
where is one BEV query that located at and are multi-view features. For each point on the BEV plane, BEVFormer will lift this point up to 3D points with different heights . Projection function will get the projection point on -th image based on Equation 1. is the deformable attention function that using query to sample features of the projection point on image feature .
By using the depth consistency of this paper, we can directly evolve SCA into Depth-Aware SCA (SCAda) by:
where is the depth consistency between 3D point and 2D point . Compared to the original SCA, our proposed SCAda is capable of generating more discriminative BEV features along the longitudinal direction. Due to the high efficiency of our depth-aware SCA, we only use back projection once instead of stacking 6 layers used by the original BEVFormer .
Experiments
Implementation Details. By adhering to common practices , we default to using ResNet-50 and an image size of . During training, we adopt the CBGS strategy and apply data augmentations at both the image and BEV levels, which include random scaling, flipping, and rotation as per BEVDet . By default, our model is trained for 20 epochs using a batch size of 64 and the AdamW optimizer with a learning rate of 2e-4. While training FB-BEV with V2-99 backbone for test set, we train the model with 30 epochs without CBGS. For training the depth net with temporal information, we use the camera-aware Depth Net in BEVDepth with a total of 118 depth categories (), and . Incorporating depth-aware spatial cross-attention, we sample the predefined heights uniformly from [-5m, 3m], use 8 attention heads, and set . The spatial shape of BEV grids is, by default, with a channel dimension of 256. The threshold for the foreground mask is set to . When introducing temporal information, we stack the BEV features of two adjacent keyframes, as done in BEVDet4D for val set, and 9 previous keyframes for test set.
2 Baselines
To assess the efficacy of our novel approach, FB-BEV, we conduct comparisons with two types of baselines that solely rely on forward and backward projection techniques, respectively. Notably, for these baselines, we maintain consistency with FB-BEV in terms of backbone, detection head, and training strategy, with the exception of the view transformation module. We reduce the number of channels and layers of FB-BEV to match the computational cost.
Forward Projection. For forward projection methods, we adopt BEVDet and BEVDepth as our baseline. Compared to BEVDet, BEVDepth uses point clouds to generate the ground truth of depth and train the depth net with the ground truth of depth.
Backward Projection. For backward projection methods, we choose BEVFormer as the baseline. Considering the difference in implementation details, we ported the view transformation module of BEVFormer to BEVDet for a fair comparison. It is worth mentioning that we discard the temporal self-attention module in BEVFormer. In this paper, we note the BEVFormer that with temporal information as BEVFormer-T.
3 Benchmark Results
Table 1 shows the 3D detection results on the nuScenes val set for our proposed FB-BEV method, as well as the two baseline methods BEVDet and BEVFormer , and other previous state-of-the-art 3D detection methods. Without using temporal information or depth supervision, our method outperforms BEVDet and BEVFormer by a significant margin of 2.4% NDS and 2.7% NDS. When introducing temporal information by stacking historical BEV features, our proposed FB-BEV still outperforms BEVDet and BEVFormer by 1.3 points. With depth supervision, our method achieves a lead of more than 1.5 points over BEVDepth. However, as the previous backward projection cannot use depth information, BEVFormer-T only brings a marginal improvement of 0.2% NDS when only using depth supervision as an auxiliary task. This confirms the limitations of existing backward projection methods. Despite achieving higher performance, our method still maintains a comparable or even lower computational cost than our baselines. As shown in Table 2, our model obtains a new state-of-the-art 62.4% NDS and outperforms previous SOLOFusion with a clear margin of 0.5 points.
4 Ablation Studies
Depth-aware Backward Projection. In Table 3, we compare the results of adopting depth-aware backward projection in FB-BEV and BEVFormer. BEVFormer obtains an improvement of 0.9% NDS with depth-aware projection. When using depth supervision as an auxiliary task, BEVFormer-T achieves a larger gain of 1.1% NDS. Without depth-aware backward projection in FB-BEV, the performance drops by about 0.9% NDS. In the past, only forward projection methods could benefit from more accurate depth prediction. With depth consistency, backward methods can also improve performance by leveraging accurate depth prediction.
Figure 6 presents visual results of FB-BEV with and without depth-aware backward projection. When the depth-aware projection is not employed, the model tends to produce incorrect results along the longitudinal direction due to depth ambiguities, as seen in the yellow boxes in Figure 6 (b). In addition, Figure 6 (c) and (d) show the depth consistency on the BEV plane for FB-BEV with and without depth-aware projection. Foreground grids exhibit higher depth consistency, which prevents background regions from erroneously identifying false foreground features. Moreover, Figure 6 (d) shows that the depth consistency varies with height for the same location on the BEV plane. Prior backward projection methods aggregated features at all heights, resulting in feature interference. However, with our proposed depth-aware backward projection, the model selectively aggregates features based on depth consistency at different heights. These visualizations provide compelling evidence for the effectiveness of our method.
Effect of FRPN. We employ FRPN to selectively optimize the foreground grids in BEV feature through B-VTM. To study its effectiveness, we conduct an experiment where we exclude FRPN and instead feed all BEV features into B-VTM. Results in Table 4 demonstrate that FRPN not only improves the inference efficiency but also improves the detection accuracy. In Figure 7, we present the depth consistency map of FB-BEV with and without FRPN. Without using FRPN, depth-aware backward projectionmay focus on some background regions due to imprecise depth predictions. On the other hand, using the foreground mask provided by FRPN, the model can selectively concentrate only on the foreground objects, thus avoiding interference from the background regions.
Effect of Reducing Sparsity. Due to the fixed discrete depth values, the forward projection method generates fixed discrete 3D projection points (projection matrix remains unchanged). As the BEV scale increases, the proportion of blank grids on BEV generated by forward projection will also increase. The rate of blank gird of BEVDet with a BEV scale and input shape is 80.5%. Thus we can observe that BEVDet performance drops on large-scale BEV. Our forward-backward projection fixes it by filling these blank grids, then obtains continuous performance gains. In addition, we can observe that current VTMs are highly efficient and are not considered a potential bottleneck against inference efficiency.
Conclusion
We present a forward-backward projection paradigm to address the limitations of current projection schemes. Our approach addresses the issue of sparse features generated by forward projection and introduces depth into backward projection to establish a more precise projection relationship. This two-stage VTM strategy is suitable for higher-resolution BEV perception and has application prospects for ultra-long-distance object detection or high-resolution occupancy perception.