MatrixVT: Efficient Multi-Camera to BEV Transformation for 3D Perception
Hongyu Zhou, Zheng Ge, Zeming Li, Xiangyu Zhang
Introduction
Vision-centric 3D perception in Bird’s-Eye-View (BEV) has recently drawn extensive attention. Apart from their outstanding performance, the compact and unified feature representation in BEV facilitates straight-forward feature fusions , and enables various downstream tasks (e.g. object detection , map segmentation , motion planning, etc.) to be applied thereon easily.
View Transformation (VT) is the key component that converts multi-camera features to BEV, which has been heavily studied in previous works . Existing VT methods can be categorized into geometry-based and learning-based methods . Among these two categories, geometry-based methods show superior performance due to the use of geometric constraints. Lift-Splat , as a representative geometry-based VT, predicts categorical depth distribution for each pixel and “lift” the corresponding features into 3D space according to the predicted depth. These feature vectors are then gathered into pre-defined grids on a reference BEV plane (i.e., “splat”) to form the BEV feature (Fig. 1, upper). The Lift-Splat-based VT has shown great potential to produce high-quality BEV features, achieving remarkable performance on object detection and map segmentation tasks on the nuScenes benchmark .
Despite the effectiveness of Lift-Splat-like VT , two issues remain. First, the “splat” operation is not universally feasible. Existing implementations of “splat” relies on either the “cumsum trick” that is highly inefficient, or customized operators that can only be used on specific devices, increases the cost of applying BEV perception. Second, the size of “lifted” multi-view image features is huge, becoming the memory bottleneck of BEV models. These two issues lead to a heavy burden on BEV methods during both the training and inference phases. As a result, the drawbacks of existing view transformers limit the broad application of autonomous driving technology
In this work, we propose a novel VT method, MatrixVT, to address the above problems. MatrixVT is proposed based on the fact that the VT can be viewed as a feature transportation process. In that case, the BEV feature can be viewed as the MatMul between the “lifted” feature and a transporting matrix, namely Feature Transporting Matrix (FTM). We thus generalize the Lift-Splat VT into a purely mathematical form and eliminate specialized operators.
However, transformation with FTM is a kind of degradation — the mapping between the 3D space and BEV grids is extreme sparse, leading to the huge size of FTM and poor efficiency. Prior works seek customized operators, successfully avoiding such sparsity. In this paper, we argue that there are other solutions to the problem of sparse mapping. First, we propose Prime Extraction. Motivated by an observation that the height (vertical) dimension of images is less informative in autonomous driving (see Sec. 3.2), we compress the image features along this dimension before VT. Second, we adopt matrix decomposition to reduce the sparsity of FTM. The proposed Ring & Ray Decomposition orthogonally decomposes the FTM into two separate matrices, each encoding the distance and direction of the ego-centric polar coordinate. This decomposition also allows us to reformulate our pipeline into a mathematically equivalent but more efficient one (Fig. 1, lower). These two techniques reduce memory footprint and calculation during VT by hundreds of times, enabling MatrixVT to be more efficient than existing methods.
The proposed MatrixVT inherits the advantages of the Lift-Splat paradigm while being much more memory efficient and fast. Extensive experimental results show that MatrixVT is 2-to-8 times faster than previous methods and saves up to 97% memory footprint among different settings. Meanwhile, the perception model with MatrixVT achieves 46.6% mAP and 56.2% NDS for object detection and 46.2% mIoU for vehicle segmentation on the nuScenes val set, which is comparable to the state-of-the-art performance . We conclude our main contributions as follows:
We propose a new description of multi-camera to BEV transformation — using the Feature Transportation Matrix (FTM), which is a more general representation.
To solve the sparse mapping problem, we propose Prime Extraction and the Ring & Ray Decomposition, boosting VT with FTM by a huge margin.
Extensive experiments demonstrate that MatrixVT yields comparable performance to the state-of-the-art method on the nuScenes object detection and map segmentation tasks while being more efficient and generally applicable.
Related Works
Perception of 3D objects and scenes takes a key role in autonomous driving and robotics, thus attracting increasing attention nowadays. Camera-based perception is the most commonly used method for varies scenarios due to its low cost and high accessibility. Comparing with 2D perception (object detection , semantic segmentation , etc.), 3D perception requires additional prediction of the depth information which is an naturally ill-posed problem .
Existing works either predict the depth information explicitly or implicitly. FCOS3D simply extend the structure of the classic 2D object detector , predicting pixel-wise depth explicitly using an extra sub-net that is supervised by LiDAR data. CaDDN propose treating depth prediction as classification task rather than regression task, and project image feature into Bird’s-Eye-View (BEV) space to achieve unified modeling of detection and depth prediction. BEVDepth and following works propose several techniques to enhance depth prediction, these works achieve outstanding performance due to precise depth. Meanwhile, methods like PON and BirdGAN use pure neural networks to transform image features into BEV space, learning object depth implicitly using segmentation or detection supervision. Currently, methods that explicitly learn depth show prior performance than implicit approaches thanks to the supervision from LiDAR data and depth modeling. In this paper, we use the DepthNet same as in BEVDepth for high-performance.
2 Perception in Bird’s-Eye-View
The concept of BEV is firstly proposed for processing LiDAR point cloud , and found effective for fusing multi-view image features. The core component of vision-based BEV paradigm is the view-transformation. OFT firstly propose mapping image feature from Perspective View into BEV using camera parameters. This method project a reference point from an BEV grid to image plane, and sample corresponding features back to the BEV grid. Following this work, BEVFormer propose using Deformable Cross Attention to sample features around the reference point. These methods fail to distinguish BEV grids that are projected to same position on the image plane, thus show inferior performance than depth-based methods.
Depth-based methods, represented by LSS and CaDDN , predict categorical depth for each pixel, the extracted image feature on a specific pixel is then projected into 3D space by doing per-pixel outer product with corresponding depth. The projected high-dimensional tensor is then “collapsed” to a BEV reference plane using convolution , Pillar Pooling , or Voxel Pooling . Lift-Splat based methods show outstanding performance for down-stream tasks, but introduces two extra problems. Firstly, the intermediate representation of image feature is large and in-efficient, making training and application of these methods difficult. Secondly, the Pillar Pooling introduces random memory access, which is slow and device demanding (extremely slow on general-purpose devices). In this paper, we propose a new depth-based view transformation to overcome these problems while retaining the ability of producing high-quality BEV features.
MatrixVT
Our MatrixVT is a simple view transformer based on the depth-based VT paradigm. In Sec. 3.1, we first revisit existing Lift-Splat transformation and introduce the concept of Feature Transporting Matrix (FTM) together with the sparse mapping problem. Then, techniques proposed to solve the problem of sparse mapping i.e., Prime Extraction (Sec. 3.2) and Ring & Ray Decomposition (Sec. 3.3), are introduced. In Sec. 3.4. we designate the novel VT method utilizing the aforementioned techniques as MatrixVT, and elaborate its overall pipeline.
The intermediate tensor can be treated as feature vectors, each vector corresponding to a geometric coordinate. The “splat” operation is then adopted using operators like Pillar Pooling , during which each feature vector is summed to a BEV grid according to geometric coordinates (see Fig. 2).
Replacing the “splat” operation with FTM eliminates the need for customized operators. However, the sparse mapping between the 3D space and BEV grids can lead to a massive and highly sparse FTM, harm the efficiency of matrix-based VT. To address the sparse mapping problem without customized operators, we propose two techniques to reduce the sparsity of FTM and speed up the transformation.
2 Prime Extraction for Autonomous Driving
The high-dimensional intermediate tensor is the primary cause of the sparse mapping and the low efficiency of Lift-Splat-like VT. An intuitive way to reduce sparsity is reducing the size of . Therefore, we propose Prime Extraction — a compression technique for autonomous driving and other scenarios where information redundancy exists.
The prime extraction is motivated by an observation: The image feature’s height dimension has a lower response variance than the width dimension. This observation indicates that this dimension contains less information than the width dimension. We thus propose to compress image features on the height dimension. Previous works have also exploited reducing the height dimension of image features , but we firstly propose compressing both image features and corresponding depth to boost VT.
In Sec. 4.4, we will show that the extracted Prime Feature and Prime Depth effectively retain valuable information from the raw feature and produce BEV features of the same high quality as the raw feature. Moreover, the Prime Extraction technique can be individually adopted to existing Lift-Splat-like VTs to enhance their efficiency at almost no performance cost.
3 “Ring and Ray” Decomposition
However, this decomposition do not reduce the FLOPs during VT and introduces the Intermediate Feature in Fig. 5 (blue) whose size if huge and depends on feature channel . To reduce the calculation and memory footprint during VT, we combine Eq. 3 to Eq. 5 and rewrite them in a mathematically equivalent form (see Appendix 1.2 for proof):
With this reformulation, we reduce the calculation during VT from to ; the memory footprint is also reduced from to . Under common setting (), the Ring & Ray Decomposition reduces calculation by 46 times and saves 96% memory footprint.
4 Overall Pipeline
With above techniques, we reduce the calculation and memory footprint using FTM by hundreds times, making VT with FTM not only feasible but also efficient. Given the multi-view images for a specific scene, we conclude the overall pipeline of MatrixVT as follows:
We first use an image backbone to extract image features from each image.
Then, a depth predictor is adopted to predict categorical depth distribution for each feature pixel to obtain the depth prediction.
After that, we send each image feature and corresponding depth to the Prime Extraction module, obtaining the Prime Feature and the Prime Depth, which is the compressed feature and depth.
Finally, with the Prime Feature, Prime Depth, and the pre-defined Ring & Ray Matrices, we use Eq. 6 (see also Fig. 1, lower) to obtain the BEV feature.
Experiments
In this section, we compare the performance, latency, and memory footprint of MatrixVT and other existing VT methods on the nuScenes benchmark .
We conduct our experiments based on BEVDepth , the current state-of-the-art detector on the nuScenes benchmark. In order to conduct a fair comparison of performance and efficiency, we re-implement BEVDepth according to their paper. Unless otherwise specified, we use ResNet-50 and VoVNet2-99 pre-trained on DD3D as the image backbone and SECOND FPN as the image neck and BEV neck. The input image adopts pre-processing and data augmentations same as in . We use BEV feature size for low input resolution and for high resolution on detection. Segmentation experiments use BEV resolution as in LSS . We use the DepthNet to predict categorical depth from 2m to 58m in nuScenes, with uniform 112 division. During training, CBGS and model EMA are adopted. Models are trained to converge since MatrixVT converges a little slower than other methods, but no more than 30 epochs.
2 Comparison of Performances
We conduct experiments under several settings to evaluate the performance of MatrixVT in Tab. 1. We first adopt the ResNet family as the backbone without applying multi-frame fusion. MatrixVT achieves 33.6% and 49.7% mAP with ResNet-50 and ResNet-101, which is comparable to the BEVDepth and surpasses other methods by a large margin. We then test the upper bound of MatrixVT by replacing the backbone with V2-99 pre-trained on external data and applying multi-frame fusion. Under this setting, MatrixVT achieves 46.6% mAP and 56.2% NDS, which is also comparable to the BEVDepth.
2.2 Map Segmentation
We also conduct experiments on map segmentation tasks to validate the quality of the BEV feature generated by MatrixVT. To achieve this, we simply put a U-Net-like segmentation head on the BEV feature. For fair comparison, we put the same head on the BEV feature of BEVDepth for experiments, and results are reported as “BEVDepth-Seg” in Tab. 2. It is worth noting that previous works achieve the best segmentation performance under different settings (different resolution, head structure, etc.); thus, we report the highest performance of each method. As can be seen from Tab. 2, the map segmentation performance of MatrixVT surpasses most existing methods on all three sub-tasks and is comparable to our baseline, BEVDepth.
3 Efficient Transformation
We compare the efficiency of View Transformation in two dimensions: Speed and Memory Consumption. Note that we measure the latency and memory footprint (using fp32) of view transformers since these metrics are affected by backbone and head design. We take the CPU as a representative general-purpose device where customized operators are unavailable. For a fair comparison, we measure and compare the other two Lift-Splat-like view transformers. The LS-BEVDet is the accelerated transformation used in BEVDet (with default parameters); the LS-BEVDepth uses the CUDA operator proposed in BEVDepth that is not available on CPU and other platforms. To demonstrate the characteristics of different methods, we define six transformation settings, namely S1S6, varies in image feature size, BEV feature size, and feature channels that are closely related to model performance.
As can be seen in Fig. 6, the proposed MatrixVT boost the transformation significantly on the CPU, being 4 to 8 times faster than the LS-BEVDet. On the CUDA platform , where customized operators enable faster transformation, MatrixVT still shows a much faster speed than the LS-BEVDepth under most settings. Besides, we calculate the number of intermediate variables during VT as an indicator of extra memory footprint. For MatrixVT, these variables include the Ring Matrix, Ray Matrix, and the intermediate matrix; for Lift-Splat, intermediate variables include the intermediate tensor and the predefined BEV grids. As illustrated in Fig. 6 (right), MatrixVT consumes 2 to 15 times less memory than LS-BEVDepth and 40 to 80 times less memory than LS-BEVDet.
4 Effectiveness of Prime Extraction
In this section, we validate the effectiveness of Prime Extraction by performance comparison and visualization.
As mentioned in Sec. 3.2, we argue that the features and corresponding depths can be compressed with little or no information loss. We thus individually adopt the Prime Extraction onto the BEVDepth to verify its effectiveness. Specifically, we compress image features and the depths before the common Lift-Splat. As shown in Tab. 3, the BEVDepth with Prime Extraction achieves 41.1% NDS with Res-50 and 56.1% NDS with VovNetv2-99 , which is comparable to the baseline without compression. Thus, we argue that Prime Extraction effectively retained key information from raw features.
4.2 Prime Information in Object Detection
We then delve into the mechanism of Prime Extraction by visualization. Fig. 7 shows the inputs and outputs of the Prime Extraction module. It can be seen that the Prime Extraction module trained on different tasks focusing on different information. The Prime Depth Attention in object detection focuses on foreground objects. Thus the Prime Depth retained the depth of objects while ignoring the depth of the background. Also, it can be seen from the yellow car in the second column, which is obscured by three pedestrians. Prime Extraction effectively distinguishes these objects at different depths.
Tab. 3 shows the effect of adopting Prime Extraction on the BEVDepth . We do the per-pixel outer product of the Prime Feature and Prime Depth, then apply Voxel Pooling on the obtained tensor. The improved version of BEVDepth saves about 28% percent of memory consumption while offering comparable performance.
4.3 Prime Information in Map Segmentation
For the map segmentation task, we take lane segmentation as an example — the Prime Extraction module focus on lane and road that is closely related to this task. However, since the area of the target category is a wide range covering several depth bins, the distribution of Prime Depth is uniform in the target area (see Fig. 7, 3rd and 4th columns). The observation indicates that Prime Extraction generates a new form of depth distribution that fits the map segmentation task. With the Prime Depth that is rather uniform, the same feature can be projected to multiple depth bins since they are occupied by the same category.
5 Extraction of Prime Feature
In the PFE, we propose using max pooling followed by several 1D convolutions to reduce and refine the image feature. Before reduction, the coordinate of each pixel is embedded into the feature as position embedding. We conduct experiments to validate the contribution of each design.
A possible alternative of the max pooling is the CollapseConv as in and , which merges the height dimension into the channel dimension and reduces the merged channel by linear projection. However, the design of CollapseConv brings some disadvantages. For example, the merged dimension is of size , which is high and requires extra memory transportation. To address these problems, we propose using max pooling to reduce the image feature in Prime Extraction. Tab. 4 shows that reduction using max pooling achieves even better performance than CollapseConv while eliminating these shortcomings.
We also conduct experiments to show the effectiveness of the refine sub-net after reducing in Tab. 4. The results indicate that the refine sub-net plays a vital role in adapting the reduced feature to BEV space, without which a performance drop of 0.9% mAP will occur. Finally, the experimental results in Tab. 4 have shown that the position embedding brings an improvement of 0.4% mAP, which is also preferable.
Conclusion
This paper proposes a new paradigm of View Transformation (VT) from multi-camera to Bird’s-Eye-View. The proposed method, MatrixVT, generalizes VT into a feature transportation matrix. We then propose Prime Extraction, which eliminates the redundancy during transformation, and the Ring & Ray Decomposition, which simplifies and boost the transformation. While being faster and more efficient on both specialized devices like GPU and general-purpose devices like CPU, our extensive experiments on the nuScenes benchmark indicate that the MatrixVT offers comparable performance to the state-of-the-art method.
References
Appendix
Here, we give a brief description of symbols used in the below sections in Tab. 5.
A1.1 Generation of Ring & Ray Matrices
The generation of the Ring Matrix and the Ray Matrix relies on the intrinsic and extrinsic parameters of the camera setting. These parameters determine the geometrical relationship between the “lifted” features and the BEV grids. Note that these matrices need to be generated only once for real-world applications - as the camera positions are usually fixed.
To simplify the algorithm, we take the geometry (2D coordinate, ) of the “lifted” Prime Feature and the BEV grids (2D grids, ), instead of intrinsic and extrinsic parameters as input. One of the algorithms that generate these matrices is described in Alg. 1.
A1.2 Pipeline Reformulation
In MatrixVT, we reformulate our pipeline into a new form to eliminate huge intermediate tensors. Now we prove that these two forms are mathematically equivalent. For clarity, we mark the shape of variables in their upper right corner.
Taking Eq. 8 into Eq. 7, the prime feature can be derived as: