CAPE: Camera View Position Embedding for Multi-View 3D Object Detection
Kaixin Xiong, Shi Gong, Xiaoqing Ye, Xiao Tan, Ji Wan, Errui Ding, Jingdong Wang, Xiang Bai
Introduction
3D perception from multi-view cameras is a promising solution for autonomous driving due to its low cost and rich semantic knowledge. Given multiple sensors equipped on autonomous vehicles, how to perform end-to-end 3D perception integrating all features into a unified space is of critical importance. In contrast to traditional perspective-view perception that relies on post-processing to fuse the predictions from each monocular view into the global 3D space, perception in the bird’s-eye-view (BEV) is straightforward and thus arises increasing attention due to its unified representation for 3D location and scale, and easy adaptation for downstream tasks such as motion planning.
The camera-based BEV perception is to predict 3D geometric outputs given the 2D features and thus the vital challenge is to learn the view transformation relationship between 2D and 3D space. According to whether the explicit dense BEV representation is constructed, existing BEV approaches could be divided into two categories: the explicit BEV representation methods and the implicit BEV representation methods. The former constructs an explicit BEV feature map by lifting the 2D perspective-view features to 3D space . The latter mainly follow DETR-based approaches in an end-to-end manner. Without projection or lift operation, those methods implicitly encode the 3D global information into 3D position embedding (3D PE) to obtain 3D position-aware multi-view features, which is shown in Figure 1(a).
Though learning the transformation from 2D images to 3D global space is straightforward, we reveal that the interaction in the global space for the query embeddings and 3D position-aware multi-view features hinders performance. The reasons are two-fold. For one thing, defining each camera coordinate system as the 3D local space, we find that the view transformation couples the 2D image-to-local transformation and the local-to-global transformation together. Thus the network is forced to differentiate variant camera extrinsics in the high-dimensional embedding space for 3D predictions in the global system, while the local-to-global relationship is a simple rigid transformation. For another, we believe the view-invariant transformation paradigm from 2D image to 3D local space is easier to learn, compared to directly transforming into 3D global space. For example, though two vehicles in two views have similar appearances in image features, the network is forced to learn different view transformations, as depicted in Figure 2 (a).
To ease the difficulty in view transformation from 2D image to global space, we propose a simple yet effective approach based on local view position embedding, called CAPE, which performs 3D position embedding in the local system of each camera instead of the 3D global space. As depicted in Figure 2 (b), our approach learns the view transformation from 2D image to local 3D space, which eliminates the variances of view transformation caused by different camera extrinsics.
Specially, as for key 3D PE, we transform camera frustum into 3D coordinates in the camera system using camera intrinsics only, then encoded by a simple MLP layer. As for query 3D PE, we convert the 3D reference points defined in the global space into the local camera system with camera extrinsics only, then encoded by a simple MLP layer. Inspired by , we obtain the 3D PE with the guidance of image features and decoder embeddings, for keys and queries, respectively. Given that 3D PE is in the local space whereas the output queries are defined in the global coordinate system, we adopt the bilateral attention mechanism to avoid the mixture of embeddings in different representation spaces, as shown in Figure 1(b).
We further extend CAPE to integrate multi-frame temporal information to boost the 3D object detection performance, named CAPE-T. Different from previous methods that either warp the explicit BEV features using ego-motion or encode the ego-motion into the position embedding , we adopt separated sets of object queries for each frame and encode the ego-motion to fuse the queries.
We summarize our key contributions as follows:
We propose a novel multi-view 3D detection method, called CAPE, based on camera-view position embedding, which eliminates the variances of view transformation caused by different camera extrinsics.
We further generalize our CAPE to temporal modeling, by exploiting the object queries of previous frames and leveraging the ego-motion explicitly for boosting 3D object detection and velocity estimation.
Extensive experiments on the nuScenes dataset show the effectiveness of our proposed approach and we achieve the state-of-the-art among all LiDAR-free methods on the challenging nuScenes benchmark.
Related Work
DETR is the pioneering work that successfully adopts transformers in the object detection task. It adopts a set of learnable queries to perform cross-attention and treat the matching process as a set prediction case. Many follow-up methods focus on addressing the slow convergence problem in the training phase. For example, Conditional DETR decouples the items in attention into spatial and content items, which eliminates the noises in cross attention and leads to fast convergence.
2 Monocular 3D Detection
Monocular 3D Detection task is highly related to multi-view 3D object detection since they both require restoring the depth information from images. The methods can be roughly grouped into two categories: pure image-based methods and depth-guided methods. Pure image-based methods mainly learn depth information from objects’ apparent size and geometry constraints provided by eight keypoints or pin-hole model . Depth-guided methods need extra data sources such as point clouds and depth images in the training phase [36, 49, ma2020rethinking, 35, 7]. Pseudo-LiDAR converts pixels to pseudo point clouds and then feeds them into a LiDAR-based detector . DD3D claims that pre-training paradigms could replace the pseudo-lidar paradigm. The quality of depth estimation would have a large influence on those methods.
3 Multi-View 3D Detection
Multi-view 3D detection aims to predict 3D bounding boxes in the global system from multi-cameras. Previous methods mostly extend from monocular 3D object detection. Those methods cannot leverage the geometric information in multi-view images. Recently, several methods attempt to percept objects in the global system using explicit bird’s-eye view (BEV) maps. LSS conducts view transform via predicting depth distribution and lift images onto BEV. BEVFormer exploits spatial and temporal information through predefined grid-shaped BEV queries. BEVDepth leverages point clouds as the depth supervision and encodes camera parameters into the depth sub-network.
Some methods learn implicit BEV features following DETR paradigm. These methods mainly initialize 3D sparse object queries and interact with 2D features by attention to directly perform 3D object detection. For example, DETR3D samples 2D features from the projected 3D reference points and then conducts local cross attention to update the queries. PETR proposes 3D position embedding in the global system and then conducts global cross attention to update the queries. PETRv2 extends PETR with temporal modeling and incorporate ego-motion in the position embedding. CAPE conducts the attention in image space and local 3D space to eliminate the variances in view transformation. CAPE could preserve the pros in single-view approaches and leverage the geometric information provided by multi-view images.
4 View Transformation
The view transformation from the global view to local view in 3D scenes is an effective approach to boost performances for detection tasks. This could be treated as a normalization by aligning all the views, which could facilitate the learning procedure greatly. Several LiDAR-based 3D detectors estimate local coordinates rather than global coordinates for instances in the second stage, which could fully extract ROI features. For example, PointRCNN proposes the canonical 3D box refinement in the canonical system for more precise regression. To reduce the data variability for point cloud, AziNorm proposes a general normalization in the data pre-process stage. Different from these methods, our method conduct view transformation for eliminating the extrinsic variances brought by multi cameras with camera-view position embedding.
Our Approach
We present a camera-view position embedding (CAPE) approach for multi-view 3D detection and construct the position embeddings in each camera coordinate system.
Architecture. We adopt the multi-view DETR framework, a multi-view extension of DETR, depicted in Figure 3. The input multi-view images, , are processed with the encoders to extract the image embeddings,
The decoder is similar to DETR decoder, with a stack of decoder layers that is composed of self-attention, cross-attention , and feed-forward network (FFN). The -th decoder layer is formulated as follows,
Here, and are the output decoder embedding of the th and the th layer, respectively. are D reference points following the design in .
Self-attention is the same as the normal DETR, and takes the sum of and the position embedding of the reference points as input. Our work lies in learning 3D position embeddings. In contrast to PETR and PETRv2 that form the position embedding in the global coordinate system, we focus on learning position embeddings in the camera coordinate system for cross-attention.
Key Position Embedding Construction. We take one camera view (one image) as an example and describe how to construct the position embeddings for one view. Each D position in the image plane corresponds to D coordinates along the predefined depth bins in the frustum: . The D coordinate in the image frustum is transformed to the camera coordinate system,
where is the intrinsic matrix for -th camera. The transformed D coordinates are mapped into the single embedding,
Here, is a -dimensional vector, with . is instantiated by a multi-layer perceptron (MLP) of two layers.
Query Position Embedding Construction. We use a set of learnable D reference points in the global space to form the object queries. We transform the D points into the camera coordinate system for each view,
where is the extrinsic parameter matrix for the -th camera denoting the coordinate transformation from the global (LiDAR) system to the camera-view system. The transformed D coordinates are then mapped into the query position embedding,
The decoder embedding is updated by aggregating information from all views:
where is the soft-max, note that projection layers are omitted for simplicity. More details can be seen in the appendix. Feature-guided Key and Query Position Embeddings. Similar to PETRv2 , we make use of the image features to guide the key position embedding computation by learning the scaling weights and update Eq.4 as:
where is a two-layers MLP, denotes the element-wise multiplication and is the image features at the corresponding position. It is assumed to provide some informative guidance (e.g., depth).
On the query side, inspired by conditional DETR , we use the decoder embedding to guide the query position embedding computation and update Eq.6 as:
The extrinsic parameter matrix is used to transform the spatial decoder embedding to camera coordinate system, for alignment with the reference point position embeddings. Specifically, is instantiated as:
Here and are two-layers MLP.
Temporal modeling with ego-motion embedding. We utilize the previous frame information to boost the detection for the current frame. The reference points of the previous frame are transformed from the reference points of the current frame using the ego-motion matrix ,
Considering the same moving objects may have different 3D positions in two frames, we build the separated decoder embeddings that can represent different position information for each frame. The interaction between the decoder embeddings of two frames for -th decoder layer is formulated as follows:
We elaborate on the interaction between queries in two frames in Figure 5. Given that and are not in the same ego coordinate system, we inject the ego-motion information into the decoder embedding for spatial alignment. Then we update decoder embeddings with channel attention weights generated from the concatenation of the decoder embeddings of two frames. Compared with using one set of queries learning objects in different frames, queries in our method have a stronger capability in positional learning.
Heads and Losses. The detection heads consist of the classification branch that predicts the probability of object classes and the regression branch that regresses 3D bounding boxes. The regression branch predicts the relative offsets w.r.t. the coordinates of 3D reference points in the global system. As for the loss function, we adopt focal loss for classification and L1 loss for regression following prior works . The label assignment strategy here is the Hungarian algorithm . Suppose that is the assignment function, the loss for 3D object detection for the model without temporal modeling can be summarized as:
where and denote the set of ground truths and predictions respectively. is a hyper-parameter to balance losses. As for the network with temporal modeling, different from other methods supervise predictions only on the current frame, we predict results and supervise them on previous frames as auxiliary losses to enhance the temporal consistency. We use the center location and velocity of the ground truth on the current frame to generate the ground truths on the previous frame.
and denote the losses for the current and previous frame separately. is a hyper-parameter to balance losses.
Experiments
We evaluate CAPE on the large-scale nuScenes dataset. This dataset is composed of 1000 scene videos, with 700/150/150 scenes for training, validation, and testing set, respectively. Each sample consists of RGB images from 6 cameras and has 360 ° horizontal FOV. There are 20s video frames for each scene and 3D annotations are provided with every 0.5s. We report nuScenes Detection Score (NDS), mean Average Precision (mAP), and five True Positive (TP) metrics: mean Average Translation Error (mATE), mean Average Scale Error (mASE), mean Average Orientation Error (mAOE), mean Average Velocity Error (mAVE), mean Average Attribute Error (mAAE).
2 Implementation Details
We follow the PETR to report the results. We stack six transformer layers and adopt eight heads in the multi-head attention. Following other methods, CAPE is trained with the pre-trained model FCOS3D on validation dataset and with DD3D pre-trained model on the test dataset as initialization. We use regular cropping, resizing, and flipping as data augmentations. The total batch size is eight, with one sample per GPU. We set = 0.1 to balance the loss weight between the current frame and the previous frame and set = 2.0 to balance the loss weight between classification and regression. For validation dataset setting, we train CAPE for 24 epochs on 8 A100 GPUs with a starting learning rate of 2 that decayed with cosine annealing policy. For test dataset setting, we adopt denoise for faster convergence. We train 24 epochs with CBGS on the single-frame setting. Then we load the single-frame model of CAPE as the pre-trained model for multi-frame training of CAPE-T and train 60 epochs without CBGS.
3 Comparison with State-of-the-art
We show the performance comparison in the nuScenes test set in Tab. 1. We first compare the CAPE with state-of-the-art methods on the single-frame setting and then compare CAPE-T(the temporal version of CAPE) with methods that leverage temporal information. As for the model on the single-frame setting, as far as we know, CAPE(NDS=52.0) could achieve the first place on nuScenes benchmark compared with vision-based methods with the single-frame setting. As for the model with temporal modeling, CAPE-T still outperforms all listed methods. CAPE achieves on NDS and on mAP. We point out that using LiDAR as supervision could largely improve the mATE metric, thus it’s not fair to compare methods (w/wo LiDAR supervision) together. Nevertheless, even compared with contemporary methods leveraging LiDAR as supervision, CAPE outperforms BEVDepth on NDS and on mAP and achieves comparable results to BEVStereo .
We further show the performance comparison on the nuScenes validation set in Tab. 2 and Tab. 3. It could be seen that CAPE surpasses our baseline to a large margin and performs well compared with other methods.
4 Ablation Studies
In this section, we validate the effectiveness of our designed components in CAPE. We use the resolution for all single-frame experiments and the resolution for all multi-frames experiments.
Effectiveness of camera view position embedding. We validate the effectiveness of our proposed camera view position embedding in Tab 4. In Setting(a), we simply adopt PETR with feature-guided position embedding as our baseline. When we adopt the bilateral attention mechanism with 3D position embedding in the LiDAR system in Setting(b), the performance can be improved by on NDS. When using camera 3D position embedding without bilateral attention mechanism in Setting(c), the model could not converge and only get NDS. It indicates that the 3D position embedding in the camera system should be decoupled with output queries in the LiDAR system. The best performances could be achieved when we use camera view position embeddings along with the bilateral attention mechanism. For fair comparison with 3D position embedding in the LiDAR system, camera view position embedding improves in NDS and in mAP (see Setting(b) and (d)).
Effectiveness of feature-guided position embedding. Tab. 5 shows the effect of feature-guided position embedding in both queries and keys. Different from the FPE in PETRv2, our K-FPE here is formed under the local camera-view coordinate system instead of the global coordinate system. Compared with Setting(a) and (b), we find the Q-FPE increases the location accuracy (see the improvement of mAP and mATE), but decreases the orientation performance in mAOE, mainly owing to the lack of image appearance information in the local view attention. Compared with Setting(a) and (c), using K-FPE could improve on NDS and on mAOE, which benifits from more precise depth and orientation information in image appearance features. Compared with Setting(c) and (d), adding Q-FPE could bring and gain on NDS and mAP further. The benefit brought by Q-FPE could be explained by 3D anchor points being refined by input queries in high-dimensional embedding space. It could see that using both Q-FPE and K-FPE achieves the best performances.
Effectiveness of the temporal modeling approach. We show the necessity of using a set of queries for each frame in Tab. 6. It could be observed that decomposing queries into different frames could improve on NDS and on mAP. With the multi-group design, one object query could correspond to one instance on each frame respectively, which is proper for DETR-based paradigms. Similar conclusion is also observed in 2D instance segmentation tasks . Since we use multi groups of queries, auxiliary supervision on previous frames can be added to better align object queries between frames. Meanwhile, the generated ground truth on the previous frame could be treated as a type of data augmentation to avoid overfitting. When we adopt the previous loss, the mAP increases , which proves the validity of supervision on multi-frames. Compared with Setting(a) and (c), our temporal modeling approach could improve on NDS and on mAP on the validation dataset.
Effectiveness of different fusion approaches. The fusion module is used to fuse different frames for temporal modeling. We explore some common fusion approaches for temporal fusion in Tab. 7. We first try a simple fusion approach “concat with MLP”, and achieve on NDS, which has improvement compared with sharing queries. Considering queries on each frame have similar semantic information and different positional information, we propose the fusion model inspired by the channel attention. As is seen in Tab. 7, our proposed fusion approach “channel att” achieves higher performance compared to simple concatenation operation. We claim that the performance gain is not from the increased parameters since only three fully-connected layers are added in our model. Since queries are defined in each frame’s system and the ego motion occurs between frames, we encode the ego-motion matrix as a high dimensional embedding to align queries in the current frame’s system. With ego-motion embeddings, our fusion approach could further improve on NDS and on mAP.
5 Visualization
We show the visualization of attention maps in Figure 6 from 4 heads out of 8 heads. We display attention maps after the soft-max normalized operation. From the up-bottom way in each row, there are local view attention maps, global view attention maps, and overall attention maps separately. We obverse and draw three conclusions from visualization results. Firstly, local view attention mainly tends to highlight the neighbor of objects, such as the front, middle, and bottom of the objects, while global view attention pays more attention to the whole of images, especially on the ground plane and the same type of objects, as shown in Figure 6 (a). This phenomenon indicates that local view attention and global view attention are complementary to each other. Secondly, as shown in Figure 6 (a) and (c), it can be seen that overall attention maps are highly similar to local view attention maps, which means the local view attention maps play the dominant role compared to global view attention. Thirdly, we could obverse that the overall attention maps are further concentrating on the foreground objects in a fine-grained way, which implies superior localization accuracy.
6 Robustness Analysis
We evaluate the robustness of our method on camera extrinsic interference in this section. Camera extrinsic interference is an unavoidable dilemma caused by calibration errors, vehicle jittering, etc. We imitate extrinsics noises on the rotation with different noisy levels following PETRv2 for a fair comparison. Specifically, we randomly sample an angle within the specific range and then multiply the generated noisy rotation matrix by the camera extrinsics in the inference. We present the performance drop of metric mAP on both PETRv2 and CAPE-T in Fig.7. We could see that our method has more robust performances on all noisy levels compared to PETRv2 when facing extrinsics interference. For example, in the noisy level setting , CAPE-T drops 1.31% while PETRv2 drops 2.39%, which shows the superiority of camera-view position embeddings.
Conclusion
In this paper, we study the 3D positional embeddings of sparse query-based approaches for multi-view 3D object detection and propose a simple yet effective method CAPE. We form the 3D position embedding under the local camera-view system rather than the global coordinate system, which largely reduces the difficulty of the view transformation learning. Furthermore, we extend our CAPE to temporal modeling by exploiting the fusion between separated queries for temporal frames. It achieves state-of-the-art performance even without LiDAR supervision, and provides a new insight of position embedding in multi-view 3D object detection.
Limitation and future work. The computation and memory cost would be unaffordable when it involves the temporal fusion of long-term frames. In the future, we will dig deeper into more efficient spatial and temporal interaction of 2D and 3D features for autonomous driving systems.