Polar Parametrization for Vision-based Surround-View 3D Detection

Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Qian Zhang, Chang Huang, Wenyu Liu

Introduction

In the field of autopilot, surround-view camera system has been attached great importance and popularized by both industry and academia, given its low assembly cost and rich semantic information. And 3D detection based on such vision system has become the critical technique of environmental perception. The passed several years have witnessed tremendous progress in vision-based 3D detection . Most methods parameterize object’s position in two manners, 1) Image-based Parametrization ( Image-based Param. for short) and 2) Cartesian Parametrization (Cartesian Param. for short).

For Image-based Param. (Fig. 1 left), object’s position is defined in the pixel coordinate frame, parameterized by a 33-tuple (u,v,d)(u,v,d), where (u,v)(u,v) is the pixel coordinate and dd is the object’s depth relative to camera. With camera’s extrinsic and intrinsic parameters, we can transform (u,v,d)(u,v,d) to a 3D point to localize the object in 3D space. Image-based Param. is widely used in monocular methods . To process surround-view images, they independently take each view’s image as input to regress object’s (u,v,d)(u,v,d). Predicted objects of different views are then projected to the same 3D space for inter-camera merging. Post-processing among views (e.g., NMS) is adopted to filter out duplicated predictions. However, Image-based Param. raises some problems:

Estimating depth from single image is inherently an ill-posed inverse problem. The predicted depth is of large error and thus object’s positioning accuracy is poor. For the surround-view camera system, adjacent views overlap with each other. The correlation among views can be leveraged to promote perception performance, which is neglected in Image-based Param.

Post-processing among views is tricky and unstable. When the predictions from different views do not overlap in 3D space, NMS fails to filter out duplicated predictions.

Cartesian Parametrization.

For Cartesian Param. (Fig. 1 mid), object’s position is parameterized by 3D cartesian coordinate (x,y,z)(x,y,z). Correspondingly, the perception range is a rectangular region, denoted as {(x,y),∣x∣<Xmax⁡,∣y∣<Ymax⁡}\{(x,y),|x|<X_{\max},|y|<Y_{\max}\}. Cartesian Param. is adopted in some recent studies . Compared with monocular methods, they consider the correlation among views, i.e., taking all views together as input to jointly regress object’s 3D coordinate (x,y,z)(x,y,z).

But Cartesian Param. also raise problems. We take a special case for example to illustrate them. As shown in Fig. 2, assume that object AA appears in different views at timestamp t1t_{1} and t2t_{2}. They have the same pixel coordinate (u,v)(u,v) and radial distance dd. Their image patterns are also the same.

Cartesian Param. leads to ambiguity in label assignment because of its rectangular perception range. Ground-truth objects out of the perception range are ignored in label assignment. In the case of Fig. 2, though At1A_{t_{1}} and At2A_{t_{2}} are at the same distance dd and have the same image patterns, At1A_{t_{1}} is ignored while At2A_{t_{2}} is kept. At1A_{t_{1}} and At2A_{t_{2}} are not treated equally in the training phase, which is contradictory and harmful to the convergence of network.

Cartesian Param. neglects the view symmetry of surround-view cameras. From the perspective of machine learning, the detector aims at approximating a function FF that maps from image patterns XX to the prediction targets YY, i.e., Y=F(X)Y=F(X). In the case of Fig. 2, for Image-based Param., At1A_{t_{1}} and At2A_{t_{2}} have the same image patterns XX and prediction targets (u,v,d)(u,v,d), and can be treated as the same mapping in FF. At1A_{t_{1}} and At2A_{t_{2}} are symmetrical in terms of view. Image-based Param. leverages the view symmetry as inductive bias of the detector, making it easier to approximate FF. While for Cartesian Param., object At1A_{t_{1}} and At2A_{t_{2}} correspond to different prediction targets ((xt1,yt1,zt1)(x_{t_{1}},y_{t_{1}},z_{t_{1}}) for At1A_{t_{1}} and (xt2,yt2,zt2)(x_{t_{2}},y_{t_{2}},z_{t_{2}}) for At2A_{t_{2}}). Since the view symmetry is neglected, the mappings for At1A_{t_{1}} and At2A_{t_{2}} are learned separately. Thus, FF is more complicated and optimization becomes harder.

In this work, we propose a surround-view 3D DEtection TRansformer, named PolarDETR. To adapt to the view symmetry of surround-view camera system, we parameterize object’s position by polar (cylindrical)’Cylindrical’ is more proper for describing 3D space. But considering ’polar’ better reflects the insight of this work and Z dimension is not critical, we adopt ’polar’ for method description coordinate (r,α,z)(r,\alpha,z) and decomposes velocity to radial velocity vradv_{rad} and tangential velocity vtanv_{tan}. Besides, we reformulate perception range, label assignment and loss function in polar coordinate system. We term it as Polar Parametrization (Polar Param. for short).

Polar Param. establishes explicit associations between image patterns and prediction targets. Radial distance rr is associated with object’s size in image. Azimuth α\alpha is associated with the index of pixel. We add 3D positional encoding to each pixel which contains explicit clues about azimuth. Radial velocity vradv_{rad} is associated with the changing rate of object’s size. Tangential velocity vtanv_{tan} is associated with the object’s movement in image plane (similar to optical flow). With these explicit associations, the mapping is simplified and the detector achieves better convergence and performance.

Besides, PolarDETR enables center-context feature aggregation to enhance the information interaction between object queries and images, and adopts pixel ray as positional encoding to provide 3D spatial priors and help to predict azimuth α\alpha.

We evaluate PolarDETR on the challenging nuScenes benchmark. PolarDETR achieves promising performance-speed trade-off on different backbone configurations. And we submit PolarDETR for official evaluation on the nuScenes test set. PolarDETR ranks 1st on the highly competitive 3D detection and tracking leaderboard at the submission time (Mar. 4th, 2022). Besides, thorough ablation studies are provided to validate the effectiveness of PolarDETR.

Related Work

Polar coordinate is leveraged in some previous works for data representation. In the field of point cloud semantic segmentation, some grid-based methods partition 3D space into polar grids and transform LiDAR point clouds into the grid representation, in order to ease the problem of long-tailed distribution. In the field of image instance segmentation, PolarMask formulates the instance mask as a set of contour points represented in the polar coordinate system centered at the instance. Differently, in order to adapt to the view symmetry of surround-view cameras, PolarDETR parameterizes object’s position in polar coordinate system and reformulates label assignment and loss function accordingly.

2 DETR-based Object Detection

Recently, DETR formulates object detection as a set prediction problem and exploits a standard transformer , which contains a simple encoder-decoder structure without hand-craft designs and achieves promising performance. Based on Hungarian algorithm , the set of object queries can be matched with targets in a one-to-one manner, which helps remove the non-maximum suppression (NMS). Deformable DETR motivated by adopts multi-scale deformable attention in the transformer and achieves faster convergence and better performance than DETR. also address the convergence problem in DETR by incorporating spatial information into object queries. Several works propose to prune the redundant tokens in transformers to minimize the computation cost of DETR. In terms of 3D object detection, DETR3D extends DETR for 3D domain. Inspired by Deformable DETR, DETR3D projects 3D object queries to 2D reference points to aggregate features from all views. DETR3D builds up a simple DETR-based pipeline for 3D object detection, while it suffers from (1) ambiguity in label assignment, (2) neglecting view symmetry and (3) insufficient contextual information. To solve these problems, PolarDETR adopts Polar Parametrization to ease optimization and enables center-context feature aggregation to enhance the feature interaction.

3 Vision-based 3D Object Detection

Vision-based 3D detection is a basic perception task in autonomous driving. Early studies are mainly based on KITTI dataset, which provides front-view object annotations. KITTI boosts the development of monocular 3D object detection methods . Recently, with nuScenes dataset available, which contains 360∘360^{\circ} annotations around the ego vehicle, new 3D detection paradigms have been proposed. Some works still follows the monocular pipeline to detect objects, and then project the multi-view detection results to the same coordinate system and adopt NMS to merge results. DETR3D extends DETR for 3D object detection. BEVDet projects surround-view image features to Bird-Eye-View (BEV) space, and set detection head on BEV features. We presents a new paradigm specially designed for surround-view camera systems, in which the view symmetry is exploited to ease optimization and boost performance.

PolarDETR

Fig. 3 illustrates the proposed PolarDETR, which follows the DETR paradigm. Given surround-view images I={I1,…,IK}\mathcal{I}=\{{I}_{1},\dots,{I}_{K}\} (K denotes the number of views), a shared CNN backbone extracts image features Fimg={F1,…,FK}\mathcal{F}_{\text{img}}=\{{F}_{1},\dots,{F}_{K}\}. A set of object queries Q={q1,…,qN}\mathcal{Q}=\{{q}_{1},\dots,{q}_{N}\} (N denotes the number of queries) is used to detect objects. Specifically, each object query encodes the semantic features and positional information of the corresponding object. And a series of decoder layers aggregate features from surround-view feature maps and update queries iteratively. The feed forward network (FFN) follows the decoder layers and predicts polar box encodings BencB_{\text{enc}}, polar velocity components (vrad,vtan)(v_{rad},v_{tan}), and class labels based on queries.

2 Polar Parametrization

In PolarDETR, object’s position is parameterized by the polar coordinate. FFN outputs a 99-tuple polar box encoding, denoted as,

We then decode BencB_{\text{enc}} to get the predicted box vector BpredB_{\text{pred}} represented in polar coordinate system, i.e.,

where rr, α\alpha, zz are the radial distance, azimuth and height, indicating the position of object’s geometry center in 3D space. (l,w,h)(l,w,h) and θ\theta denote the size and orientation of bounding box, respectively. Zmax⁡Z_{\max} and Zmin⁡Z_{\min} denote the perception range along the ZZ dimension. Rmax⁡R_{\max} denotes the maximum perception distance. σ\sigma is sigmoid function. To make sure the continuity of the regression space , we parameterize both α\alpha and θ\theta by a 2-D (sin⁡(⋅),cos⁡(⋅))(\sin(\cdot),\cos(\cdot)) pair.

For Polar Param., object’s position is decoupled into the radial distance and the azimuth. The radial distance is symmetrical in different views. It’s highly correlated with the object size in image, and can be learned from image patterns. The azimuth is relative to the indices of cameras and pixels which capture the object, and can be learned from the positional encodings. Compared with Cartesian Param., which localizes objects by regressing the cartesian coordinate (x,y)(x,y), Polar Param. makes more sense.

Most methods decompose velocity along the cartesian axes and get velocity components vxv_{x} and vyv_{y}. Differently, PolarDETR decomposes velocity to radial velocity vradv_{rad} and tangential velocity vtanv_{tan}. Radial velocity vradv_{rad} is associated with the changing rate of object’s size. Tangential velocity vtanv_{tan} is associated with the object’s movement in image plane (similar to optical flow). PolarDETR establishes explicit associations between image patterns (input) and velocity (prediction target), resulting in more accurate velocity estimation.

3 Decoder Layer

Decoder layers iteratively aggregate features and update queries. Each decoder layer begins with a multi-head self-attention module (MHSA) for inter-query information interaction. Then a linear layer extracts object’s 3D position from each query, i.e.,

We decode (br,bsin⁡α,bcos⁡α,bz)(b_{r},b_{\sin\alpha},b_{\cos\alpha},b_{z}) with Eq. (2) and get object’s 3D center point ci3D=(r,α,z)c_{i}^{\text{3D}}=(r,\alpha,z).

Then we aggregate features from surround-view feature maps through a two-step center-context procedure.

Firstly, we project the 3D center point ci3Dc_{i}^{\text{3D}} to each image and get a set of 2D center points {ci1,...,ciK}\{c_{i}^{1},...,c_{i}^{K}\}, which denote object’s centers in all views. I.e.,

where Kk\mathbf{K}^{k} and Rtk\mathbf{Rt}^{k} are the projection matrices of view kk derived from camera’s intrinsics and extrinsics respectively. Based on {ci1,...,ciK}\{c_{i}^{1},...,c_{i}^{K}\}, we sample center features {fci1,...,fciK}\{f_{c_{i}^{1}},...,f_{c_{i}^{K}}\} from surround-view feature maps {F1,…,FK}\{{F}_{1},\dots,{F}_{K}\} through bilinear interpolation. A 3D point may be invisible in some views and the projected 2D points are out of range of the images. In this case, the corresponding point features are set to zero.

Considering that center features are not informative enough for localizing objects, we further include context features to enhance the interaction between queries and surround-view features. Specifically, based on both the center features {fci1,...,fciK}\{f_{c_{i}^{1}},...,f_{c_{i}^{K}}\} and the query embeddings qiq_{i}, we generate a set of context points {pi1,...,piK}\{p_{i}^{1},...,p_{i}^{K}\} by predicting the offset relative to the center points. I.e.,

Pixel Ray.

Motivated by , we introduce pixel ray as 3D spatial priors. As shown in Fig. 4, pixel ray travels from camera’s optical center through the pixel to the corresponding 3D point. It encodes the correspondence between 2D image pixel and 3D point, and contains explicit clues about azimuth. We leverage pixel ray as extra positional encodings. Specifically, for each center or context point, the unit direction vector of corresponding pixel ray drayd_{\text{ray}} is concatenated with point features in the channel dimension (refer to Eq. (6)).

Query Update.

Then we aggregate center and context features with drayd_{\text{ray}} to update query embeddings, i.e.,

The updated query embeddings encode more accurate positional information of the object and contribute to better feature aggregation in the next decoder layer.

4 Perception Range, Label Assignment and Loss Function

Polar Param. reformulates perception range, label assignment and loss function in polar coordinate system.

As discussed above, for Cartesian Param. the rectangular perception range introduces ambiguity into label assignment. For Polar Param., the concerned perception range is a circular region with radius Rmax⁡R_{\max} centered at the ego vehicle. It avoids ambiguity and fits with the common sense. It’s worth noting that the evaluation region of nuScenes benchmark corresponds with the circular perception range of Polar Param.

Label Assignment.

Label assignment between predictions and ground-truth (GT) objects is also performed based on polar coordinate. 3D box annotations are first transformed to polar coordinate representation, i.e.,

Then we adopt bipartite matching to uniquely assign predictions with ground-truth boxes. Assume there exist NN predictions and MM GTs. The pair-wise matching cost between prediction ii and GT jj is defined as,

Loss Function.

We reformulate the loss function based on polar coordinate. The bipartite matching loss consists of two parts: a focal loss for class labels and a L1\mathcal{L}_{1} loss for polar box parameters (r,sin⁡α,cos⁡α,z,l,w,h,sin⁡θ,cos⁡θ)(r,\sin\alpha,\cos\alpha,z,l,w,h,\sin\theta,\cos\theta) and polar velocity components (vrad,vtan)(v_{rad},v_{tan}). The terms about azimuth, i.e., sin⁡α\sin\alpha and cos⁡α\cos\alpha, are also scaled up by kscalingk_{\text{scaling}} to balance the error distribution between tangential and radial direction.

5 Temporal Information

Temporal information is important for velocity estimation and occlusion cases. We straightforwardly and elegantly extend PolarDETR to PolarDETR-T, which takes streaming camera frames as input. The 3D center point ci3Dc_{i}^{\text{3D}} are projected to past frames for fetching image features. Taking frame t−nt-n for example,

where cik (t−n)c_{i}^{k\ (t-n)} denotes the corresponding 2D point in the frame t−nt-n, and Pose(t−n)\mathbf{Pose}^{(t-n)} denotes the pose transformation matrix which reflects the movement of the ego-vehicle in the time interval [t−n,t][t-n,t]. We sample center and context features from past frames in the same manner of current frame as mentioned above. And all the sampled features (both current and past ones) are aggregated together to update query embeddings.

For a running system with streaming input, for efficient inference, we can cache image feature maps of the past frames. For each moment, we only need to process images of current frame tt with backbone to get Fimg(t)\mathcal{F}_{\text{img}}^{(t)}. Feature maps of past frames {Fimg(t−1),Fimg(t−2),...}\{\mathcal{F}_{\text{img}}^{(t-1)},\mathcal{F}_{\text{img}}^{(t-2)},...\} are directly fetched from cache, avoiding duplicated computation. Since most computation cost lies in backbone, PolarDETR-T can run at a similar FPS compared with PolarDETR.

Experiments

We validate the effectiveness of PolarDETR on the large-scale nuScenes dataset, which is currently the most popular benchmark for vision-based methods. NuScenes contains 10001000 driving sequences, with 700700, 150150 and 150150 sequences for training, validation and testing, respectively. Each sequence is approximately 20-second long and provides 66 surround-view images per frame with the resolution of 1600×9001600\times 900. We submit test set results to the online server for official evaluation to get the leaderboard results. And other experiments are evaluated on the val set.

2 Experimental Settings

We implement PolarDETR with PyTorch framework and MMDetection3D toolbox. The results are based on three backbone configurations. We adopt ImageNet pretraining for ResNet-50, FCOS3D pretraining for ResNet-101, and DD3D pretraining for VoVNet. Unless specified, we adopt ResNet-50 as backbone for ablation experiments. We train PolarDETR on eight RTX3090 GPUs with the total batch size 88. Inference speeds of all models are measured on one RTX3090 GPU. We adopt mixed precision (float32 and float16) to accelerate training and disable it when measuring the inference speeds. We do not use test-time augmentation or model ensemble during inference. We adopt the implementation of for the context point sampling. The number of attention heads is set to 88 as the common practice. Results of PolarDETR-T are based on only one past frame. Using more past frames are feasible and brings more gain. More details about experimental settings will be available in the code.

In addition, we simply extend PolarDETR for 3D object tracking by leveraging the tracking-by-detection algorithm . Specifically, we project objects of the current frame back to the previous frame with the predicted velocity, and then match them with the tracked objects by closest distance matching. We evaluate the tracking performance on nuScenes tracking benchmark.

3 Metrics

For 3D object detection, we report the official metrics: NuScenes Detection Score (NDS), mean Average Precision (mAP) and true positive (TP) metrics (including Average Translation Error (ATE), Average Scale Error (ASE), Average Orientation Error (AOE), Average Velocity Error (AVE) and Average Attribute Error (AAE)). The main metric, NDS, is a weighted sum of the other metrics for comprehensively judging the detection capacity, defined as,

Tracking.

For 3D object tracking, we report the official metrics: Average Multi Object Tracking Accuracy (AMOTA), Average Multi Object Tracking Precision (AMOTP), False Positives (FP), False Negatives (FN), Identity Switches (IDS), Track Initialization Duration (TID) and Longest Gap Duration (LGD). AMOTA serves as the main metric, which penalizes ID switches, false positive, and false negatives and is averaged among various recall thresholds.

4 Main Results

In Tab. 1, we compare PolarDETR with other state-of-the-art methods on three backbone configurations. On both ResNet-50 and ResNet-101, PolarDETR significantly outperforms DETR3D and BEVDet while achieving comparable inference speed. And on ResNet-101, PolarDETR outperforms FCOS3D and PGD in terms of both performance and speed. On VoVNet, PolarDETR significantly outperforms DETR3D . With temporal information, PolarDETR-T achieves much higher results than PolarDETR, especially in terms of mAVE.

NuScenes Leaderboard.

Tab. 2 shows the 3D detection leaderboard of nuScenes benchmark for camera modality at the submission time (Mar. 4th, 2022). For fair comparison, we adopt the pretrained VoVNet (refer to DD3D ) as backbone, the same with the other top methods on the leaderboard. PolarDETR ranks 1st on this highly competitive leaderboard. And it’s worth noting that our implementation is compact and elegant, without test-time augmentation, multi-model ensemble or other tricks.

Tab. 3 shows the tracking leaderboard of nuScenes benchmark for camera modality (Mar. 4th, 2022). At the submission time, PolarDETR ranks 1st on the leaderboard and outperforms other methods by a large margin. We only adopt a simple algorithm to generate tracking results. The tracking performance could be further improved with well-designed tracking algorithm.

5 Ablation Study

We provide ablation experiments to validate the key components of PolarDETR. As shown in Tab. 4, Polar Param., context point and pixel ray respectively improve NDS by 1.5%1.5\%, 0.9%0.9\% and 0.6%0.6\% and bring negligible overhead. It’s worth noting that Polar Param. only reformulates the optimization problem without introducing any computational budget. The significant improvement brought by Polar Param. validates its effectiveness.

Polar Velocity Decomposition.

Tab. 5 presents the ablation study about the velocity decomposition. We compare cartesian and polar decomposition based on temporal input (frame t−1t-1 and tt). Polar decomposition results in much lower mAVE (mean Average Velocity Error) and achieves much accurate velocity estimation.

Scaling Factor.

Tab. 6 shows the ablations about the scaling factor kscalingk_{\text{scaling}}. Without numerically scaling up azimuth, i.e., kscaling=1k_{\text{scaling}}=1, the performance is poor because of the numerical unbalance between the tangential and radial direction. We use coarse grid search to tune kscalingk_{\text{scaling}}, and find kscaling=20k_{\text{scaling}}=20 corresponds to relatively good results. With finer tuning, higher performance can be expected.

Context Point.

Tab. 7 shows the ablations about the number of context points. Without context points, NDS is relatively low (0.390)0.390) because of lack of contextual information. Performance improves with the number increasing. But the gain get saturated with 44 context points (NDS=0.403\text{NDS}=0.403). And further increasing the number has negative effects. By default 44 context points are adopted in PolarDETR.

In Tab. 8, we further explore the impact of input for generating context points. As mentioned in Fig. 3 and Eq. (5), context points are predicted based on a combination of query embeddings qiq_{i} and center features {fci1,...,fciK}\{f_{c_{i}^{1}},...,f_{c_{i}^{K}}\}. The results prove that both terms contribute to better context feature aggregation and higher performance.

Decoder Layer.

Ablations about the number of decoder layers are shown in Tab. 9. With only 11 decoder layer (no iteration), results are poor. In PolarDETR, we adopt 66 decoder layers, which corresponds to relatively good results.

Qualitative Results

Visualizations about center and context points are in Fig. 6. Center points focus on object’s center for localization while context points focus on a wider region to capture more information. They complement each other and contribute to better feature aggregation.

Detection Results.

Qualitative results about final predictions are shown in Fig. 5. Blue and green boxes respectively denote predictions and GTs. Highly occluded or far-away objects may result in bad cases, which is a common problem faced by all detectors. Expect these challenging cases, PolarDETR achieves stable and satisfactory detection results.

Conclusion

In this paper, we present Polar Param. to exploit the view symmetry of surround-view camera system. Polar Param. establishes explicit associations between image patterns and prediction targets, superior to Image-based Param. and Cartesian Param.. Based on Polar Param., PolarDETR achieves promising performance in terms of both 3D detection and 3D tracking on the challenging nuScenes benchmark. And Polar Param. can be extended to other perception task and even planning task. We leave it as further work.

References