Sparse4D: Multi-view 3D Object Detection with Sparse Spatial-Temporal Fusion

Xuewu Lin, Tianwei Lin, Zixiang Pei, Lichao Huang, Zhizhong Su

Introduction

Multi-view visual 3D perception plays a critical role in autonomous driving systems, especially for low-cost deployment. Compared with Lidar modality, cameras can provide valuable visual cues for long-range distance detection and vision-only element identification. However, without explicit depth cues, 3D perception from 2D images is an ill-posed issue, leading to a long-standing challenge of how to properly fuse multi-camera images to address 3D perception tasks such as 3D detection. There are two mainstream categories of recent methods: the BEV-based methods and the sparse-based methods.

BEV-based methods address 3D detection via converting multi-view image features into an unified BEV space, and achieve excellent performance promotion. However, besides the advantages of the BEV fashion, there still exist some unavoidable disadvantages as follows: (1) The image-to-BEV perspective transformation requires dense feature sampling or rearrangement, which is complex and computationally expensive for low-cost edge devices deployment; (2) The maximum perception range is limited by the size of BEV feature map, making it difficult to trade off among perception range, efficiency and precision; (3) The height dimension is compressed in BEV feature with losing of texture cues. Thus BEV features are incompetent for some perception tasks such as sign detection.

Different from BEV based methods, the sparse based algorithms do not require a dense perspective transformation module, but directly sample sparse feature for 3D anchors refinement, thus can alleviate above issues. Among them, the most representative sparse 3D detection method is DETR3D . However, its model capacity is limited since DETR3D only sample feature of a single 3D reference point for each anchor query. Recently, SRCN3D utilizes RoI-Align to sample multi-view feature, but is not efficient enough and cannot precisely align feature points from different views. Meanwhile, existing sparse 3D detection methods have not taken advantage of rich temporal context, and have a significant performance gap compared with state-of-the-art BEV based methods.

In this work, we devote our best effect to expand the limit of sparse based 3D detection. To address existing issues, we introduce a novel framework named Sparse4D, which utilizes multiple keypoints distributed in the region of 3D anchor box to sample feature. Compared with the single point manner and the RoI-Align manner , our sampling manner has two main advantages: (1) can efficiently extract rich and complete context inside each anchor box; (2) can be simply extend to temporal dimension as 4D keypoints, then can effectively align temporal information. With 4D keypoints, as illustrated in Fig. 1, Sparse4D first performs multi-timestamp, multi-view and multi-scale for each keypoint. These sampled features then go through a hierarchical fusion module to generate high-quality instance feature for 3D box refinement. Further, to alleviate the ill-posed issue of camera-based 3D detection and improve the perceptual performance, we explicitly add an instance-level depth reweight module, where the instance feature is reweighted by depth confidence sampled from predicted depth distribution. This module is trained in a sparse way without additional Lidar point cloud surpervision.

In summary, our work have four main contributions:

To the best of our knowledge, our proposed Sparse4D is the first sparse multi-view 3D detection algorithm with temporal context fusion, which can efficiently and effectively align spatial and temporal visual cues to achieve precise 3D detection.

We propose a deformable 4D aggregation module that can flexibly complete the sampling and fusion of multi-dimensional (point, timestamp, view and scale) features.

We introduce a depth reweight module to alleviate the ill-posed issue in image-based 3D perception system.

On the challenging benchmark - nuScenes dataset, Sparse4D outperforms all existing sparse based algorithms and most BEV-based algorithms on 3D detection task, and also performs well on tracking task.

Related Work

Early object detection methods used dense predictions as output, and then utilized non-maxima suppression (NMS) to process those dense predictions. DETR introduces a new detection paradigm that utilizes set-based loss and transformer to directly predict sparse detection results. DETR performs cross attention between object-query and global image context, leading to heavy computation cost and difficulty in convergence. Due to the use of global cross attention, DETR cannot be regarded as a pure sparse method. Deformable DETR then modifies DETR and proposes a local cross attention based on reference points, which accelerates the model convergence and reduced computational complexity. Sparse R-CNN proposes another sparse detection framework based on the idea of region proposal. The network structure is extremely simple and effective, showing the feasibility and superiority of sparse detection. As the extension of 2D detection, many 3D detection methods have recently paid more attention to these sparse paradigms, such as MoNoDETR , DETR3D , Sparse R-CNN3D , SimMOD , etc.

2 Monocular 3D Object Detection

The monocular 3D detection algorithm takes a single image as input and outputs the 3D bounding box of the objects. Since the image does not contain depth information, this problem is ill-posed, and is more challenging compared with 2D detection. FCOS3D and SMOKE is extended based on a single-stage 2D detection network, using a fully convolution network to directly regress the depth of each object. convert the 2D image into the 3D pseudo point cloud signal with monocular depth estimation results, and then use the LiDAR-based detection network to complete the 3D detection. OFT and CaDDN transform the dense 2D image feature into BEV space with the help of the view transformation module and then send the BEV feature to the detector to complete 3D object detection. The difference is that OFT uses the 3D to 2D inverse projection relationship to complete the feature space transformation, while CaDDN is based on the 2D to 3D projection, which is more like a pseudo-LiDAR method.

3 Multi-view 3D Object Detection

Dense algorithms are the main research direction of multi-view 3D detection, which use dense feature vectors for view transformation, feature fusion or box prediction. Currently, BEV-based methods are the main part of dense algorithms. BEVFormer adopts deformable attention to complete the BEV feature generation and dense spatial-temporal feature fusion. BEVDet uses lift-splat operation to achieve the view transformation. On the basis of BEVDet, BEVDepth adds explicit depth supervision, which significantly improves the accuracy of the detection. BEVStereo and SOLOFusion introduce temporal stereo technology into 3D detection, further improving the depth estimation effect. PETR utilizes 3D position encoding and global cross attention for feature fusion, but the global cross attention is computationally expensive. Like vanilla DETR , PETR cannot be regarded as a purely sparse method. DETR3D is a representative work of sparse methods, which performs feature sampling and fusion based on sparse reference points. Graph DETR3D follows DETR3D and introduces a graph network to achieve better spatial feature fusion, especially for multi-view overlapping regions.

Methodology

As shown in Fig. 2, Sparse4D conforms to an encoder-decoder structure. The image encoder is used to extract image features with shared weights, which contains a backbone (e.g., ResNet and VoVNet ) and a neck (e.g., FPN ). Given NN view input images at time tt, the image encoder extracts multi-view multi-scale feature maps as It={It,n,s∣1≤s≤S,1≤n≤N}I_{t}=\left\{I_{t,n,s}|1\leq s\leq S,1\leq n\leq N\right\}. To exploit temporal context, we extract image feature of recent TT frames as image feature queue I={It}t=tst0I=\left\{I_{t}\right\}_{t=t_{s}}^{t_{0}}, where ts=t0−(T−1)t_{s}=t_{0}-(T-1).

All 3D anchors are set in a unified 3D coordinate system(e.g. central LiDAR coordinate).

In each refinement module, we first adopt self-attention to realize the interaction between instances, with the embedding of anchor parameters added before and after. Then, we conduct deformable 4D aggregation (Sec. 3.2) to fuse multi-view, multi-scale, multi-timestamp and multi-keypoint features. Furthermore, we introduce a depth reweight module (Sec. 3.3) to alleviate the ill-posedness issue in image-based 3D detection. Finally, a regression head is used to refine the current anchor via predicting the offset between ground truth and the current anchor.

2 Deformable 4D Aggregation

The quality of instance features have a critical impact on the overall sparse perception system. To address this, as demonstrated on Fig. 3, we introduce the deformable 4D aggregation module to obtain high-quality instance features with sparse feature sampling and hierarchy feature fusion.

where Ryaw\textbf{R}_{yaw} denotes the rotation matrix of yawyaw.

Temporal features are crucial for 3D detection and can improve depth estimation accuracy. Therefore, after getting the 3D keypoints of the current frame, we extend them to 4D to prepare for temporal fusion. For a past timestamp tt, we first build a constant velocity model to shift each 3D keypoints in the 3D coordinate system of the current frame.

where dtd_{t} is the time interval between two adjacent frames. Then, we use the ego vehicle motion information to convert Pm,t′P_{m,t}^{\prime} to the coordinate system of the past tt frame.

where Rt0→t\textbf{R}_{t_{0}\rightarrow t} and Tt0→t\textbf{T}_{t_{0}\rightarrow t} represent the rotation matrix and translation of the ego vehicle from current frame t0t_{0} to frame tt, respectively. In this way, we can finally construct 4D keypoints as Pm={Pm,t}t=tst0P_{m}=\left\{P_{m,t}\right\}_{t=t_{s}}^{t_{0}}.

Sparse Sampling. Based on the above 4D keypoints PP and the image feature maps queue FF, sparse features with strong representation ability can be efficiently sampled. First, the 4D keypoints are projected onto the feature maps through the transformation matrix Tcam\textbf{T}^{\rm cam}.

Then, we conduct multi-scale feature sampling for each view and each timestamp via bilinear interpolation:

Hierarchy Fusion. To generate high quality instance feature, we fuse the above features vectors fmf_{m} in a hierarchical manner. As shown in Fig. 3(c), for each keypoint, we first aggregate features in different view and scale with predicted weights and then conduct temporal fusion with sequence linear layers. Finally, for each anchor instance, we fuse multi-point features to generate instance feature.

Specifically, given instance feature FmF_{m} with anchor box embedding added, we first predict group weighting coefficients through a linear layer Ψ\Psi as:

where GG is the number of groups to divide features by channels. With this, we can aggregate channels of different groups with different weights, which is similar to group convolution . We sum the weighted feature vectors for each group along the scale and view dimensions, and then concatenate the groups to obtain the new features fm,k,t′f^{{}^{\prime}}_{m,k,t}.

The multi-keypoint features fm,k′′f_{m,k}^{{}^{\prime\prime}} after temporal fusion will be summed to complete the final feature aggregation and get the updated instance feature as:

3 Depth Reweight Module

This 3D to 2D transformation (Eq. 5) has a certain ambiguity, that is, different 3D points may correspond to the same 2D coordinates. For different 3D anchors, the same features may be sampled (see Fig. 4), which increases the difficulty of neural network fitting. To alleviate this problem, we incorporate an explicit depth estimation module Ψdepth\Psi_{depth}, which consists of multiple MLPs with residual connections. For each aggregated feature Fm′F_{m}^{\prime}, we estimate a discrete depth distribution, and use the depth of center point of 3d anchor box to sample the corresponding confidence CmC_{m}, which will be used to reweight the instance feature.

In this way, for those instances whose 3D center points are far from the ground truth in the depth direction, even if the 2D image coordinates are very close to the ground truth, the corresponding depth confidence tends to zero. Thus the corresponding instance feature Fm′′F_{m}^{\prime\prime} is punished after reweighting also tend to 0. Incorporating an explicit depth estimation module can help the visual perception system to further improve the perception accuracy. Also, the depth estimation module can be designed and optimized as a separate part to facilitate model performance.

4 Training

We sample video clips with TT frames to train the detector end to end. The time interval between consecutive frames is randomly sampled in {dt,2dt}\left\{d_{t},2d_{t}\right\} (dt≈0.5d_{t}\approx 0.5). Following DETR3D , the Hungarian algorithm is used to match each ground truth with one predicted value. The loss includes three parts: classification loss, bounding box regression loss and depth estimation loss:

where λ1\lambda_{1}, λ2\lambda_{2} and λ3\lambda_{3} are weight terms to balance the gradient. We adopt focal loss for classification, L1L_{1} loss for bounding box regression, and binary cross entropy loss for depth estimation. In depth reweight module, we directly use the depth of the labeled bounding box center as the ground truth to supervise per-instance depth. Since we only estimate per-instance depth rather than dense depth, the training process gets rid of the dependence on LiDAR data.

Experiment

We evaluate our method on the nuScenes benchmark. The nuScenes dataset contains data for 1000 scenes, of which 700, 150, and 150 scenes are used for training, validation, and testing, respectively. Each scene is a 20 second video clip at 2 frames per second. Each frame has image data from 6 cameras, and enough annotations such as the category, 3D bounding box, and ID of objects.

For the 3D detection task, evaluation metrics include mean Average Precision (mAP), mean Average Error of Translation (mATE), Scale (mASE), Orientation (mAOE), Velocity (mAVE), Attribute (mAAE) and nuScenes Detection Score (NDS), where NDS is a weighted average of other metrics. For the object tracking task, Average Multi-Object Tracking Accuracy (AMOTA), Average Multi-Object Tracking Precision (AMOTP) and Recall are the three main evaluation metrics. Please refer to for details.

2 Implementation Details

The initial {x,y,z}\left\{x,y,z\right\} parameters of the 3D anchors are obtained by performing K-Means clustering on the training set, and the other parameters are all initialized with fixed values {1,1,1,0,1,0,0,0}\left\{1,1,1,0,1,0,0,0\right\}. The instance feature uses random initialization. By default, the number of 3D anchors and instance features MM is set to 900, the number of cascade refinement modules is 6, the number of feature map scales SS from the neck is 4, the number of fixed keypoints KFK_{F} is 7, the number of learnable keypoints KLK_{L} is 6, the input image size is 640×1600640\times 1600, and the backbone is ResNet101.

Sparse4D is trained with AdamW optimizer . The initial learning rates of backbone and other parameters are 2e-5 and 2e-4, respectively. The decay strategy is cosine annealing . The initial network parameters come from pre-trained FCOS3D . For the experiments on the nuScenes test set, the network was trained for 48 epochs, and the rest of the experiments were only trained for 24 epochs unless otherwise specified. In order to save GPU memory, we detach the feature maps of all historical frames and the fusion features ft′f_{t}^{\prime} of a random part of historical frames during the training phase. CBGS and test time augmentation were not used in all experiments.

3 Ablation Studies and Analysis

Depth Reweight Module and Learnable Keypoints. By adding the depth reweight module or learnable keypoints, we compare and analyze the before-and-after changes in metrics on the nuScenes validation dataset, see Tab. 2. It can be seen that the addition of these two modules has a certain promotion effect on the model performance, and the impact on the metric NDS is similar, 0.33%0.33\% and 0.35%0.35\%, respectively. When these two structures are added together, all metrics will increase, among which mAP increases by 0.38%0.38\% and NDS increases by 0.79%0.79\%.

Motion Compensation. When generating 4D keypoints, we consider both the ego vehicle motion and the object motion. From Tab. 2, we can see that even without any motion compensation, after adding temporal information, the model performance still improved to a certain extent, in which mAVE increased by 20.8%20.8\% and NDS increased by 2.3%2.3\%. However, the overall perceptual performance of this model is still low. After adding ego motion compensation, the model effect is significantly improved, especially mAP and mAVE, which are increased by 4.2%4.2\% and 28.4%28.4\% respectively, and the comprehensive metric NDS is increased by 6.4%6.4\%. On this basis, considering the motion of the object to be detected, the detection accuracy is not improved, but the error of the speed estimation will be reduced by 6.9%6.9\%, thus increasing the NDS by about 0.7%0.7\%.

Number of Refinements. The number of iterative refinements also has a significant impact on detection performance. In this regard, we designed two sets of experiments for analysis. In the first set of experiments, we train a model with 6 refinement modules and compute the detection metrics output by each refinement module. As can be seen from Fig. 5(a), as the times of refinements increases, the overall metrics show an increasing trend, and the growth rate gradually decreases. Compared with the first one, the output accuracy of the second refinement module is significantly increased, but the detection effect between the fifth refinement module and the sixth module is not much different. In the second set of experiments, we train multiple models whose number of refinement modules is increased from 2 to 14. When the number of modules is 10, the NDS is the highest at 38.1%38.1\%, as shown in Fig. 5(b).

Number of Historical Frames. We train and infer Sparse4D with varying numbers of historical frames and find that model performance continues to grow as the number of frames increases(Fig. 5(c)). Even if the number of frames increases to 10 (equivalent to 5 seconds in history), there is still a small increase compared to 8 frames. There may still be room for improvement in Sparse4D’s performance if more frames are added. However, due to the limitations of our training device’s memory (V100, 32G), it was not possible to try more frames.

FLOPs and Parameters. In this experiment, the input image size is set to 900x1600 and ResNet101 is used as backbone, and the experimental results are shown in Tab. 3. When T=1T=1, the FLOPs of our model is 1019.2G, and the parameter amount is 58.1M. Compared with DETR3D, the amount of calculation and parameter are only increased by 2.3%2.3\% and 9.0%9.0\%, respectively, and the mAP and NDS are increased by 3.6%3.6\% and 2.6%2.6\%. Compared with Lift-splat and BEVFormer-S, we have certain advantages in algorithm metrics, FLOPs and parameter amount. After adding 3 history frames, Sparse4D with temporal fusion only increases 9.3%9.3\% FLOPs and 1.4%1.4\% parameters, and achieves a very noticeable improvement, 5.4%5.4\% mAP and 9.0%9.0\% NDS.

4 Main Results

The comparison of the results on the nuScenes validation set is shown in Tab. 4. Among all the non-temporal models, Sparse4D gets the highest mAP and NDS. Compared with the baseline of the sparse methods, DERT3D, we improve mAP and NDS by 3.3%3.3\% and 1.7%1.7\%, respectively. Compared to the baseline of the BEV-based methods, BEVFormer, we lead by 0.7%0.7\% on mAP and 0.3%0.3\% on NDS. We also compared Sparse4D with other SOTA temporal algorithms, and still obtained the best NDS and mAP. When T=4T=4, Sparse4D outperforms BEVDepth by 2.4%2.4\% on mAP and 0.6%0.6\% on NDS. When TT is increased from 4 to 9, the mAP and NDS of Sparse4D are improved by 0.9%0.9\% and 0.6%0.6\% respectively. Moreover, adding 24 epochs to training can further improve NDS by 0.3%0.3\%.

We compare Sparse4D with other SOTA algorithms on the nuScenes test set (online leaderboard). As shown in the Tab. 5, with DD3D pre-trained VoVNet-99, Sparse4D achieves 51.1%51.1\% and 59.5%59.5\% on mAP and NDS metrics, respectively, outperforming all non-BEV methods including PETRv2. Compared with baseline DETR3D, our method has achieved significant improvements. The mAP and NDS have increased by 9.9%9.9\% and 11.6%11.6\%, respectively, greatly improving the competitiveness of sparse methods. In addition, Sparse4D is also superior to the dense BEV based methods including UVTR , BEVdet , BEVFormer and BEVDistill , especially in mAP, which is 1.5%1.5\% higher than BEVDistill.

5 Extend to 3D Object Tracking

Based on the tracking-by-detection framework , Sparse4D is easily extended to a tracker. We use the instance features and bounding boxes output by the last refinement module to extract identity features, and use a lightweight sub-network to estimate the correlation matrix between historical trajectories and current objects. Then, the matching relationship between the historical trajectory and the current object will be obtained using the Hungarian matching algorithm. As shown in Tab. 6, Sparse4D obtains 0.519 AMOTA and 1.078 AMOTP on nuScenes test set, which is ahead of most learning-based methods.

Conclusion

In this work, we propose a new method, Sparse4D, which achieves feature-level fusion of multi-timestamp and multi-view through a deformable 4D aggregation module, and uses iterative refinement to achieve 3D box regression. Sparse4D can provide excellent perceptual performance, and it outperforms all existing sparse algorithms and most BEV-based algorithms on the nuScenes leaderboard.

We believe that Sparse4D still has a lot of room for improvement. For example, in the depth reweight module, multi-view stereo (MVS) technology can be added to obtain more accurate depth. Camera parameters can also be considered in the encoder to improve 3D generalization . Therefore, we hope that Sparse4D can become a new baseline for sparse 3D detection. In addition, the framework of Sparse4D can also be extended to other tasks, such as HD map construction, occupancy estimation, 3D reconstruction, etc.

References