MonoDTR: Monocular 3D Object Detection with Depth-Aware Transformer

Kuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. Hsu

Introduction

Three-dimensional (3D) object detection is a fundamental problem and enables various applications such as autonomous driving. Previous methods have achieved superior performance based on the accurate depth information from multiple sensors, such as LiDAR signal or stereo matching . In order to lower the sensor costs, some image-only monocular 3D object detection methods have been proposed and made impressive progress relying on geometry constraints between 2D and 3D. However, the performance is still far from satisfactory without the aid of depth cues.

Recently, several works have tried to produce estimated depth from the pre-trained depth estimation models to assist monocular 3D object detection. Pseudo-LiDAR-based approaches convert estimated depth maps into 3D point clouds to imitate LiDAR signals, followed by the existing LiDAR-based detector for 3D object detection (see Figure 1(a)). Some fusion-based approaches apply several fusion strategies to combine features extracted from depths and images to detect objects (see Figure 1(b)). These methods, though better localize objects with the help of estimated depth, may suffer from the risk of learning 3D detection on inaccurate depth maps. Also, the additional computational cost of the depth estimator makes it impractical for real-world applications .

To address the above issues, we propose MonoDTR, a novel end-to-end depth-aware transformer network for monocular 3D object detection (see Figure 1(c)). A depth-aware feature enhancement (DFE) module is introduced to learn depth-aware features with the auxiliary depth supervision, which avoids obtaining inaccurate depth priors from the pre-trained depth estimator. Furthermore, the DFE module is lightweight yet effective in assisting 3D object detection without constructing the complicated architecture to extract features from off-the-shelf depth maps. It significantly reduces computational time compared with previous depth-assisted methods (see Table 1).

In addition, unlike previous fusion-based methods (\eg, D4LCN and DDMP-3D) that apply carefully designed convolution kernels for context- and depth-aware features, we develop the first transformer-based fusion module to globally integrate the image and depth information. The transformer encoder-decoder structure has been proven to capture long-range dependency effectively; thus, we apply it to model the relationship between context- and depth-aware features. To better represent the property of the 3D object, we utilize depth-aware features to replace the commonly used object queries as input of the transformer decoder, which can provide more meaningful cues for 3D reasoning. Furthermore, we introduce a novel depth positional encoding (DPE) to involve depth-aware hints to the transformer, achieving better performance than conventional pixel-wise positional encodings.

We summarize our contributions as follows:

We propose a novel framework, MonoDTR, learning depth-aware features via auxiliary supervision to assist monocular 3D object detection, which avoids introducing high computational cost and inaccurate depth priors from using the off-the-shelf depth estimator.

We present the first depth-aware transformer module to integrate context- and depth-aware features efficiently. A novel depth positional encoding (DPE) is proposed to inject depth positional hints into the transformer.

Experimental results on the KITTI dataset show that our approach outperforms state-of-the-art monocular-based methods and achieves real-time detection. Furthermore, the proposed depth-aware modules can be easily plug-and-play in existing image-only frameworks to improve performance.

Related Work

Image-only monocular 3D object detection. Recently, several works only adopt a single image for 3D object detection . Due to the lack of depth information from images, these methods mainly rely on geometric consistency to predict objects. Deep3Dbox solves orientation prediction by proposed novel MultiBin loss and enforces constraint between 2D and 3D boxes with geometric prior. M3D-RPN generates 3D object proposals with 2D bounding box constraints and proposes a depth-aware convolution to predict 3D objects. OFTNet introduces an orthographic feature transform to map image-based features into a 3D voxel space. Besides, MonoPair explores spatial pair-wise relationship between objects to improve detection performance. M3DSSD presents a two-step feature alignment approach to solve the feature mismatching problem. Furthermore, some works predict keypoints of the 3D bounding box as an intermediate task to recover the location of objects. However, the above purely monocular methods fail to accurately localize objects due to the lack of depth cues.

Depth-assisted monocular 3D object detection. To further improve the performance, many approaches propose using depth information to aid 3D object detection . Some prior works transform image into pseudo-LiDAR representation by leveraging off-the-shelf depth estimator and calibration parameters, followed by the existing LiDAR-based 3D detector to predict objects, leading to progressive improvement. PatchNet reveals that the success of pseudo-LiDAR comes from the coordinate transformation and organizes it into the image representation, which can benefit from the powerful CNNs networks. D4LCN and DDMP-3D focus on developing the fusion-based approach between image and estimated depth with carefully designed convolutional networks. Besides, CaDDN learns categorical depth distributions for each pixel to construct bird’s-eye-view (BEV) representations and recovers bounding boxes from the BEV projection. However, most of the abovementioned methods directly using pre-trained depth estimators suffer from additional computational cost and only achieve limited performance caused by inaccurate depth priors.

Transformer. Transformer was firstly introduced in sequential modeling and has considerable improvement in natural language processing (NLP) tasks. The self-attention mechanism is the core component in the transformer with its capability of capturing the long-range dependencies. Recently, transformer architecture has been successfully leveraged in the computer vision field, such as image classification and human-object interaction . In addition, DETR proposes developing object detection with the transformer without relying on many hand-designed components used in traditional pipelines.

Though the transformer can perform well in most visual tasks, its usage in monocular 3D object detection has not been explored. In the image-based 3D detection task, the object size at far and near distance in the image varies significantly due to the perspective projection , which makes it challenging to utilize the learned object query mentioned in DETR to fully represent the object property. Thus, in this paper, we propose to globally integrate context- and depth-aware features with transformers and inject depth hints into the transformer for better 3D reasoning.

Proposed Approach

2 Depth-Aware Feature Enhancement Module

Existing depth-assisted methods , using off-the-shelf depth estimators, suffer from the risk of introducing inaccurate depth priors and extra computation burden. To alleviate this, we propose a depth-aware feature enhancement (DFE) module for depth reasoning as in Figure 3. The precise depth map is utilized for auxiliary supervision in the training stage, making the DFE module implicitly learn the depth-aware features. Compared with previous works that apply an additional backbone or complicated architectures to encode depths, we generate depth-aware features to assist 3D object detection with a lightweight module, significantly reducing the computation budget.

Depth prototype representation learning. To further enhance the capability of depth representation, we augment the feature of each pixel by introducing the central representation of the corresponding depth category (bin), inspired from the class center in . The feature center of each depth category (regarded as the depth prototype) can be computed by aggregating the depth-aware features of each pixel belonging to a specified category. In practice, we first apply a group convolution to the predicted depth map D\mathbf{D} to merge the adjacent depth categories (bins), reducing the class number from DD to D′=D/rD^{\prime}=D/r with scale rr. It helps to share similar depth cues and reduce computation. The representation of depth prototype Fd\mathbf{F}_{d} can be generated by gathering the feature of all pixels X′\mathbf{X^{\prime}} weighted by their probability to the depth category dd:

Feature enhancement with depth prototype. Now we can reconstruct new depth-aware features based on the depth prototype representation, which allows each pixel to understand the presentation of the depth category from the global view. The reconstructed feature F′\mathbf{F}^{\prime} is calculated as:

Consequently, we obtain the enhanced depth feature by concatenating the initial depth-aware feature X\mathbf{X} and the reconstructed features F′\mathbf{F}^{\prime}, followed by a simple 1 ×\times 1 convolution layer, as shown in Figure 3(c).

3 Depth-Aware Transformer

Inspired by the tremendous success of the transformer on modeling the long-range relationships, we exploit the transformer encoder-decoder architecture to construct the depth-aware transformer (DTR) module to globally integrate the context- and depth-aware features.

Transformer decoder. The decoder is also built upon the standard transformer architecture. We propose utilizing the depth-aware features as the input of the decoder instead of learnable embeddings (object query) , which is different from the common usage in previous encoder-decoder vision transformer works . The main reason is that, in the monocular 3D object detection task, the camera views at near and far distances often cause significant changes in object scale due to the perspective projection . It makes the simple learnable embedding hard to fully represent the object’s property and handle complex scale variant situations. In contrast, plentiful distance-aware cues are hidden in the depth-aware features. Thus, we propose adopting depth-aware features as the input of the transformer decoder. To this end, the decoder can take the power of cross-attention modules in the transformer to efficiently model the relationship between context- and depth-aware features, leading to better performance.

Computation reduction. The standard self-attention layer in Equation 3 leads to O(N2)\mathcal{O}(N^{2}) time and memory, which damages the computational budget. To mitigate this issue, more recent works make efforts on accelerating the attention operation. Among these methods, Linear transformer proposes to approximate softmax operation with the linear dot product of features. Specifically, the similarity function in original transformer can be formulated as: sim(q,k)=exp(q⊤kC){\rm sim}(q,k)={\rm exp}(\frac{q^{\top}k}{\sqrt{C}}). It is replaced with sim(q,k)=ϕ(q)ϕ(k){\rm sim}(q,k)=\phi(q)\phi(k) in , where ϕ(x)=elu(x)+1\phi(x)=\text{elu}(x)+1, and elu(⋅)\text{elu}(\cdot) is the exponential linear unit activation function. To this end, ϕ(K)⊤\phi(K)^{\top} and VV can be combined first to reduce computation to O(N)\mathcal{O}(N). We refer the readers to for more details. In our transformer, we consider applying the linear attention described in to replace the vanilla self-attention for higher inference speed.

4 2D-3D Detection and Loss

Anchor definition. We adopt the single-stage detector with the pre-defined 2D-3D anchors to regress the bounding box. Each pre-defined anchor consists parameters with 2D bounding box [x2d,y2d,w2d,h2d][x_{2d},y_{2d},w_{2d},h_{2d}] and 3D bounding box [xp,yp,z,w3d,h3d,l3d,θ][x_{p},y_{p},z,w_{3d},h_{3d},l_{3d},\theta]. [x2d,y2d][x_{2d},y_{2d}] and [xp,yp][x_{p},y_{p}] represent the 2D box center and 3D object center projected to image plane. [w2d,h2d][w_{2d},h_{2d}] and [w3d,h3d,l3d]{[w_{3d},h_{3d},l_{3d}]} represent the physical dimension of 2D and 3D bounding box, respectively. zz denotes the depth of 3D object center. θ\theta is the observation angle. During training, we project all ground truth into the 2D space to calculate the intersection over union (IoU) with all 2D anchors. The anchor with IoU greater than 0.5 is chosen to assign with the corresponding 3D box for optimization.

Output transformation. Similar to prior works , we follow Yolov3 to predict [tx,ty,tw,th]2d[t_{x},t_{y},t_{w},t_{h}]_{2d} and [tx,ty,tw,th,tl,tz,tθ]3d[t_{x},t_{y},t_{w},t_{h},t_{l},t_{z},t_{\theta}]_{3d} for each anchor, which aims at parameterizing the residual value for 2D and 3D bounding box, and also predict the classification scores clscls. The output bounding box can be restored based on the anchor and the network prediction as follows:

where (⋅)^\hat{(\cdot)} denotes the recovered parameters of the 3D object. Note that we apply the same anchor center for 2D box center [x2d,y2d][x_{2d},y_{2d}] and 3D projection center [xp,yp][x_{p},y_{p}].

Loss function. The overall loss L\mathcal{L} contains a classification loss Lcls\mathcal{L}_{cls} for objectness and class, a bounding box regression loss Lreg\mathcal{L}_{reg} to optimize Equation 4, and a depth loss Ldep\mathcal{L}_{dep} with auxiliary depth supervision described in Section 3.2:

We adopt the focal loss to balance the samples for the classification task, and the smoothed-L1 loss for the regression task. For the depth categorical prediction described in Section 3.2, we utilize the focal loss :

where P\mathcal{P} is the pixel region on the image with the valid depth labels, and D^\mathbf{\hat{D}} is the depth bins ground truth generated from LiDAR (more details are provided in the supplementary material).

Experiments

Dataset. We evaluate the proposed approach on the challenging KITTI 3D object detection dataset , which is the most commonly used benchmark for the 3D object detection task. It contains 7481 images for training and 7518 images for testing. We follow to divide training samples into the training set (3712) and the validation set (3769). The ablation studies are conducted based on this split.

Evaluation metric. The average precision (AP) is used as the metric for evaluation in both 3D object detection and bird’s eye view (BEV) detection tasks. We utilize 40 recall positions metric AP40 instead of original AP11 to avoid the bias . The difficulty of the detection in the benchmark is divided into three levels (”Easy”, ”Mod.”, ”Hard”) according to size, occlusion, and truncation. All methods are ranked based on AP3D{}_{3\text{D}} of moderate setting (Mod.) same as the KITTI benchmark. The thresholds of Intersection over Union (IoU) are 0.7, 0.5, 0.5 for cars, cyclists, and pedestrians categories following the official setting.

Implementation details. We use Adam optimizer to train our network for 120 epochs with batch size 4. The learning rate starts at 0.0001 and decays with a cosine annealing schedule. We apply 48 anchors on each pixel of the feature map with 3 aspect ratios of {0.5,1.0,1.5}\{0.5,1.0,1.5\}, and 12 scales in height following the exponential function 24×2i/4,i={0,...,15}24\times 2^{i/4},i=\{0,...,15\}. For 3D anchor parameters, we calculate the mean and variance statistics of 3D ground truth in the training dataset as prior statistical knowledge of each anchor. Following , we crop the top 100 pixels of each image to reduce inference time, and all images are resized to 288 ×\times 1280. In the training stage, we apply random horizontal mirroring as data augmentation. In the inference stage, we drop the predictions with a confidence score lower than 0.75 and adopt Non-Maximum Suppression (NMS) with IoU 0.4 to reduce redundancy.

2 Main Results

Results of the Car category on the KITTI test set. As shown in Table 1, we compare our MonoDTR with several state-of-the-art monocular 3D object detection methods on the KITTI test set. It can be observed that our approach achieves better performance than other methods in terms of the moderate level of the two tasks, which is the most important metric in the benchmark. Furthermore, it is worth noting that our approach outperforms other depth-assisted methods by large margins. For instance, compared to top three depth-assist methods, DFRNet , CaDDN and DDMP-3D , our MonoDTR obtains 2.59/1.76/2.38, 2.82/1.98/1.27 and 2.28/2.61/2.93 improvements in AP3D{}_{3\text{D}} at IoU threshold 0.7 on three settings, which indicates the effectiveness of the proposed depth-aware modules.

Results of the Car category on the KITTI validation set. We also conduct experiments on the KITTI validation dataset under different IoU thresholds and tasks as listed in Table 2. Our approach obtains superior performance over several image-only methods, benefiting from the auxiliary depth supervision. Specifically, compared with GUPNet , our method achieves significant improvements of 6.41/4.99/4.61 in AP3D{}_{3\text{D}} and 7.26/5.41/5.02 in APBEV{}_{\text{BEV}} at IoU threshold 0.5 on the easy, moderate, and hard settings.

Results of Pedestrians and Cyclists categories on the KITTI test set. We further present the performance of pedestrians and cyclists categories in Table 3. Detecting these two categories is more challenging than cars due to their smaller size and non-rigid body, making it difficult to precisely locate the position. Overall, our model significantly outperforms all methods on pedestrian category with a considerable margin. For the cyclist 3D detection, we also achieve competitive results to CaDDN and obtain better performance than other methods.

Running time analysis. We measure the average running time for processing the whole validation set with batch size 1 on a single Nvidia Tesla v100 GPU. As shown in Table 1, our model can achieve real-time performance at 27 FPS, which confirms the efficiency of our approach. Compared with state-of-the-art depth-assisted methods, our MonoDTR runs 17×\times and 4.8×\times faster than CaDDN and DDMP-3D , respectively. The main reasons can be summarized as follows: (1) CaDDN builds the bird’s eye view representation from predicted depth maps to perform 3D detection, which applies more complicated architecture to generate precise depth predictions, leading to slow speed. (2) Fusion-based methods often utilize two separate backbones for extracting features of image and depth, which is time-consuming. Note that the depth estimator also takes additional inference time, which is not included in Table 1. On the contrary, our model learns depth-aware features through the lightweight DFE module with auxiliary supervision, which reduces running time significantly.

3 Ablation Study

Effectiveness of each proposed components. In Table 4, we conduct an ablation study to analyze the effectiveness of the proposed components: (a) Baseline: only using context-aware features for 3D object detection, \ie, without all proposed depth-aware modules. (b) Replacing depth-aware features with object query in the transformer, \iebaseline + DETR-like transformer. (c) Replacing depth-aware features with features extracted from depth images generated by DORN . (d) Integrating context- and depth-aware features with the convolutional concatenate operation. (e) Full model without depth prototype enhanced feature F′\mathbf{F^{\prime}}. (f) MonoDTR (full model).

Firstly, we can observe from (b→\rightarrowf) that utilizing depth-aware features to replace the object query in the transformer can provide meaningful depth hints and improve the performance. Besides, compared to our end-to-end training framework (f), simply utilizing depth priors from the pretrained depth estimator (c) leads to worse results. Next, we demonstrate that applying our depth-aware transformer (DTR) module (f) can more effectively integrate context- and depth-aware features than simple convolutional concatenation (d). Furthermore, utilizing our proposed depth prototype enhancement module can boost the performance (e→\rightarrowf). Finally, by applying all the designed modules, our full model (f) achieves significant improvement compared to the baseline (a). Also, an in-depth analysis in Figure 5 suggests that our method surpasses the baseline under different IoU thresholds and object depths. These results prove the effectiveness of our depth-aware modules.

Comparison with different positional encodings. We investigate the effectiveness of the proposed depth positional encoding (DPE) in Table 5. Compared with several commonly used positional encodings, including absolute positional encoding (APE) , conditional positional encoding (CPE) , sinusoidal positional encoding , and without using positional encoding (No PE), our proposed DPE achieves better performance on KITTI validation set. We believe that encoding the depth-aware cues is more effective for learning the position representation of 3D tasks than pixel-level encodings.

Plugging into the existing image-only methods. Our proposed approach is flexible to extend to existing image-only 3D object detectors to improve the depth reasoning capability. We respectively plug our depth-aware modules into three popular monocular 3D object detectors: M3D-RPN , GAC , and MonoDLE , based on their official codeshttps://github.com/garrickbrazil/M3D-RPNhttps://github.com/Owen-Liuyuxuan/visualDet3Dhttps://github.com/xinzhuma/monodle. In practice, we take the features from the above models (before the detection head) as the initial features, and utilize our proposed modules (DFE, DTR, and DPE modules) to generate final integrated features, followed by their original detection head to detect 3D objects. As shown in Table 6, with the aid of our proposed depth-aware modules, these detectors achieve further improvements on the KITTI validation set, which demonstrates the flexibility and efficiency of our approach.

4 Qualitative Results

We provide the qualitative examples on the KITTI validation set in Figure 6. Compared with the baseline model without the aid of depth-aware modules, the predictions from MonoDTR are much closer to the ground truth. It shows that the proposed depth-aware modules can help to locate the object precisely. More qualitative results are included in the supplementary material.

Conclusion

In this paper, we propose a depth-aware transformer network for monocular 3D object detection. The proposed lightweight DFE module implicitly learns depth-aware features in an end-to-end fashion to avoid obtaining inaccurate depth priors and high computational cost from an off-the-shelf depth estimator. We also introduce the depth-aware transformer to globally integrate context- and depth-aware features, while the novel depth positional encoding (DPE) is designed to inject depth hints into the transformer. Comprehensive experiments on the KITTI dataset validate that our model achieves real-time detection and outperforms previous state-of-the-art monocular-based methods.

This work was supported in part by the Ministry of Science and Technology, Taiwan, under Grant MOST 110-2634-F-002-051, Qualcomm Technologies, Inc., and Mobile Drive Technology Co., Ltd (MobileDrive). We are grateful to the National Center for High-performance Computing.

References

Appendix A Depth-Aware Transformer

Transformer architecture. The detailed architecture of depth-aware transformer (DTR) is shown in Figure 7. The encoder aims to generate the encoded context-aware features, while the decoder produces the fused feature from context- and depth-aware features through the multiple self-attention layers. Besides, we supplement two features with the proposed depth positional encoding (DPE) before passing them to the transformer, enabling better 3D reasoning.

Effectiveness of linear attention. Table 7 shows the results of different self-attention layers on the KITTI dataset, where we can observe that applying linear attention can achieve almost 4 ×\times faster than vanilla self-attention with comparable performance. Thus, we adopt the linear attention in our transformers for real-time applications.

Appendix B Auxiliary Depth Supervision

Depth ground truth generation. We project the LiDAR signals into the image plane to generate the sparse ground truth depth map. Then we apply linear-increasing discretization (LID) method to convert continuous depth dd to discretized depth bins. The LID is defined as follows:

where ii is the depth bin index. The number of depth bins DD is set as 96, and the range of depth [dmin⁡,dmax⁡d_{\min},d_{\max}] is set as . Note that the pixels with the depth value outside the range will be marked as invalid and not used for optimization during training.

Different discretization methods. In Table 8, we investigate the effectiveness of different discretization methods for depth auxiliary supervision. In addition to the LID method, the continuous depth can be discretized using uniform discretization (UD) with fixed bin size: dmax⁡−dmin⁡D\frac{d_{\max}-d_{\min}}{D}, or spacing-increasing discretization (SID) with the increasing bin size in the log space. It can be observed that using the LID strategy can achieve better performance, so we apply it as our discretization method.

Appendix C Results on nuScenes Dataset

Table 9 shows the experimental results of deploying our proposed approach on nuScenes val set. Under the same configurations (\eg, backbone and training schedule), our model achieves better performance than two 3D object detection baselines (FCOS3D , and PGD ), which demonstrates the effectiveness of our approach.

Appendix D Qualitative Visualization

More visualization results. In Figure 8, we provide some qualitative results on the KITTI dataset for multiple-category predictions. In Figure 9, we show the qualitative comparisons of the baseline (without proposed depth-aware modules) and our MonoDTR (full model). It can be observed that our MonoDTR can generate higher quality bounding boxes benefit from the aid of depth cues.

Failure case. We show a representative failure case in Figure 10. The lower-quality 3D bounding box is caused by the inaccurately predicted object depth, which is typical in most monocular 3D object detection tasks.

Appendix E Broader Impacts

Our work aims to develop the monocular 3D object detection approach for autonomous driving. The proposed model may generate inaccurate object depth prediction, leading to incorrect downstream decision-making and potential traffic accidents. Furthermore, we provide a new perspective of leveraging learned depth-aware features to assist monocular 3D object detection. Although considerable progress has been made with our proposed lightweight depth-aware feature extraction module, we believe it is worth further exploring how to learn depth-aware features to effectively improve detection performance.