M3D-RPN: Monocular 3D Region Proposal Network for Object Detection

Garrick Brazil, Xiaoming Liu

Introduction

Scene understanding in 3D plays a principal role in designing effective real-world systems such as in urban autonomous driving and robotics . Currently, the foremost methods on 3D detection rely extensively on expensive LiDAR sensors to provide sparse depth data as input. In comparison, monocular image-only 3D detection is considerably more difficult due to an inherent lack of depth cues. As a consequence, the performance gap between LiDAR-based methods and monocular approaches remains substantial.

Prior work on monocular 3D detection have each relied heavily on external state-of-the-art (SOTA) sub-networks, which are individually responsible for performing point cloud generation , semantic segmentation , 2D detection , or depth estimation . A downside to such approaches is an inherent disconnection in component learning as well as system complexity. Moreover, reliance on additional sub-networks can introduce persistent noise, contributing to a limited upper-bound for the framework.

In contrast, we propose a single end-to-end region proposal network for multi-class 3D object detection (Fig. 1). We observe that 2D object detection performs reasonably and continues to make rapid advances . The 2D and 3D detection tasks each aim to ultimately classify all instances of an object; whereas they differ in the dimensionality of their localization targets. Intuitively, we expect the power of 2D detection can be leveraged to guide and improve the performance of 3D detection, ideally within a unified framework rather than as separate components. Hence, we propose to reformulate the 3D detection problem such that both 2D and 3D spaces utilize shared anchors and classification targets. In doing so, the 3D detector is naturally able to perform on par with the performance of its 2D counterpart, from the perspective of reliably classifying objects. Therefore, the remaining challenge is reduced to 3D localization within the camera coordinate space.

To address the remaining difficultly, we propose three key designs tailored to improve 3D estimation. Firstly, we formulate 3D anchors to function primarily within the image-space and initialize all anchors with prior statistics for each of its 3D parameters. Hence, each discretized anchor inherently has a strong prior for reasoning in 3D, based on the consistency of a fixed camera viewpoint and the correlation between 2D scale and 3D depth. Secondly, we design a novel depth-aware convolutional layer which is able to learn spatially-aware features. Traditionally, convolutional operations are preferred to be spatially-invariant in order to detect objects at arbitrary image locations. However, while it is likely beneficial for low-level features, we show that high-level features improve when given increased awareness of their depth and while assuming a consistent camera scene geometry. Lastly, we optimize the orientation estimation θ\theta using 3D →\rightarrow 2D projection consistency loss within a post-optimization algorithm. Hence, helping correct anomalies within θ\theta estimation while assuming a reliable 2D bounding box.

To summarize, our contributions are the following:

We formulate a standalone monocular 3D region proposal network (M3D-RPN) with a shared 2D and 3D detection space, while using prior statistics to serve as strong initialization for each 3D parameter.

We propose depth-aware convolution to improve the 3D parameter estimation, thereby enabling the network to learn more spatially-aware high-level features.

We propose a simple orientation estimation post-optimization algorithm which uses 3D projections and 2D detections to improve the θ\theta estimation.

We achieve state-of-the-art performance on the urban KITTI benchmark for monocular Bird’s Eye View and 3D Detection using a single multi-class network.

Related Work

2D Detection: Many works have addressed 2D detection in both generic and urban scenes . Most recent frameworks are based on seminal work of Faster R-CNN due to the introduction of the region proposal network (RPN) as a highly effective method to efficiently generate object proposals. The RPN functions as a sliding window detector to check for the existence of objects at every spatial location of an image which match with a set of predefined template shapes, referred to as anchors. Despite that the RPN was conceived to be a preliminary stage within Faster R-CNN, it is often demonstrated to have promising effectiveness being extended to a single-shot standalone detector . Our framework builds upon the anchors of a RPN, specially designed to function in both the 2D and 3D spaces, and acting as a single-shot multi-class 3D detector.

LiDAR 3D Detection: The use of LiDAR data has proven to be essential input for SOTA frameworks for 3D object detection applied to urban scenes. Leading methods tend to process sparse point clouds from LiDAR points or project the point clouds into sets of 2D planes . While the LiDAR-based methods are generally high performing for a variety of 3D tasks, each is contingent on the availability of depth information generated from the LiDAR points or directly processed through point clouds. Hence, the methods are not applicable to camera-only applications as is the main purpose of our monocular 3D detection algorithm.

Image-only 3D Detection: 3D detection using only image data is inherently challenging due to an overall lack of reliable depth information. A common theme among SOTA image-based 3D detection methods is to use a series of sub-networks to aid in detection. For instance, uses a SOTA depth prediction with stereo processing to estimate point clouds. Then 3D cuboids are exhaustively placed along the ground plane given a known camera projection matrix, and scored based upon the density of the cuboid region within the approximated point cloud. As a follow-up, adjusts the design from stereo to monocular by replacing the point cloud density heuristic with a combination of estimated semantic segmentation, instance segmentation, location, spatial context and shape priors, used while exhaustively classifying proposals on the ground plane.

In recent work, uses an external SOTA object detector to generate 2D proposals then processes the cropped proposals within a deep neural network to estimate 3D dimensions and orientation. Similar to our work, the relationship between 2D boxes and 3D boxes projected onto the image plane is then exploited in post-processing to solve for the 3D parameters. However, our model directly predicts 3D parameters and thus only optimizes to improve θ\theta, which converges in ∼8\mathord{\sim}8 iterations in practice compared with 6464 iterations in . Xu et al. utilize an additional network to predict a depth map which is subsequently used to estimate a LiDAR-like point cloud. The point clouds are then sampled using 2D bounding boxes generated from a separate 2D RPN. Lastly, a R-CNN classifier receives an input vector consisting of the sampled point clouds and image features, to estimate the 3D box parameters.

In contrast to prior work, we propose a single network trained only with 3D boxes, as opposed to using a set of external networks, data sources, and composed of multiple stages. Each prior work use external networks for at least one component of their framework, some of which have also been trained on external data. To the best of our knowledge, our method is the first to generate 2D and 3D object proposals simultaneously using a Monocular 3D Region Proposal Network (M3D-RPN). In theory, M3D-RPN is complementary to prior work and may be used to replace the proposal generation stage. A comparison between our method and prior is further detailed in Fig. 2.

M3D-RPN

Our framework is comprised of three key components. First, we detail the overall formulation of our multi-class 3D region proposal network. We then outline the details of depth-aware convolution and our collective network architecture. Finally, we detail a simple, but effective, post-optimization algorithm for increased 3D→\to2D consistency. We refer to our method as Monocular 3D Region Proposal Network (M3D-RPN), as illustrated in Fig. 3.

The core foundation of our proposed framework is based upon the principles of the region proposal network (RPN) first proposed in Faster R-CNN , tailored for 3D. From a high-level, the region proposal network acts as sliding window detector which scans every spatial location of an input image for objects matching a set of predefined anchor templates. Then matches are regressed from the discretized anchors into continuous parameters of the estimated object.

The θ3D\theta_{\text{3D}} represents the observation viewing angle . Compared to the Y-axis rotation in the camera coordinate system, the observation angle accounts for the relative orientation of the object with respect to the camera viewing angle rather than the Bird’s Eye View (BEV) of the ground plane. Therefore, the viewing angle is intuitively more meaningful to estimate when dealing with image features. We encode the remaining 3D dimensions [w,h,l]3D[w,h,l]_{\text{3D}} as given in the camera coordinate system.

The mean statistic for each zPz_{\mathbf{P}} and [w,h,l,θ]3D[w,h,l,\theta]_{\text{3D}} is pre-computed for each anchor individually, which acts as strong prior to ease the difficultly in estimating 3D parameters. Specifically, for each anchor we use the statistics across all matching ground truths which have ≥0.5\geq 0.5 intersection over union (IoU) with the bounding box of the corresponding [w,h]2D[w,h]_{\text{2D}} anchor. As a result, the anchors represent discretized templates where the 3D priors can be leveraged as a strong initial guess, thereby assuming a reasonably consistent scene geometry. We visualize the anchor formulation as well as precomputed 3D priors in Fig. 4.

where xPx_{\mathbf{P}} and yPy_{\mathbf{P}} denote spatial center location of each box. The transformed box b2D′{b}^{\prime}_{\text{2D}} is thus defined as [x,y,w,h]2D′[x,y,w,h]^{\prime}_{\text{2D}}. The following 77 outputs represent transformations denoting the projected center [tx,ty,tz]P[t_{x},t_{y},t_{z}]_{\mathbf{P}}, dimensions [tw,th,tl]3D[t_{w},t_{h},t_{l}]_{\text{3D}} and orientation tθ3Dt_{\theta_{\text{3D}}}, which we collectively refer to as b3Db_{\text{3D}}. Similar to 2D, the transformation is applied to an anchor with parameters [w,h]2D[{w},{h}]_{\text{2D}}, zPz_{\mathbf{P}}, and [w,h,l,θ]3D[{w},{h},{l},{\theta}]_{\text{3D}} as follows:

Hence, b3D′b^{\prime}_{\text{3D}} is then denoted as [x,y,z]P′[x,y,z]^{\prime}_{\mathbf{P}} and [w,h,l,θ]3D′[w,h,l,\theta]^{\prime}_{\text{3D}}. As described, we estimate the projected 3D center rather than camera coordinates to better cope with the convolutional features based exclusively in the image space. Therefore, during inference we back-project the projected 3D center location from the image space [x,y,z]P′[x,y,z]^{\prime}_{\mathbf{P}} to camera coordinates [x,y,z]3D′[x,y,z]^{\prime}_{\text{3D}} by using the inverse of Eqn. 1.

Loss Definition: The network loss of our framework is formed as a multi-task learning problem composed of classification LcL_{c} and a box regression loss for 2D and 3D, respectfully denoted as Lb2DL_{b_{\text{2D}}} and Lb3DL_{b_{\text{3D}}}. For each generated box, we check if there exists a ground truth with at least ≥0.5\geq 0.5 IoU, as in . If yes then we use the best matched ground truth for each generated box to define a target with τ{\tau} class index, 2D box b^2D\hat{b}_{\text{2D}}, and 3D box b^3D\hat{b}_{\text{3D}}. Otherwise, τ\tau is assigned to the catch-all background class and bounding box regression is ignored. A softmax-based multinomial logistic loss is used to supervise for LcL_{\text{c}} defined as:

We use a negative logistic loss applied to the IoU between the matched ground truth box b^2D\hat{b}_{\text{2D}} and the transformed b2D′b^{\prime}_{\text{2D}} for Lb2DL_{b_{\text{2D}}}, similar to , defined as:

The remaining 3D bounding box parameters are each optimized using a Smooth L1L_{1} regression loss applied to the transformations b3Db_{\text{3D}} and the ground truth transformations g^3D\hat{g}_{\text{3D}} (generated using b^3D\hat{b}_{\text{3D}} following the inverse of Eqn. 3):

Hence, the overall multi-task network loss LL, including regularization weights λ1\lambda_{1} and λ2\lambda_{2}, is denoted as:

2 Depth-aware Convolution

Spatial-invariant convolution has been a principal operation for deep neural networks in computer vision . We expect that low-level features in the early layers of a network can reasonably be shared and are otherwise invariant to depth or object scale. However, we intuitively expect that high-level features related to 3D scene understanding are dependent on depth when a fixed camera view is assumed. As such, we propose depth-aware convolution as a means to improve the spatial-awareness of high-level features within the region proposal network, as illustrated in Fig. 3.

The depth-aware convolution layer can be loosely summarized as regular 2D convolution where a set of discretized depths are able to learn non-shared weights and features. We introduce a hyperparameter bb denoting the number of row-wise bins to separate a feature map into, where each learns a unique kernel kk. In effect, depth-aware kernels enable the network to develop location specific features and biases for each bin region, ideally to exploit the geometric consistency of a fixed viewpoint within urban scenes. For instance, high-level semantic features, such as encoding a feature for a large wheel to detect a car, are valuable at close depths but not generally at far depths. Similarly, we intuitively expect features related to 3D scene understanding are inherently related to their row-wise image position.

An obvious drawback to using depth-aware convolution is the increase of memory footprint for a given layer by ×b\times b. However, the total theoretical FLOPS to perform convolution remains consistent regardless of whether kernels are shared. We implement the depth-aware convolution layer in PyTorch by unfolding a layer LL into bb padded bins then re-purposing the group convolution operation to perform efficient parallel operations on a GPUIn practice, we observe a 10−20%10-20\% overhead for reshaping when implemented with parallel group convolution in PyTorch ..

3 Network Architecture

The backbone of our network uses DenseNet-121 . We remove the final pooling layer to keep the network stride at 1616, then dilate each convolutional layer in the last DenseBlock by a factor of 22 to obtain a greater field-of-view.

We connect two parallel paths at the end of the backbone network. The first path uses regular convolution where kernels are shared spatially, which we refer to as global. The second path exclusively uses depth-aware convolution and is referred to as local. For each path, we append a proposal feature extraction layer using its respective convolution operation to generate Fglobal\mathbf{F}_{\text{global}} and Flocal\mathbf{F}_{\text{local}}. Each feature extraction layer generates 512512 features using a 3×33\times 3 kernel with 11 padding and is followed by a ReLU non-linear activation. We then connect the 1212 outputs to each F\mathbf{F} corresponding to c,[tx,ty,tw,th]2D,[tx,ty,tz]P,[tw,th,tl,tθ]3Dc,[t_{x},t_{y},t_{w},t_{h}]_{\text{2D}},[t_{x},t_{y},t_{z}]_{\mathbf{P}},[t_{w},t_{h},t_{l},t_{\theta}]_{\text{3D}}. Each output uses a 1×11\times 1 kernel and are collectively denoted as Oglobal\mathbf{O}_{\text{global}} and Olocal\mathbf{O}_{\text{local}}. To leverage the depth-aware and spatial-invariant strengths, we fuse each output using a learned attention α\alpha (after sigmoid) applied for i=1…12i=1\dots 12 as follows:

4 Post 3D→→\rightarrow2D Optimization

We optimize the orientation parameter θ\theta in a simple but effective post-processing algorithm (as detailed in Alg. 1). The proposed optimization algorithm takes as input both the 2D and 3D box estimations b2D′b^{\prime}_{\text{2D}}, [x,y,z]P′[x,y,z]^{\prime}_{\mathbf{P}}, and [w,h,l,θ]3D′[w,h,l,\theta]^{\prime}_{\text{3D}}, as well as a step size σ\sigma, termination β\beta, and decay γ\gamma parameters. The algorithm then iteratively steps through θ\theta and compares the projected 3D boxes with b2D′b^{\prime}_{\text{2D}} using a L1L_{1} loss. The 3D→\rightarrow2D box-project function is defined as follows:

where \mathbf{P}^{\raisebox{0.60275pt}{\scriptscriptstyle-1}} is the inverse projection after padding $,and, and\phidenotesanindexforaxisdenotes an index for axis[x,y,z].Wethenusetheprojectedboxparameterizedby. We then use the projected box parameterized by\rho=[x_{\text{min}},y_{\text{min}},x_{\text{max}},y_{\text{max}}]andthesourceand the sourceb^{\prime}_{\text{2D}}tocomputeato compute aL_{1}loss,whichactsasthedrivingheuristic.Whenthereisnoimprovementtothelossusingloss, which acts as the driving heuristic. When there is no improvement to the loss using\theta\pm\sigma,wedecaythestepby, we decay the step by\gammaandrepeatwhileand repeat while\sigma\geq\beta$.

5 Implementation Details

We implement our framework using PyTorch and release the code at http://cvlab.cse.msu.edu/project-m3d-rpn.html. To prevent local features from overfitting on a subset of the image regions, we initialize the local path with pretrained global weights. In this case, each stage is trained for 50k50k iterations. We expect higher degrees of data augmentation or an iterative binning schedule, e.g., b=2ib=2^{i} from i=0…log⁡2(bfinal)i=0\dots\log_{2}(b_{\text{final}}), could enable more ease of training at the cost of more complex hyperparameters.

We use a learning rate of 0.0040.004 with a poly decay rate using power 0.90.9, a batch size of 22, and weight decay of 0.90.9. We set λ1=λ2=1\lambda_{1}=\lambda_{2}=1. All images are scaled to a height of 512512 pixels. As such, we use b=32b=32 bins for all depth-aware convolution layers. We use 1212 anchor scales ranging from 3030 to 400400 pixels following the power function of 30⋅1.265i30\cdot 1.265^{i} for i=0…11i=0\dots 11 and aspect ratios of [0.5,1.0,1.5][0.5,1.0,1.5] to define a total of 3636 anchors for multi-class detection. The 3D anchor priors are learned using these templates with the training dataset as detailed in Sec. 3.1. We apply NMS on the box outputs in the 2D space using a IoU criteria of 0.40.4 and filter boxes with scores <0.75<0.75. The 3D →\to 2D optimization uses settings of σ=0.3π\sigma=0.3\pi, β=0.01\beta=0.01, and γ=0.5\gamma=0.5. Lastly, we perform random mirroring and online hard-negative mining by sampling the top 20%20\% high loss boxes in each minibatch.

We note that M3D-RPN relies on 3D box annotations and a known projection matrix P\mathbf{P} per sequence. For extension to a dataset without these known, it may be necessary to predict the camera intrinsics and utilize weak supervision leveraging 3D-2D projection geometry as loss constraints.

Experiments

We evaluate our proposed framework on the challenging KITTI dataset under two core 3D localization tasks: Bird’s Eye View (BEV) and 3D Object Detection. We comprehensively compare our method on the official test dataset as well as two validation splits , and perform analysis of the critical components which comprise our framework. We further visualize qualitative examples of M3D-RPN on multi-class 3D object detection in diverse scenes (Fig. 5).

The KITTI dataset provides many widely used benchmarks for vision problems related to self-driving cars. Among them, the Bird’s Eye View (BEV) and 3D Object Detection tasks are most relevant to evaluate 3D localization performance. The official dataset consists of 7,4817{,}481 training images and 7,5187{,}518 testing images with 2D and 3D annotations for car, pedestrian, and cyclist. For each task we report the Average Precision (AP) under 33 difficultly settings: easy, moderate and hard as detailed in . Methods are further evaluated using different IoU criteria per class. We emphasize our results on the official settings of IoU ≥0.7\geq 0.7 for cars and IoU ≥0.5\geq 0.5 for pedestrians and cyclists.

We conduct experiments on three common data splits including val1 , val2 , and the official test split . Each split contains data from non-overlapping sequences such that no data from an evaluated frame, or its neighbors, have been used for training. We focus our comparison to SOTA prior work which use image-only input. We primarily compare our methods using the car class, as has been the focus of prior work . However, we emphasize that our models are trained as a shared multi-class detection system and therefore also report the multi-class capability for monocular 3D detection, as detailed in Tab. 3.

Bird’s Eye View: The Bird’s Eye View task aims to perform object detection from the overhead viewpoint of the ground plane. Hence, all 3D boxes are first projected onto the ground plane then top-down 2D detection is applied. We evaluate M3D-RPN on each split as detailed in Tab. 1.

M3D-RPN achieves a notable improvement over SOTA image-only detectors across all data splits and protocol settings. For instance, under criteria of IoU ≥0.7\geq 0.7 with val1, our method achieves 21.18%21.18\% (↑7.55%\uparrow 7.55\%) on moderate, and 17.90%17.90\% (↑6.30%\uparrow 6.30\%) on hard. We further emphasize our performance on test which achieves 18.36%18.36\% (↑8.74%\uparrow 8.74\%) and 16.24%16.24\% (↑8.02%\uparrow 8.02\%) respectively on moderate and hard settings with IoU ≥0.7\geq 0.7, which is the most challenging setting.

3D Object Detection: The 3D object detection task aims to perform object detection directly in the camera coordinate system. Therefore, an additional dimension is introduced to all IoU computations, which substantially increases the localization difficulty compared to BEV task. We evaluate our method on 3D detection with each split under all commonly studied protocols as described in Tab. 2. Our method achieves a significant gain over state-of-the-art image-only methods throughout each protocol and split.

We emphasize that the current most difficult challenge to evalaute 3D localization is the 3D object detection task. Similarly, the moderate and hard settings with IoU ≥0.7\geq 0.7 are the most difficult protocols to evaluate with. Using these settings with val1, our method notably achieves 17.06%17.06\% (↑11.37%\uparrow 11.37\%) and 15.2115.21 (↑9.82%\uparrow 9.82\%) respectively. We further observe similar gains on the other splits. For instance, when evaluated using the testing dataset, we achieve 15.70%15.70\% (↑10.52\uparrow 10.52) and 13.32%13.32\% (↑8.64\uparrow 8.64) on the moderate and hard settings despite being trained as a shared multi-class model and compared to single model methods . When evaluated with less strict criteria such as IoU ≥0.5\geq 0.5, our method demonstrates smaller but reasonable margins (∼3−6%\mathord{\sim}3-6\%), implying that M3D-RPN has similar recall to prior art but significantly higher precision overall.

Multi-Class 3D Detection: To demonstrate generalization beyond a single class, we evaluate our proposed 3D detection framework on the car, pedestrian and cyclist classes. We conduct experiments on both the Bird’s Eye View and 3D Detection tasks using the KITTI test dataset, as detailed in Tab. 3. Although there are not monocular 3D detection methods to compare with for multi-class, it is noteworthy that the performance on pedestrian outperforms prior work performance on car, which usually has the opposite relationship, thereby suggesting a reasonable performance. However, M3D-RPN is noticeably less stable for cyclists, suggesting a need for advanced sampling or data augmentation to overcome the data bias towards car and pedestrian.

2D Detection: We evaluate our performance on 2D car detection (detailed in Tab. 4). We note that M3D-RPN performs less compared to other 3D detection systems applied to the 2D task. However, we emphasize that prior work use external networks, data sources, and include multiple stages (e.g., Fast , Faster R-CNN ). In contrast, M3D-RPN performs all tasks simultaneously using only a single-shot 3D proposal network. Hence, the focus of our work is primarily to improve 3D detection proposals with an emphasis on the quality of 3D localization. Although M3D-RPN does not compete directly with SOTA methods for 2D detection, its performance is suitable to facilitate the tasks in focus such as BEV and 3D detection.

2 Ablations

For all ablations and experimental analysis we use the KITTI val1 dataset split and evaluate utilizing the car class. Further, we use the moderate setting of each task which includes 2D detection, 3D detection, and BEV (Tab. 5).

Depth-aware Convolution: We propose depth-aware convolution as a method to improve the spatial-awareness of high-level features. To better understand the effect of depth-aware convolution, we ablate it from the perspective of the hyperparameter bb which denotes the number of discrete bins. Since our framework uses an image scale of 512512 pixels with network stride of 1616, the output feature map can naturally be separated into 51216=32\frac{512}{16}=32 bins. We therefore ablate using bins of $$ as described in Tab. 5.

We additionally ablate the special case of b=1b=1, which is the equivalent to utilizing two global streams. We observe that both b=1b=1 and b=4b=4 result in generally worse performance than the baseline without local features, suggesting that arbitrarily adding deeper layers is not inherently helpful for 3D localization. However, we observe consistent improvements when b=32b=32 is used, achieving a large gain of 3.71%3.71\% in APBEV{}_{\text{BEV}}, 1.98%1.98\% in AP3D{}_{\text{3D}}, and 1.51%1.51\% in AP2D{}_{\text{2D}}.

We breakdown the learned α\alpha weights after sigmoid which are used to fuse the global and local outputs (Tab. 6). Lower values favor local branch and vice-versa for global. Interestingly, the classification cc output learns the highest bias toward local features, suggesting that semantic features in urban scenes have a moderate reliance on depth position.

Post 3D→\rightarrow2D Optimization: The post-optimization algorithm encourages consistency between 3D boxes projected into the image space and the predicted 2D boxes. We ablate the effectiveness of this optimization as detailed in Tab. 5. We observe that the post-optimization has a significant impact on both BEV and 3D detection performance. Specifically, we observe performance gains of 4.48%4.48\% in APBEV{}_{\text{BEV}} and 4.09%4.09\% in AP3D{}_{\text{3D}}. We additionally observe that the algorithm converges in approximately 88 iterations on average and adds minor 1313 ms overhead (per image) to the runtime.

Efficiency: We emphasize that our approach uses only a single network for inference and hence involves overall more direct 3D predictions than the use of multiple networks and stages (RPN with R-CNN) used in prior works . We note that direct efficiency comparison is difficult due to a lack of reporting in prior work. However, we comprehensively report the efficiency of M3D-RPN for each ablation experiment, where bb and post-optimization are the critical factors, as detailed in Tab. 5. The runtime efficiency is computed using NVIDIA 10801080ti GPU averaged across the KITTI val1 dataset. We note that depth-aware convolution incurs 2−20%2-20\% overhead cost for b=1…32b=1\dots 32, caused by unfolding and reshaping in PyTorch .

Conclusion

In this work, we present a reformulation of monocular image-only 3D object detection using a single-shot 3D RPN, in contrast to prior work which are comprised of external networks, data sources, and involve multiple stages. M3D-RPN is uniquely designed with shared 2D and 3D anchors which leverage strong priors closely linked to the correlation between 2D scale and 3D depth. To help improve 3D parameter estimation, we further propose depth-aware convolution layers which enable the network to develop spatially-aware features. Collectively, we are able to significantly improve the performance on the challenging KITTI dataset on both the Birds Eye View and 3D object detection tasks for the car, pedestrian, and cyclist classes.

Acknowledgment: Research was partially sponsored by the Army Research Office under Grant Number W911NF-18-1-0330. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the Army Research Office or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for Government purposes notwithstanding any copyright notation herein.

References

M3D-RPN: Monocular 3D Region Proposal Network for Object Detection — p7