Center3D: Center-based Monocular 3D Object Detection with Joint Depth Understanding

Yunlei Tang, Sebastian Dorn, Chiragkumar Savani

Introduction

3D object detection is currently one of the most challenging topics for both industry and academia. Applications of related developments can easily be found in the areas of robotics, AI-based medical surgery, autonomous driving ,, etc. The goal is to have agents with the ability to identify, localize, react, and interact with objects in their surroundings. 2D object detection approaches ,, ,, achieved impressive results in the last decade. In contrast, inferring associated 3D properties from a 2D image turned out to be a challenging problem in computer vision, due to the intrinsic scale ambiguity of 2D objects and the lack of depth information. Hence many approaches involve additional sensors like LiDAR ,, or radar , to measure depth. Modern LiDAR sensors have proven themselves a reliable choice with high depth accuracy. On the other hand, there are reasons to prefer monocular-based approaches too. LiDAR has reduced range in adverse weather conditions, while visual information of a simple RGB camera is more dense and also more robust under rain, snow, etc. Another reason is that cameras are currently significantly more economical than high precision LiDARs and are already available in e.g. robots, vehicles, etc. Additionally the processing of single RGB images is much more efficient and faster than processing 3D point clouds in terms of CPU and memory utilization.

These compelling reasons have led to research exploring the possibility of 3D detection solely from monocular images ,,,,,,,. The network structure of most 3D detectors starts with a 2D region proposal based (also called anchor based) approach, which enumerates an exhaustive set of predefined proposals over the image plane and classifies/regresses only within the region of interest (ROI). Mousavian et al. uses a two-stage 2D detector and adds specific heads for regression of 3D properties. The resulting 3D cuboid is then fine tuned to ensure that it tightly fits inside the associated 2D bounding box. GS3D modifies Faster RCNN to propose a 3D guidance based on a 2D bounding box and orientation, which guides to frame a 3D cuboid by refinement. OFT-Net maps 2D image features into an orthographic bird’s-eye view (BEV) by an orthographic feature transformation, and infers the depth in a reprojected 3D space. Mono3D focuses on the generation of proposed 3D boxes, which are scored by features like contour and shape under the assumption that all vehicles are placed on the ground plane. MonoGRNet consists of parameter-specific subnetworks. All further regressions are guided by the detected 2D bounding box. M3D-RPN demonstrates a single-shot model with a standalone 3D RPN, which generates 2D and 3D proposals simultaneously. Additionally, the specific design of depth-aware convolutional layers improved the network’s 3D understanding. With the help of an external network, Multi-Fusion estimates a disparity map and subsequently a LiDAR point cloud to improve 3D detection. Due to multi-stage or anchor-based pipelines, most of them perform slowly.

Most recently, to overcome the disadvantages above, 2D anchor-free approaches have been used by researchers ,,,. They model objects with keypoints like centers, corners or points of interest of 2D bounding boxes. Anchor-free approaches are usually one-stage, thus eliminating the complexity of designing a set of anchor boxes and fine tuning hyperparameters. The recent work of CenterNet: Objects as Points proposed a possibility to associate a 2D anchor free approach with a 3D detection.

Nevertheless, the performance of CenterNet is still restricted by the fact that a 2D bounding box and a 3D cuboid are sharing the same center point. In this paper we analyze the difference between the center points of 2D bounding boxes and the projected 3D center points of objects, which are almost never at the same image position. We directly regress the 3D centers from 2D centers to locate the objects in the image plane and in 3D space separately. Furthermore, we show that the weakness of monocular 3D detection is mainly caused by imprecise depth estimation. This causes the performance gap between LiDAR-based and monocular image-based approaches. By examining depth estimation in monocular images, we show that a combination of classification and regression explores visual clues better than using only a single approach. An overview of our approach is shown in Figure 1.

We introduce two approaches to validate this conclusion: (1) Motivated by DORN we consider depth estimation as a sequential classification with residual regression. According to the statistics of the instances in the KITTI dataset, a novel discretization strategy is used. (2) We divide the whole depth range of objects into two bins, foreground and background, either with overlap or associated. Classifiers indicate which depth bin or bins the object belongs to. With the help of Eigen’s transformation , two regressors are trained to gather specific features for closer and farther away objects, respectively. For illustration see the depth part in Figure 1.

Compared to CenterNet, our approach improved the AP of easy, moderate, hard objects in BEV from 31.531.5, 29.729.7, 28.128.1 to 55.8\mathbf{55.8}, 42.8\mathbf{42.8}, 36.6\mathbf{36.6}, in 3D space from 19.519.5, 18.618.6, 16.616.6 to 49.1\mathbf{49.1}, 38.9\mathbf{38.9}, 33.5\mathbf{33.5}, which is comparable with state-of-the-art approaches. Center3D achieves the best speed-accuracy trade-off on the KITTI dataset in the field of monocular 3D object detection. Details are given in Table 1.

Center3D

The 3D detection approach of CenterNet described in is the basis of our work. It models an object as a single point: the center of its 2D bounding box. For each input monocular RGB image, the original network produces a heatmap for each category, which is trained with focal loss . The heatmap describes a confidence score for each location, the peaks in this heatmap thus represent the possible keypoints of objects. All other properties are then regressed and captured directly at the center locations on the feature maps respectively. For generating a complete 2D bounding box, in addition to width and height, a local offset will be regressed to capture the quantization error of the center point caused by the output stride. For 3D detection and localization, the additional abstract parameters, i.e. depth, 3D dimensions and orientation, will be estimated separately by adding a head for each of them. However, the reconstruction of the lacking space information is ill-posed and challenging because of the inherent ambiguity ,. Following the output transformation of Eigen et al. for depth estimation, CenterNet converts the feature output into an exponential area to suppress the depth space.

In Table 1, we show the reproduced evaluation results of CenterNet on the KITTI dataset, which splits all annotated objects into easy, moderate and hard targets. The results are extended by the approaches introduced in Section 2.2 and 2.3. As CenterNet, we also only focus on the performance in vehicle detection according to standard training and validation splits in literature . To numerically compare our results with other approaches we use intersection over union (IoU) based on 2D bounding boxes (AP), orientation (AOP), and bounding boxes in Bird’s-eye view (BEV AP). From the first row in Table 1, we see that the 2D performance of CenterNet is very good. In contrast, the APs in BEV and especially in 3D perform poorly. This is caused by the difference between the center point of the visible 2D bounding box in the image and the projected center point of the complete object from physical 3D space. This is illustrated in the first two rows of Figure 2. Here a qualitative comparison and demonstration of the influences, caused by the shift of center locations during different driving scenarios, is shown. A center point of the 2D bounding box for training and inference is enough for detecting and decoding 2D properties, e.g. width and height, while all additionally regressed 3D properties, e.g. depth, dimension and orientation, should be consistently decoded from the projected 3D center of the object. The gap between 2D and 3D decreases for faraway objects and for objects which appear in the center area of the image plane. However the gap between 2D and 3D center points becomes significant for objects that are close to the camera or on the image boundary. Due to perspective projection, this offset will increase as vehicles get closer. Close objects are especially important for technical functions based on perception (e.g. in autonomous driving or robotics).

2 Enriching Depth Information

This section first introduces two novel approaches to infer depth cues over monocular images: First, we adapt the advanced DORN approach from pixel-wise to instance-wise depth estimation. We introduce a novel linear-increasing discretization (LID) strategy to divide continuous depth values into discrete ones, which distributes the bin sizes more evenly than spacing-increasing discretization (SID) in DORN. Additionally, we employ a residual regression for refinement of both discretization strategies. Second, with the help of a reference area (RA) we describe the depth estimation as a joint task of classification and regression (DepJoint) in exponential range.

Usually a faraway object with higher depth value and less visible features will induce a higher loss, which could dominate the training and increases uncertainty. On the other hand these targets are usually less important for functions based on object detection. This is also the motivation for the SID strategy to discretize the given continuous depth interval [dmin,dmax][d_{\text{min}},d_{\text{max}}] in log space and hence down-weight farther away objects, see Eq. 1. However, such a discretization often yields too dense bins within unnecessarily close range, where objects barely appear (as shown in Figure 3 first row). According to the histogram in Figure 4 most instances of the KITTI dataset are between 5 m5\ m and 80 m80\ m. Assuming that we discretize the range between dmin=1 md_{\text{min}}=1\ m and dmax=91 md_{\text{max}}=91\ m into N=80N=80 sub-intervals, 2929 bins will be involved within just 5 m5\ m. Thus, we use the LID strategy to ensure the lengths of neighboring bins increase linearly instead of log-wise. For this purpose, assume the length of the first bin is δ\delta. Then the length of the next bin is always δ\delta longer than the previous bin. Now we can encode an instance depth dd in lint=⌊l⌋l_{\text{int}}=\lfloor l\rfloor ordinal bins according to LID and SID respectively. Additionally, we reserve and regress the residual decimal part lres=l−lintl_{\text{res}}=l-l_{\text{int}} for both discretization strategies:

where Pni\mathcal{P}_{n}^{i} is the probability that the ii-th instance is farther away than the nn-th bin, and SmL1SmL1 represents the smooth L1 loss function . During inference, the amount of activated bins will be counted up as l^inti\hat{l}_{\text{int}}^{i}. We refine the result by taking into account the residual part, l^=l^inti+l^resi\hat{l}=\hat{l}_{\text{int}}^{i}+\hat{l}_{\text{res}}^{i}, and decode the result by inverse-transformation of Equation 1.

2.2 DepJoint

where d^bi\hat{d}_{b}^{i} represents the regression output for the bb-th bin and ii-th instance. Training is only applied on 2D centers of bounding boxes. During inference the weighted average will be decoded as the final result:

where PBin biP_{\text{Bin b}}^{i} denotes the normalized probability of d^i\hat{d}^{i}.

3 Offset3D: Bridging 2D to 3D

As described in Section 2.1, the significant difference in performance between 2D and 3D object localization results from the gap between the centers of 2D bounding boxes c2Di=(x2Di,y2Di)\mathbf{c_{\text{2D}}^{i}}=(x_{\text{2D}}^{i},y_{\text{2D}}^{i}) and the 3D projected center points of cuboids from physical space c3Di=(x3Di,y3Di)\mathbf{c_{\text{3D}}^{i}}=(x_{\text{3D}}^{i},y_{\text{3D}}^{i}). Instinctively we would like to anchor the objects by projected 3D center points c3Di\mathbf{c_{\text{3D}}^{i}} instead of c2Di\mathbf{c_{\text{2D}}^{i}}. Rather than using a regression of width and height (x2i−x1i,y2i−y1i)(x_{2}^{i}-x_{1}^{i},y_{2}^{i}-y_{1}^{i}) for the 2D task, we can regress the distances from four boundaries to the 3D centers as (x3Di−x1i,x2i−x3Di,y3Di−y1i,y2i−y3Di)(x_{\text{3D}}^{i}-x_{1}^{i},x_{2}^{i}-x_{\text{3D}}^{i},y_{\text{3D}}^{i}-y_{1}^{i},y_{2}^{i}-y_{\text{3D}}^{i}) to decode the additional 3D parameters properly. This strategy has some natural limits, though. First, the detector can not locate objects very close to the camera with 3D center points outside the image (e.g. second column in Figure 2). Secondly, there is an ambiguity in the distance with respect to width and height. Hence we split the 2D and 3D tasks into separate parts. We still locate an object with 2D center c2Di=(x2Di,y2Di)\mathbf{c_{\text{2D}}^{i}}=(x_{\text{2D}}^{i},y_{\text{2D}}^{i}), which is definitively included in the image, and determine the 2D bounding box of the visible part with wiw^{i} and hih^{i}. For the 3D task we relocate the projected 3D center c3Di=(x3Di,y3Di)\mathbf{c_{\text{3D}}^{i}}=(x_{\text{3D}}^{i},y_{\text{3D}}^{i}) by adding two head layers on top of the backbone and regress the offset Δci=(x3Di−x2Di,y3Di−y2Di)\mathbf{\Delta c^{i}}=(x_{\text{3D}}^{i}-x_{\text{2D}}^{i},y_{\text{3D}}^{i}-y_{\text{2D}}^{i}) from 2D to 3D centers. Given the projection matrix P\mathbf{P} in KITTI, we can now determine the 3D location C=(X,Y,Z)\mathbf{C}=(X,Y,Z) by converting the transformation in homogeneous coordinates. Similarly, we generate 8 corners of a cuboid by decoding object dimensions based on the 3D location, which is actually the center of the object in the word coordinate system.

4 Reference Area

Conventionally the regressed values of a single instance will be trained and accessed only on a single center point, which reduces the calculation. However, it also restricts the perception field of regression and affects reliability. To overcome these disadvantages, we apply the approach used by Eskil et al. and Krishna et al. . Instead of relying on a single point, a Reference Area (RA) based on the 2D center point is defined within the 2D bounding box, whose width and height are set accordingly with a proportional value γ\gamma. All values within this area contribute to regression and classification. If RAs overlap, the area closest to the camera dominates, since only the closest instance is completely visible on the monocular image. An example using RAs is shown in Figure 6. During inference all predictions in the related RA will be weighted equally.

Experiments

We performed experiments on the KITTI object detection benchmark , which contains 7481 training images with annotation and calibration. KITTI was chosen from the public available datasets, e.g. ,,,, since it is the widely used benchmark by previous works for monocular 3D detection, including our baseline, CenterNet. All instances are divided into easy, moderate and hard targets according to visibility in the image . We follow the standard training/validation split strategy in literature , which leads to 3712 images for training and 3769 images for validation. Like most previous work, and in particular CenterNet, we only consider the “Car” category. By default, KITTI evaluates 2D and 3D performance with AP at eleven recalls from 0.00.0 to 1.01.0 at different IoU thresholds. For a fair comparison with CenterNet, parameters stay unchanged if not stated otherwise. In particular we keep the modified Deep Layer Aggregation (DLA)-34 as the backbone. Regarding different approaches, we add specific head layers, which consist of one 3×33\times 3 convolutional layer with 256 channels, ReLu activation and a 1×11\times 1 convolution with desired output channels at the end. We trained the network from scratch in PyTorch on 2 GPUs (1080Ti) with batch sizes 7 and 9. We trained the network for 70 epochs with an initial learning rate of 1.25e−41.25e^{-4}, which drops by a factor of 10 at 45 and 60 epochs if not specified otherwise.

2 Offset3D

With Offset3D we bridge the gap between 2D and 3D center points by adding 2 specific layers to regress the offset Δci\mathbf{\Delta c}^{i}. For demonstration we perform an experiment, which is indicated as CenterNet(ct3d) in Table 1. It models the object with a projected 3D center point with 4 distances to boundaries. The visible object, whose 3D center point is out of the image, is ignored during training. As the second row in Table 1 shows, for easy targets CenterNet(ct3d) increases the BEV AP by 48.6%48.6\% and the 3D AP by 104.6%104.6\% compared to the baseline of CenterNet. This is achieved by the proper decoding of 3D parameters based on an appropriate 3D center point. However, as discussed in Section 2.3, simply modeling an object with a 3D center will hurt 2D performance, since some 3D centers are not attainable, although the object is still partly visible in the image.

The Offset3D approach is able to balance the trade-off between a tightly fitting 2D bounding box and a proper 3D location. Center3D is showing that the regression of offsets is regularizing the spatial center, while also preserving the stable precision in 2D (the 7th row in Table 1). With a higher learning rate of 2.4e−42.4e^{-4}, BEV AP for easy targets improves from 31.5%31.5\% to 50.7%50.7\%, and 3D AP increases from 19.5%19.5\% to 43.1%43.1\%, which performs comparably to the state of the art. Since Offset3D is also the basis for all further experiments, we treat the performance as our new baseline for comparison.

3 LID

Table 2 shows both LID and SID with different number of bins improved even instance-wise with additional layers for ordinal classification, when a proper number of bins is used. Our discretization strategy LID shows a considerably higher precision in 3D evaluation, comprehensively when 80 and 100 bins are used. A visualization of inferences of both approaches from BEV is shown in Figure 3. LID only preforms worse than SID in the 40 bin case, where the number of intervals is not enough for instance-wise depth estimation. Furthermore we verify the necessity of the regression of residuals by comparing the last two rows in Table 2. The performance of LID in 3D will be deteriorate drastically if this refinement module is removed..

4 DepJoint and Reference Area

In this section, we evaluate the performance of DepJoint and RA, and finally discuss the compatibility of both approaches.

The RA enriches the cues for inference and utilizes more visible surface features of instances. Focusing especially on 3D performance, we evaluate RA over depth estimation and explore the sensitivity to the size of the RA. Table 3 shows the improvement of models supported by the RA for 3D detection. As we can see in Table 3, performance does not simply improve with higher proportion scale, γ\gamma, since the best AP is achieved with γ=0.4\gamma=0.4 . We attribute this characteristic to the misalignment of estimated 2D bounding boxes and the position of the RA. The output features of depth in the RA respond exactly to the input pixels on their location. However, the localization of the RA during inference is not the same as in ground truth. As a result, a shift of the estimated location or the size of the RA involves the feature values, which in fact belong to another instance, especially when objects overlap. A weighted average during decoding could reduce but generally not eliminate this error. Furthermore, we also determine the effectiveness of RA over the offset prediction ΔC\mathbf{\Delta C}. Therefore, the choice of loss weight to balance the decreasing losses between tasks is tricky. Finally λoff=0.025\lambda_{\text{off}}=0.025 performs best. A smaller value of λoff\lambda_{\text{off}} downweights the magnitude of the offset loss and does not influence the training of other losses. Instead, it will be fine tuned in late period.

4.2 DepJoint

In the following sections we combine the DepJoint approach with the concept of RA. As discussed above we set γ=0.4\gamma=0.4 as default for RA. Additionally, we apply dmin=0d_{\text{min}}=0 and dmax=60d_{\text{max}}=60 for all experiments. Table 4 shows that DepJoint is more robust against loss weighting during training in comparison with Eigen’s approach. As introduced in Section 2.2.2, we can divide the whole depth range into two overlapping or associated bins. For the associated strategy, the thresholds α/β\alpha/\beta of 0.3/0.30.3/0.3 and 0.4/0.40.4/0.4 show the best performance for the following reason: usually more distant objects show less visible features in the image. Hence, we want to set both thresholds a little lower, to assign more instances to the distant bins and thereby suppress the imbalance of visual clues between the two bins.

5 Comparison to the State of the Art

Table 1 shows the comparison with state-of-the-art monocular 3D detectors on the KITTI dataset. The LID approach follows the settings described in Section 3.3. Different from the experiments above, all RAs are applied for both depth dd and offset ΔC\Delta C estimation (γ=0.4\gamma=0.4). Our first model, Center3D(+ra), combines the reference area approach with Eigen’s transformation for depth estimation. The loss weights are set to λdep=1\lambda_{\text{dep}}=1 and λoffset=0.025\lambda_{\text{offset}}=0.025, while the learning rate is 2.4e−42.4e^{-4}. Finally, we evaluated the performance of a model employing both DepJoint and RA approaches, Center3D(+dj+ra). The training with λdep=1\lambda_{\text{dep}}=1, λoffset=0.1\lambda_{\text{offset}}=0.1 and learning rate 1.25e−41.25e^{-4} yields the best experimental result.

As Table 1 shows, all Center3D models perform at comparable 3D performance with respect to the best approaches currently available. Center3D(+dj+ra) achieved state-of-the-art performance with BEV APs of 55.8%55.8\% and 42.8%42.8\% for easy and moderate targets, respectively. For hard objects in particular, Center3D(+ra) outperforms all other approaches. The RA improved the BEV AP by 4.9%4.9\% (from 35.3%35.3\% to 40.2%40.2\%) and 1.0%1.0\% (from 33.0%33.0\% to 34.0%34.0\%).

Center3D preserves the advantages of an anchor-free approach. It performs better than most other approaches in 2D AP, especially on easy objects. Most importantly, it infers on the monocular input image with the highest speed (around three times faster than M3D-RPN, which performs similarly to Center3D in 3D). Therefore, Center3D is able to fulfill the requirement of a real-time detection. Additionally, it achieved the best trade-off between speed and performances not only in 2D but also in 3D.

Conclusion

In this paper we introduced Center3D, a one-stage anchor-free monocular 3D object detector, which models and detects objects with center points of 2D bounding boxes. We recognize and highlight the importance of the difference between centers in 2D and 3D by regressing the offset directly, which transforms 2D centers to 3D centers. In order to improve depth estimation, we further explored the effectiveness of joint classification and regression when only monocular images are given. Both classification-dominated (LID) and regression-dominated (DepJoint) approaches enhance the AP in BEV and 3D space. Finally, we employed the concept of RAs by regressing in predefined areas, to overcome the sparsity of the feature map in anchor-free approaches. Center3D performs comparably to state-of-the-art monocular 3D approaches with significantly improved runtime during inference. Center3D achieved the best trade-off between 3D precision and inference speed.

Acknowledgement

We gratefully acknowledge Heinz Koeppl for his support of this work as well as Christian Wildner for his advice. Furthermore, we thank Andrew Chung and Mentar Mahmudi for helpful comments and discussions.

References