Categorical Depth Distribution Network for Monocular 3D Object Detection

Cody Reading, Ali Harakeh, Julia Chae, Steven L. Waslander

Introduction

Perception in 3D space is a key component in fields such as autonomous vehicles and robotics, enabling systems to understand their environment and react accordingly. LiDAR and stereo sensors have a long history of use for 3D perception tasks, showing excellent results on 3D object detection benchmarks such as the KITTI 3D object detection benchmark due to their ability to generate precise 3D measurements.

Monocular based 3D perception has been pursued simultaneously, motivated by the potential for a low-cost, easy-to-deploy solution with a single camera . Performance on the same 3D object detection benchmarks lags significantly relative to LiDAR and stereo methods, due to the loss of depth information when scene information is projected onto the image plane.

To combat this effect, monocular object detection methods often learn depth explicitly, by training a monocular depth estimation network in a separate stage. However, depth estimates are consumed directly in the 3D object detection stage without an understanding of depth confidence, leading to networks that tend to be overconfident in depth predictions. Over-confidence in depth is particularly an issue at long range , leading to poor localization. Further, depth estimation is separated from 3D detection during the training phase, preventing depth map estimates from being optimized for the detection task.

Depth information in image data can also be learned implicitly, by directly transforming features from images to 3D space and finally to bird’s-eye-view (BEV) grids . Implicit methods, however, tend to suffer from feature smearing, wherein similar image features can exist at multiple locations in the projected space. Feature smearing increases the difficulty of localizing objects in the scene.

Related Work

Monocular Depth Estimation. Monocular depth estimation is performed by generating a single depth value for every pixel in an image. As such, many monocular depth estimation methods are based on architectures used in well-studied pixel-to-pixel mapping problems such as semantic segmentation. As an example, fully convolutional networks (FCNs) were introduced for semantic segmentation, and were subsequently adopted for monocular depth estimation . The atrous spatial pyramid pooling (ASPP) module was also first proposed for semantic segmentation in DeepLab and subsequently used for depth estimation in DORN and BTS . Further, many methods jointly perform depth estimation and segmentation in an end-to-end manner. We follow the design of the semantic segmentation network DeepLabV3 for estimating categorical depth distributions for each pixel in the image. BEV Semantic Segmentation. BEV segmentation methods attempt to predict BEV semantic maps of 3D scenes from images. Images can be used to either directly estimate BEV semantic maps or to estimate a BEV feature representation as an intermediate step for the segmentation task. In particular, Lift, Splat, Shoot predicts categorical depth distributions in an unsupervised manner, in order to generate intermediate BEV representations. In this work, we predict categorical depth distributions via supervision with ground truth one-hot encodings to generate more accurate depth distributions for object detection.

Monocular 3D Detection. Monocular 3D object detection methods often generate intermediate representations to assist in the 3D detection task. Based on these representations, monocular detection can be divided into three categories: direct, depth-based, and grid-based methods. Direct Methods. Direct methods estimate 3D detections directly from images without predicting an intermediate 3D scene representation. Rather, direct methods can incorporate the geometric relationship between the 2D image plane and 3D space to assist with detections. For example, object keypoints can be estimated on the image plane, in order to assist in 3D box construction using known geometry . M3D-RPN introduces depth-aware convolutions that divides the input row-wise and learns non-shared kernels for each region, to learn location specific features that correlate to regions in 3D space. Shape estimation can be performed for objects in the scene to create an understanding of 3D object geometry. Shape estimates can be supervised from labeled vertices of 3D CAD models , from LiDAR scans , or directly from input data in a self-supervised manner . A drawback for direct methods is that detections are generated directly from 2D images, without access to explicit depth information, usually resulting in reduced performance in localization relative to other methods. Depth-Based Methods. Depth-based methods perform the 3D detection task using pixel-wise depth maps as an additional input, where the depth maps are precomputed using monocular depth estimation architectures . Estimated depth maps can be used in combination with images to perform the 3D detection task . Alternatively, depth maps can be converted to 3D point clouds, commonly known as Pseudo-LiDAR , which are either used directly or combined with image information to generate 3D object detection results. Depth-based methods separate depth estimation from 3D object detection during the training stage, leading to the learning of sub-optimal depth maps used for the 3D detection task. Accurate depth should be prioritized for pixels belonging to objects of interest, and is less important for background pixels, a property that is not captured if depth estimation and object detection are trained independently. Grid-Based Methods. Grid-based methods avoid estimating raw depth values by predicting a BEV grid representation, to be used as input for 3D detection architectures. Specifically, OFT populates a voxel grid by projecting voxels into the image plane and sampling image features, to be transformed into a BEV representation. Multiple voxels can be projected to the same image feature, leading to repeated features along the projection ray and reduced detection accuracy.

CaDDN addresses all identified issues by jointly performing depth estimation and 3D object detection in an end-to-end manner, and leverages the depth estimates to generate meaningful bird’s-eye-view representations with accurate and localized features.

Methodology

CaDDN learns to generate BEV representations from images by projecting image features into 3D space. 3D object detection is then performed with the rich BEV representation using an efficient BEV detection network. An overview of CaDDN’s architecture is shown in Figure 2.

Our network learns to produce BEV representations that are well-suited for the task of 3D object detection. Taking an image as input, we construct a frustum feature grid using the estimated categorical depth distributions. The frustum feature grid is transformed into a voxel grid using known camera calibration parameters, and then collapsed to a bird’s-eye-view feature grid.

Let (u,v,c)(u,v,c) represent a coordinate in image features F\mathbf{F} and (u,v,di)(u,v,d_{i}) represent a coordinate in categorical depth distributions D\mathbf{D}, where (u,v)(u,v) are the feature pixel location, cc is the channel index, and did_{i} is the depth bin index. To generate a frustum feature grid G\mathbf{G}, each feature pixel F(u,v)\mathbf{F}(u,v) is weighted by its associated depth bin probabilities in D(u,v)\mathbf{D}(u,v) to populate the depth axis did_{i}, visualized in Figure 3. Feature pixels can be weighted by depth probability using the outer product, defined as:

2 BEV 3D Object Detection

To perform 3D object detection on the BEV feature grid, we adopt the backbone and detection head of the well-established BEV 3D object detector PointPillars , as it has been shown to provide accurate 3D detection results with a low computational overhead. For the BEV backbone, we increase the number of 3x3 convolution + BatchNorm + ReLU layers in the downsample blocks from (4, 6, 6) used in the original PointPillars to (10, 10, 10) for Block1, Block2, and Block3 respectively. Increasing the number of convolutional layers expands the learning capacity in our BEV network, important for learning from lower quality features produced by images compared to higher quality features originally produced by LiDAR point clouds. We use the same detection head as PointPillars to generate our final detections.

3 Depth Discretization

The continuous depth space is discretized in order to define the set of DD bins used in the depth distributions D\mathbf{D}. Depth discretization can be performed with uniform discretization (UD) with a fixed bin size, spacing-increasing discretization (SID) with increasing bin sizes in log⁡\log space, or linear-increasing discretization (LID) with linearly increasing bin sizes. Depth discretization techniques are visualized in Figure 5.

We adopt LID as our depth discretization as it provides balanced depth estimation for all depths . LID is defined as:

where dcd_{c} is the continuous depth value, [dmin⁡,dmax⁡][d_{\min},d_{\max}] is the full depth range to be discretized, DD is the number of depth bins, and did_{i} is the depth bin index.

4 Depth Distribution Label Generation

We require depth distribution labels D^\mathbf{\hat{D}} in order to supervise our predicted depth distributions. Depth distribution labels are generated by projecting LiDAR point clouds into the image frame to create sparse dense maps. Depth completion is performed to generate depth values at each pixel in the image. We require depth information at each image feature pixel, so we downsample the depth maps of size WI×HIW_{I}\times H_{I} to the image feature size WF×HFW_{F}\times H_{F}. The depth maps are converted to bin indices using the LID discretization method described in Section 3.3, followed by a conversion into a one-hot encoding to generate the depth distribution labels. A one-hot encoding ensures the depth distribution labels are sharp, essential to encourage sharpness in our depth distribution predictions via supervision.

5 Training Losses

Generally, classification is performed by predicting categorical distributions, and encouraging sharpness in the distribution in order to select the correct class . We leverage classification to encourage a single correct depth bin when supervising the depth distribution network, using the focal loss :

Experimental Results

To demonstrate the effectiveness of CaDDN we present results on both the KITTI 3D object detection benchmark and the Waymo Open Dataset .

The KITTI 3D object detection benchmark is divided into 7,481 training samples and 7,518 testing samples. The training samples are commonly divided into a train set (3,712 samples) and a val set (3,769 samples) following , which is also adopted here. We compare CaDDN with existing methods on the test set by training our model on both the train and val sets. We evaluate on the val set for ablation by training our model on only the train set.

2 Waymo Dataset Results

To the best of our knowledge, no monocular methods have reported results on Waymo. In order to provide a baseline, we extend the official implementation of M3D-RPN to support the Waymo Open Dataset . Table 2 shows the results of both the M3D-RPN baseline and CaDDN on the Waymo validation set. Our method significantly outperforms M3D-RPN with margins on AP/APH of +4.69%/+4.65% and +4.15%/+4.12% on the LEVEL_1 and LEVEL_2 difficulties respectively for an IoU criteria of 0.7.

3 Ablation Studies

4 Depth Distribution Uncertainty

To validate that our depth distributions contain meaningful uncertainty information, we compute the Shanon entropy for each estimated categorical depth distribution in D\mathbf{D}. We label each distribution with its associated ground truth depth bin and foreground/background classification. For each group, we compute the entropy statistics which are shown in Figure 6. We observe that entropy generally increases as a function of depth, where depth estimates are challenging, indicating our distributions describe meaningful uncertainty information. Our network produces the lowest distribution entropy at pixels with ground truth depth of around 6 meters. We attribute the high entropy at depths closer than 6 meters to the small number of pixels at shorter ranges in the training set. Finally, we note that the foreground depth distribution estimates have slightly higher entropy than background pixels, a phenomenon that can also be attributed to training set imbalance.

Conclusion

We have presented CaDDN, a novel monocular 3D object detection method that estimates accurate categorical depth distributions for each pixel. The depth distributions are combined with the image features to generate bird’s-eye-view representations that retain depth confidence, to be exploited for 3D object detection. We have shown that estimating sharp categorical distributions centered around the correct depth value, and jointly performing depth estimation and object detection is vital for 3D object detection performance, leading to a \nth1 place ranking on the KITTI dataset among all published methods at the time of submission.

References

Appendix A Additional Results

A.2 Ablation Studies

Depth Discretization. Table 6 shows the detection peformance of CaDDN with each of the depth discretization methods outlined in Section 3.3. We observe that LID offers the highest performance, leading to its use for our method. Feature Resolution. Table 7 shows the detection peformance of CaDDN when modifying the image feature extraction layer in the Image Backbone (See Figure 2). Experiment 7 shows the performance when image features are extracted from Block1. Experiments 7, 7, and 7 shows reduced performance when smaller resolution image features are extracted. Smaller spatial resolutions in the image features cause oversampling in the frustum to voxel grid transformation, leading to many voxel features V\mathbf{V} with similar features (See Section 3.1).

Appendix B Additional Details