Is Pseudo-Lidar needed for Monocular 3D Object detection?
Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li, Adrien Gaidon
Introduction
Detecting and accurately localizing objects in 3D space is crucial for many applications, including robotics, autonomous driving, and augmented reality. Hence, monocular 3D detection is an active research area , owing to its potentially wide-ranging impact and the ubiquity of cameras. Leveraging exciting recent progress in depth estimation , pseudo-lidar detectors first use a pre-trained depth network to compute an intermediate pointcloud representation, which is then fed to a 3D detection network. The strength of pseudo-lidar methods is that they monotonically improve with depth estimation quality, e.g., thanks to large scale training of the depth network on raw data.
However, regressing depth from single images is inherently an ill-posed inverse problem. Consequently, errors in depth estimation account for the major part of the gap between pseudo-lidar and lidar-based detectors, a problem compounded by generalization issues that are not fully understood yet . Simpler end-to-end monocular 3D detectors seem like a promising alternative, although they do not enjoy the same scalability benefits of unsupervised pre-training due to their single stage nature.
In this work, we aim to get the best of both worlds: the scalability of pseudo-lidar with raw data and the simplicity and generalization performance of end-to-end 3d detectors. To this end, our main contribution is a novel fully convolutional single-stage 3D detection architecture, DD3D (for Dense Depth-pre-trained 3D Detector), that can effectively leverage monocular depth estimation for pre-training (see Figure 1). Using a large corpus of unlabeled raw data, we show that DD3D scales similarly to pseudo-lidar methods, and that depth pre-training improves upon pre-training on large labeled 2D detection datasets like COCO , even with the same amount of data.
Our method sets a new state of the art on the task of monocular 3D detection on KITTI-3D and nuScenes with significant improvements compared to previous state-of-the-art methods. The simplicity of its training procedure and its end-to-end optimization allows for effective use of large-scale depth data, leading to impressive multi-class detection accuracy.
Related work
Monocular 3D detection. A large body of work in image-based 3D detection builds upon 2D detectors, and aims to lift them to 3D using various cues from object shapes and scene geometry. These priors are often injected by aligning 2D keypoints with their projective 3D counterparts , or by learning a low-dimensional shape representation . Other methods try to leverage geometric consistency between 2D and 3D structures, formulating the inference as a constrained optimization problem . These methods commonly use additional data (e.g. 3D CAD models, instance segmentation), or assume rigidity of objects. Inspired by Lidar-based methods , another line of work employs view transformation (i.e. birds-eye-view) to overcome the limitations of perspective range view, e.g. occlusion or variability in size . These methods often require precise camera extrinsics, or are accurate only in close vicinity.
End-to-end 3D detectors. Alternatively, researchers have attempted to directly regress 3D bounding boxes from CNN features . Typically, these approaches extend standard 2D detectors (single-stage or two-stage ) by adding heads that predict various 3D cuboid parameterizations . The use of depth-aware convolution or dense 3D anchors enabled higher accuracy. Incorporating uncertainty in estimating the 3D box (or its depth) has also been shown to greatly improve detection accuracy . In , the authors proposed a disentangled loss for 3D box regression that helps stabilize training. DD3D also falls in the end-to-end 3D detector category, however, our emphasis is on learning a good depth representation via large-scale self-supervised pre-training on raw data, which leads to robust 3D detection.
Pseudo-Lidar methods. Starting with the pioneering work of , these methods leverage advances in monocular depth estimation and train Lidar-based detectors on the resulting pseudo pointcloud, producing impressive results . Recent methods have improved upon by correcting the monocular depth with sparse Lidar readings , decorating the pseudo pointcloud with colors , using 2D detection to segment foreground pointcloud regions , and structured sparsification of the monocular pointcloud . Recently, has shown a bias in the PL results on the representative KITTI-3D benchmark, while this paradigm remains the state-of-the-art. Multiple researchers showed that inaccurate depth estimation is a major source of error in PL methods . In this work, we build our reference PL method based on and to investigate the benefits of using large-scale depth data and compare it with DD3D.
Monocular depth estimation. Estimating per-pixel depth is a key task for both DD3D (as a pre-training task) and PL (as the first of a two-stage pipeline). It is itself a thriving research area: the community has pushed toward accurate dense depth prediction via supervised as well as self-supervised methods. We note that supervised monocular depth training used in this work require no annotations from human, allowing us to scale our methods to a large amount of raw data.
Dense depth pre-training for 3D detection
Given a single image and its camera intrinsics matrix as input, the goal of monocular 3D detection is to generate a set of multi-class 3D bounding boxes relative to camera coordinates. During inference, DD3D does not require any additional data, such as per-pixel depth estimates, 2D bounding boxes, segmentations, or 3D CAD models. DD3D is also camera-aware: it scales the depth of putative 3D boxes according to the camera intrinsics.
DD3D is a fully convolutional single-stage network that extends FCOS to perform 3D detection and dense depth prediction. The architecture (see Figure 2) is composed of a backbone network and three sub-networks (or heads) that are shared among all multi-scale features. The backbone takes an RGB image as input, and computes convolutional features at different scales. As in , we adopt a feature pyramid network (FPN) as the backbone.
Three head networks are applied to each feature map produced by the backbone, and perform independent prediction tasks. The classification module predicts object category. It produces real values, where is the number of object categories. The 2D box module produces class-agnostic bounding boxes and center-ness by predicting offsets from the feature location to the sides of each bounding box and a scalar associated with center-ness. We refer the readers to for more details regarding the 2D detection architecture.
3D detection head. This head predicts 3D bounding boxes and per-pixel depth. It takes FPN features as input, and applies four 2D convolutions with kernels that generate real values for each feature location. These are decoded into 3D bounding box, per-pixel depth map, and 3D prediction confidence, as described below:
is the quaternion representation of allocentric orientation of the 3D bounding box. It is normalized and transformed to an egocentric orientation . Note that we predict orientations with the full 3 degrees of freedom.
represent the depth predictions. is decoded to the z-component of 3D bounding box centers and therefore only associated with foreground features, while is decoded to the monocular depth to the closest surface and is associated with every pixel. To decode them to metric depth, we unnormalized them using per-level parameters as follows:
where are network output, are predicted depths, are learnable scaling factors and offsets defined for each FPN level, is the pixel size computed from the focal lengths, and , and is a constant.
We note that this design of using camera focal lengths endows DD3D with camera-awareness, by allowing us to infer the depth not only from the input image, but also from the pixel size. We found that this is particularly useful for stable training. Specifically, when the input image is resized during training, we keep the ground-truth 3D bounding box unchanged, but modify the camera resolution as follows:
where and are resize rates, and is the new camera intrinsic matrix that is used in Eqs. 2 and 4. Finally, collectively represent low-resolution versions of the dense depth maps computed from each FPN features. To recover the full resolution of the dense depth maps, we apply bilinear interpolation to match the size of the input image.
represent offsets from the feature location to the 3D bounding box center projected onto the camera plane. This is decoded to the 3D center via unprojection:
where is the feature location in image space, is a learnable scaling factor assigned to each FPN level.
represents the deviation in the size of the 3D bounding box from the class-specific canonical size, i.e. . As in , is the canonical box size for each class, and is pre-computed from the training data as its average size.
represents the confidence of the 3D bounding box prediction . It is transformed into a probability: , and multiplied by the class probability computed from the classification head to account for the relative confidence to the 2D confidence. The adjusted probability is used as the final score for the candidate.
2 Losses
We adopt the classification loss and 2D bounding box regresssion loss from FCOS :
where the 2D box regression loss is the IOU loss , the classification loss is the binary focal loss (i.e. one-vs-all), and the center-ness loss is the binary cross entropy loss. For 3D bounding box regression, we use the disentangled L1 loss described in , i.e.
where and are the vertices of ground-truth and candidate 3D boxes. This loss is replicated times by using only one of the predicted 3D box components (orientation, projected center, depth, and size), while replacing other three with their ground-truth values.
Also as in , we adopt the self-supervised loss for 3D confidence which uses the error in 3D box prediction to compute a surrogate target for 3D confidence (relative to the 2D confidence):
where is the temperature parameter. The confidence loss is the binary cross entropy between and . In summary, the total loss of DD3D is defined as follows:
3 Depth pre-training
During pre-training, we use per-pixel depth predictions from all FPN levels, i.e. . We consider pixels that have valid ground-truth depth from the sparse Lidar pointclouds projected onto the camera plane, and compute L1 distance from the predicted values:
where is the ground-truth depth map, is the predicted depth map from the -th level in FPN (i.e. interpolated ), and is the binary indicator for valid pixels. We observed using all FPN levels in the objective, rather than e.g. using only the highest resolution features, enables stable training, especially when training from scratch. We also observed that the L1 loss yields stable training with large batch sizes and high-resolution input, compared to SILog loss that is popular in monocular depth estimation literature.
We note that the two paths in DD3D from the input image to the 3D bounding box and to the dense depth prediction differ only in the last convolutional layer, and thus share nearly all parameters. This allows for effective transfer from the pre-trained representation to the target task.
While pre-training, the camera-awareness of DD3D allows us to use camera intrinsics that are substantially different from the ones of the target domain, while still enjoying effective transfer. Specifically, from Eqs. 1 and 2, the error in depth prediction, , caused by the difference in resolution between the two domains is corrected by the difference in pixel size, .
Pseudo-Lidar 3D detection
An alternative way to utilize a large set of image-Lidar frames is to adopt the Pseudo-Lidar (PL) paradigm and aim to improve the depth network component using the large scale data. PL is a two-stage method: first, given an image it applies a monocular depth network to predict per-pixel depth. The dense depth map is transformed into a 3D point cloud, and then a Lidar-based 3D detector is used to predict 3D bounding boxes. The modularity of PL methods enables us to quantify the role of improved depth predictors brought by a large-scale image-LiDAR dataset (see the supplementary material for additional details.)
Monocular depth estimation. The aim of monocular depth estimation is to compute the depth for each pixel . Similarly to Eq. 9, given the ground truth depth measurement acquired from Lidar pointclouds, we define a loss by the error between the predicted and the ground-truth depth. Here we instead use the SILog loss , which yields better performance than L1 for PackNet.
Network architecture. As the depth network, we use PackNet , a state-of-the-art monocular depth prediction architecture which uses packing and unpacking blocks with 3D convolutions. By avoiding feature sub-sampling, PackNet recovers fine structures in the depth map with high precision; moreover, PackNet has been shown to generalize better thanks to its increased capacity .
3D detection. To predict 3D bounding boxes from the input image and the estimated depth map, we follow the method proposed by . We first convert the estimated depth map into a 3D pointcloud similarly to Eq. 4, and concatenate each 3D point with the corresponding pixel values. This results in a 6D tensor encompassing colors along with 3D coordinates. We use an off-the-shelf 2D detector to identify proposal regions in input images, and apply a 3D detection network to each RoI region of the 6-channel image to produce 3D bounding boxes.
Backbone, detection head and 3D confidence. We follow and process each RoI with a ResNet-18 backbone that uses Squeeze-and-Excitation layers . As the RoI contains both objects as well as background pixels, the resulting features are filtered via foreground masks computed based on the associated RoI depth map . The detection head follows and operates in 3 distance ranges, producing one bounding box for each range. The final output is then selected based on the mean depth of the input RoI. Following , we modify the detection head to also output a 3D confidence value per detection, which is linked to the 3D detection loss.
Loss function. The 3D regression loss is defined as:
In addition, we define a loss that links the predicted 3D confidence with the 3D bounding box coordinates loss using a Binary Cross Entropy (BCE) formulation with target . The final PL 3D detection loss is:
Experimental Setup
KITTI-3D. The KITTI-3D detection benchmark consists of urban driving scenes with object classes. The benchmark evaluates 3D detection accuracy on three classes (Car, Pedestrian, and Cyclist) using two average precision (AP) metrics computed with class-specific thresholds on intersection-over-union (IoU) of 3D bounding boxes or Bird-Eye-View (2D) bounding boxes. We refer to these metrics as 3D AP and BEV AP. We use the revised AP metrics . The training set consists of images, the test set of images. The objects in the test set are organized into three partitions according to their difficulty level (easy, moderate, hard), and are evaluated separately. For the analysis in Section 6.2, we follow the common practice of splitting the training set into and images , and report validation results on the latter. We refer to these splits as KITTI-3D train and KITTI-3D val.
nuScenes. The nuScenes 3D detection benchmark consists of multi-modal videos with cameras covering the full -degree field of view. The videos are split into for training, for validation, and for testing. The benchmark requires to report 3D bounding boxes of object classes over 2Hz-sampled video frames. The evaluation metric, nuScenes detection score (NDS), is computed by combining the detection accuracy (mAP) computed over four different thresholds on center distance with five true-positive metrics. We report NDS and mAP, along with the three true-positive metrics that concern 3D detection, i.e. ATE, ASE, and AOE.
KITTI-Depth. We use the KITTI-Depth dataset to fine-tune the depth networks of our PL models. It contains over thousands depth maps associated with the images in the KITTI raw dataset. The standard monocular depth protocol is to use the Eigen splits . However, as described in , up to a third of its training images overlap with KITTI-3D images, leading to biased results for PL models. To avoid this bias, we generate a new split by removing training images that are geographically close (i.e. within m) to any of the KITTI-3D images. We denote this split by Eigen-clean and use it to fine-tune the depth networks of our PL models.
DDAD15M. To pre-train DD3D and our PL models, we use an in-house dataset that consists of multi-camera videos of urban driving scenes. DDAD15M is a larger version of DDAD : it contains high-resolution Lidar sensors to generate pointclouds and cameras synchronized with Hz scans. Most videos are 10-second long, which amounts to approximately M image frames in total. Unless noted differently, we use the entire dataset for pre-training.
2 Implementation detail
DD3D. We use V2-99 extended to an FPN as the backbone network. When pre-training DD3D, we first initialize the backbone with parameters pre-trained on the 2D detection task using the COCO dataset , and perform the pre-training on dense depth prediction using the DDAD15M dataset.
We use a test-time augmentation by resizing and flipping the input images. We observed gain in “Car” BEV AP on KITTI val, but no improvement on the nuScenes validation set. All metrics on DD3D in Section 6.2 are averages over training runs. We observed the variance over the runs to be small, i.e. BEV AP.
PL. When training PackNet , we use only the front camera images of DDAD15M to pre-train PackNet, and train until convergence. We then fine-tune the network on KITTI Eigen-clean split for epochs. For more details on training DD3D and PL, please refer to the supplementary material.
Results
In this section, we evaluate DD3D on the KITTI-3D and nuScenes benchmarks. DD3D is pre-trained on DDAD15M, and then fine-tuned on the training set of each dataset. We also evaluate PL on KITTI-3D. Its depth network is pre-trained on DDAD15M and fine-tuned on KITTI Eigen-clean, and its detection network is trained on the KITTI-3D train.
KITTI-3D. In Table 1 we compare the accuracy of DD3D to state-of-the-art methods on the KITTI-3D benchmark. DD3D achieves a significant improvement over all methods with 3D AP on Moderate Cars, which amounts to a improvement from the previous best method (). Qualitative visualization is presented in Figure 3. In Table 4 we evaluate DD3D also on the Pedestrian and Cyclist classes. DD3D outperforms all other approaches on the Pedestrian category, with an improvement from the previous best method ( vs 3D AP). On Cyclist DD3D achieves the second best result, reducing the gap to that uses ground-truth pointclouds to train a per-instance pointcloud reconstruction module and a two-stage regression network to refine 3D object proposals.
We next report the accuracy of our PL detector in Table 1. The accuracy of our PL detector is on par with state-of-the-art methods, however it performs significantly worse than DD3D ( vs. ). We will discuss this result in the context of generalizability of PL methods in Section 6.2.
nuScenes. In Table 2 we compare DD3D with other monocular methods reported on the nuScenes detection benchmark. The metrics are averages over all categories of the dataset. DD3D outperforms all other methods with a improvement in mAP compared to the previously best published method ( vs. mAP) as well as a improvement compared to the best (unpublished) method. We note that DD3D even surpasses PointPillars , which is a Lidar based detector.
In Table 5 we compare per-class accuracy of DD3D with other methods on the three major categories, with various thresholds on distance. In general, DD3D offers significant improvements across all categories and thresholds. In particular, DD3D performs significantly better on the stricter criteria: comparing to the previous best method , the relative improvements averaged across the 3 object classes are and for m and m thresholds, respectively.
2 Analysis
Here we provide a detailed analysis of DD3D, focusing on the role of depth pre-training and comparison with our PL approach. After pre-training the two models under various settings, we fine-tune them on KITTI-3D train on the 3D detection task and report AP metrics on KITTI-3D val.
Ablating large-scale depth pre-training. We first ablate the effect of dense depth pre-training on the DDAD15M dataset, with results reported in Table 3. When we omit the depth pre-training and directly fine-tune DD3D with the DLA-34 backbone on the detection task, we observe a loss in Car Moderate BEV AP. In contrast, when we instead remove the initial COCO-pretraining (i.e. pre-training on DDAD15M from-scratch), we observe a relatively small loss, i.e. . For the larger backbone of V2-99, the effect of removing the depth pre-training is even more significant, i.e. .
Dense depth prediction as pre-training task. To better quantify the effect of depth pre-training, we design a controlled experiment to further isolate its effect (Table 6). Starting from a single set of initial parameters (COCO), we consider two tasks for pre-training DD3D, 2D detection and dense depth prediction. The two pre-training phases use a common set of images that are annotated with both 2D bounding box labels and sparse depth maps obtained by projecting Lidar pointcloud. To further ensure that the comparison is fair, we applied the same number of training steps and batch size (K and ). The data used for this pre-training experiment is composed of images from the nuScenes training set.
The experiment shows that, even with the smaller scale of pre-training data compared to DDAD15M (K vs. M), the dense depth pre-training yields a notable difference in the downstream 3D detection accuracy ( vs. BEV AP on Car Moderate).
How does depth-pretraning scale? We next investigate how the unsupervised depth pre-training scales with respect to the size of pre-training data (Figure 4). For this experiment, we randomly subsample 1K, 5K, and 10K videos from DDAD15M to generate a total of 4 pre-training splits (including the complete dataset), that consist of 0.6M, 3M, 6M and 15M images respectively. Importantly, the downsampled splits contain fewer images as well as less diversity, as we downsample from the set of videos. We pre-train both DD3D and the PackNet of PL on each split, and subsequently train the detectors on KITTI-3D train. We note that DD3D and PL perform similarly at each checkpoint, and continue to improve as more depth data is used in pre-training, at least up to 15M images.
2.2 The limitations of PL methods.
In-domain depth fine-tuning of PL. Recall that training our PL 3D detector entails fine-tuning of the depth network in the target domain (i.e. Eigen-clean), after it is pre-trained on DDAD15M. We ablate the effect of the in-domain fine-tuning step in training the PL detector (Table 3). Note that in this experiment, the depth network (PackNet) is trained only on the pre-training domain (DDAD15M) and is directly applied on KITTI-3D without any adaptation. In this setting, we observe a significant loss in performance ( BEV AP). This indicates that in-domain fine-tuning of the depth network is crucial for PL-style 3D detectors. This poses a practical hurdle in using PL methods, since one has to curate a separate in-domain dataset specifically for fine-tuning the depth network. We argue that this is the main reason that PL methods are exclusively reported on KITTI-3D, which is accompanied by KITTI-Depth as a convenient source for in-domain fine-tuning. This is unnecessary for end-to-end detectors, as shown in Tables 1 and 3.
Limited generalizability of PL. With the large-scale depth pre-training and in-domain fine-tuning, our PL detector offers excellent performance on KITTI-3D val (Table 3). However, the gain from the depth pre-training is not transferred to the benchmark results (Table 1). While the loss in accuracy between KITTI-3D val and the test sets is consistent with other methods , PL suffers particularly more ( BEV AP), when compared to other methods including DD3D (). This reveals a subtle issue in the generalization of PL that is not well understood yet. We argue that the in-domain fine-tuning overfits to some image statistics that cause the performance gap between KITTI-3D val and KITTI-3D test, particularly more than other methods.
Conclusion
We proposed DD3D, an end-to-end single-stage 3D object detector that enjoys the benefit of Pseudo Lidar methods (i.e. scaling accuracy using large-scale depth data), but without its limitation (i.e. impractical training, issues in generalization). This is enabled by pre-training DD3D using a large-scale depth dataset, and fine-tuning on the target task end-to-end. DD3D achieves excellent accuracy on two challenging 3D detection benchmarks.
Appendix A Details of training DD3D and PL
We provide the training details used for supervised monocular depth pre-training of both DD3D and PackNet.
DD3D. During pre-training, we use as a batch size, and train for K steps until convergence. The learning rate starts at , decayed by at the K-th and K-th steps. The size of the input images (and projected depth map) is 1600 900, and we resize them to 910 512. When resizing the depth maps, we preserve the sparse depth values by assigning all non-zero depth values to the nearest-neighbor pixel in the resized image space (note that this is different from naive nearest-neighbor interpolation, where the target depth value is assigned zero, if the nearest-neighbor pixel in the original image does not have depth value.) We observed that training converges after epochs. We use the Adam optimizer with . For all supervised depth pre-training splits, we use an L1 loss between predicted depth and projective ground-truth depth.
When training as 3D detectors, the learning rate starts at , and is decayed by , when the training reaches and of the entire duration. We use a batch size of , and train for K and K steps for KITTI-3D and nuScenes, respectively. The and are initialized as the mean and standard deviation of the depth of the 3D boxes that are associated with each FPN level, as the stride size of the associated FPN level, and is fixed to . The raw predictions are filtered by non-maxima suppression (NMS) using IoU criteria on 2D bounding boxes. For the nuScenes benchmark, to address duplicated detections in the overlapping frustums of adjacent cameras, an additional BEV-based NMS is applied across all synchronized images (i.e. a sample) after converting the detected boxes to the global reference frame.
PackNet. When training PackNet , the depth network of PL, we use a batch size of 4 and a learning rate of with input resolution of . We use only front camera images of DDAD15M to pre-train PackNet, and train until convergence, and for 5 epochs over the KITTI Eigen-clean split during fine-tuning. The PL detector is trained with a learning rate of for 100 epochs, decayed by after 40 and 80 epochs, respectively. For both networks we use the Adam optimizer with and . Both DD3D and PL are implemented using Pytorch and trained on 8 V100 GPUs.
Appendix B DD3D architecture details
FPN is composed of a bottom-up feed-forward CNN that computes feature maps with a subsampling factor of 2, and a top-down network with lateral connections that recovers high-resolution features from low-resolution ones. The FPNs yield 5 levels of feature maps. DLA-34 FPN yields three levels of feature maps (with strides of 8, 16, and 32). We add two lower resolution features (with strides of 64, 128) by applying two 2D convs with stride of 2 (see Figure 2). V2-99 by default produces 4 levels of features (strides = 4, 8, 16, and 32), so only one additional conv is used to complete 5 levels feature maps. Note that the final resolution of FPN features derived from DLA-34 and V2-99 network are different, strides=8, 16, 32, 64, 128 for DLA-34, strides= 4, 8, 16, 32, 64 for V2-99.
2D detection head. We closely follow the decoder architecture and loss formulation of . In addition, we adopt the positive instance sampling approach introduced in the updated arXiv version . Specifically, only the center-portion of the ground truth bounding box is used to assign positive samples in and (Eq. 5).
Appendix C Pseudo-Lidar 3D confidence head
Our PL 3D detector is based on , and outputs 3D bounding boxes with 3 heads, separated based on distance (i.e. near, medium and far). Following we modify each head to output a 3D confidence, trained through the 3D bounding box loss. Specifically, each 3D box estimation head consists of 3 fully connected layers with dimensions , where denotes the bounding box parameters as described in , and denotes the 3D bounding box confidence.
Appendix D The impact of data on Pseudo-Lidar depth and 3D detection accuracy
We evaluate depth quality against the 3D detection accuracy of the PL detector, with results shown in Figure 5. Our results indicate an almost perfect linear relationship between depth quality as measured by the abs rel metric and 3D detection accuracy for our PL-based detector.