Self-supervised Sparse-to-Dense: Self-supervised Depth Completion from LiDAR and Monocular Camera

Fangchang Ma, Guilherme Venturelli Cavalheiro, Sertac Karaman

Introduction

Depth sensing is fundamental in a variety of robotic tasks, including obstacle avoidance, 3D mapping , and localization . LiDAR, given its high accuracy and long sensing range, has been integrated into a large number of robots and autonomous vehicles. However, existing 3D LiDARs have a limited number of horizontal scan lines, and thus provide only sparse measurements, especially for distant objects (e.g., the 64-line Velodyne scan in Figure 1 (a)). Furthermore, increasing the density of 3D LiDARs measurements is cost prohibitiveCurrently, the 16- and 64-line Velodyne LiDARs cost around 4kand4k and75k, respectively. Consequently, estimating dense depth from sparse measurements (i.e., depth completion) is valuable for both academic research and large-scale industrial deployment.

Depth completion from LiDAR measurements is challenging for several reasons. Firstly, the LiDAR measurements are highly sparse and also irregularly spaced in the image space. Secondly, it is a non-trivial task to improve prediction accuracy using the corresponding color image, if available, since depth and color are different sensor modalities. Thirdly, dense ground truth depth is generally not available, and obtaining pixel-level annotations can be both labor-intensive and non-scalable.

In this work, we address all these challenges with two contributions: (1) We develop a network architecture that is able to learn a direct mapping from the sparse depth (and color images, if available) to dense depth. This architecture achieves state-of-the-art accuracy on the KITTI Depth Completion Benchmark and is currently the leading method. (2) We propose a self-supervised framework for training depth completion networks. Our framework assumes a simple sensor setup with a sparse 3D LiDAR and a monocular color camera. The self-supervised framework trains a network without the need for dense labels, and outperforms some existing methods that are trained with semi-dense annotations. Our softwarehttps://github.com/fangchangma/self-supervised-depth-completion and demonstration videohttps://youtu.be/bGXfvF261pc will be made publicly available.

Related Work

Depth completion is an umbrella term that covers a collection of related problems with a variety of different input modalities (e.g., relatively dense depth input vs. sparse depth measurements ; with color images for guidance vs. without ). The problems and solutions are usually sensor-dependent, and as a result they face vastly different levels of algorithmic challenges.

For instance, depth completion for structured light sensor (e.g., Microsoft Kinect) is sometimes also referred to as depth inpainting , or depth enhancement when noise is taken into account. The task is to fill in small missing holes in the relatively dense depth images. This problem is relatively easy, since most pixels (typically over 80%) are observed. Consequently, even simple filtering-based methods can provide good results. As a side note, the inpainting problem also finds close connection to depth denoising and depth super-resolution .

However, the completion problem becomes much more challenging when the input depth image has much lower density, because the inverse problem is ill-posed. For instance, Ma et al. addressed depth reconstruction from only hundreds of depth measurements, by assuming a strong a priori of piecewise linearity in depth signals. Another example is autonomous driving with 3D LiDARs, where the projected depth measurements on the camera image space account for roughly 4% pixels . This problem has attracted a significant amount of recent interest. Specifically, Ma and Karaman proposed an end-to-end deep regression model for depth completion. Ku et al. developed a simple and fast interpolation-based algorithm that runs on CPUs. Uhrig et al. proposed sparse convolution, a variant of regular convolution operations with input normalizations, to address data sparsity in neural networks. Eldesokey et al. improved the normalized convolution for confidence propagation. Chodosh et al. incorporated the traditional dictionary learning with deep learning into a single framework for depth completion. Compared with all these prior work, our method achieves significantly higher accuracy.

Depth completion is closely related to depth prediction from a monocular color image. Research in depth prediction dates further back to early work by Saxena et al. . Since then, depth prediction has evolved from simple handcrafted feature representations to the deep learning based approaches (see the reference therein). Most learning-based work relied on pixel-level ground truth depth training. However, ground truth depth is generally not available and cannot be manually annotated. To address such difficulties, recent focus has shifted towards seeking other supervision signals for training. For instance, Zhou et al. developed an unsupervised learning framework for simultaneous estimation of depth and ego-motion from a monocular camera, using photometric loss as a supervision. However, the depth estimation is only up-to-scale. Mahjourian et al. improved the accuracy by using 3D geometric constraints, and Yin and Shi extended the framework for optical flow estimation. Li et al. recovered the absolute scale by using stereo image pairs. In contrast, in this work we propose the first self-supervised framework that is designed specifically for depth completion. We utilize the RGBd sensor data and the well-studied, traditional model-based methods for pose estimation, in order to provide absolute-scale depth supervision.

Network Architecture

We formulate the depth completion problem as a deep regression learning problem. For ease of notation, we use d for sparse depth input (pixels without measured depth are set to zero), RGB for color images (or grayscale images), and pred for depth prediction.

The proposed network follows an encoder-decoder paradigm , as displayed in Figure 2. The encoder consists of a sequence of convolutions with increasing filter banks to downsample the feature spatial resolutions. The decoder, on the other hand, has a reversed structure with transposed convolutions to upsample the spatial resolutions.

The input sparse depth and the color image, when available, are separately processed by their initial convolutions. The convolved outputs are concatenated into a single tensor, which acts as input to the residual blocks of ResNet-34 . Output from each of the encoding layers is passed to, via skip connections, the corresponding decoding layers. A final 1x1 convolution filter produces a single prediction image with the same resolution as network input. All convolutions are followed by batch normalization and ReLU, with the exception at the last layer. At inference time, predictions below a user-defined threshold τ\tau are clipped to τ\tau. We empirically set τ=0.9m\tau=0.9m, the minimal valid sensing distance for LiDARs.

In the absence of color images, we simply remove the RGB branch and adopt a slightly different set of hyper parameters: the number of filters is reduced to half (e.g., the first residual block has 32 channels, instead of 64).

Self-supervised Training Framework

Existing work on depth completion relies on densely annotated ground truth for training. However, dense ground truth generally does not exist, and even the acquisition of semi-dense labels can be technically challenging. For instance, Uhrig et al. created an annotated depth dataset by aggregating consecutive data frames using GPS, stereo vision, and additional manual inspection. However, this method is not easily scalable. Furthermore, it produces only semi-dense annotations (∼30\sim 30% pixels) within the bottom half of the image.

In this section, we propose a model-based self-supervised training framework for depth completion. This framework requires only a synchronized sequence of color/intensity images from a monocular camera and sparse depth images from LiDAR. Consequently, the self-supervised framework does not rely on any additional sensors, manual labeling work, or other learning-based algorithms as building blocks. Furthermore, this framework does not depend on any particular choice of neural network architectures. The self-supervised framework is illustrated in Figure 3. During training, the current data frame RGBd1\texttt{RGBd}_{1} and a nearby data frame RGB2\texttt{RGB}_{2} are both used to provide supervision signals. However, at inference time, only the current frame RGBd1\texttt{RGBd}_{1} is needed as input to produce a depth prediction pred1\texttt{pred}_{1}.

The sparse depth input d1\texttt{d}_{1} itself can be used as a supervision signal. Specifically, we penalize the differences between network input and output on the set of pixels with known sparse depth, and thus encouraging an identity mapping on this set. This loss leads to higher accuracy, improved stability and faster convergence for training. The depth loss is defined as

Note that a denser ground truth (e.g., the 30% dense annotation from the KITTI depth completion benchmark ), if available, can also be used in place of the sparse input d1\texttt{d}_{1}.

As an intermediate step towards the photometric loss, the relative pose between the current frame and the nearby frame needs to be computed. Prior work assumes either known transformations (e.g., stereo ) or the use of another learned neural network for pose estimation (e.g., ). In contrast, in this framework, we adopt a model-based approach for pose estimation, utilizing both RGB and d.

Specifically, we solve the Perspective-n-Point (PnP) problem to estimate the relative transformation T1→2T_{1\to 2} between the current frame 1 and the nearby frame 2, using matched feature correspondences extracted from RGBd1\texttt{RGBd}_{1} and RGB2\texttt{RGB}_{2} respectively. Random sample consensus (RANSAC) is also adopted in conjunction with PnP to improve robustness to outliers in feature matching. Compared to RGB-based estimation which is up-to-scale, our estimation is scale-accurate and failure-aware (flag returned if no estimation is found).

Given the relative transformation T1→2T_{1\to 2} and the current depth prediction pred1\texttt{pred}_{1}, the nearby color image RGB2\texttt{RGB}_{2} can be inversely warped to the current frame. Specifically, given the camera intrinsic matrix KK, any pixel p1p_{1} in the current frame 1 has the corresponding projection in frame 2 as p2=KT1→2pred1(p1)K−1p1p_{2}=KT_{1\to 2}\texttt{pred}_{1}(p_{1})K^{-1}p_{1}. Consequently, we can create a synthetic color image using bilinear interpolation around the 4 immediate neighbors of p2p_{2}. In other words, for all pixels p1p_{1}:

warped is similar to the current RGB1\texttt{RGB}_{1} when the environment is static and there’s limited occlusion due to change of view point. Note that this photometric loss is made differentiable by the bilinear interpolation. Minimizing the photometric error reduces the depth prediction error, only when the depth prediction is close enough to the ground truth (i.e., when the projected point p2p_{2} differs from the true correspondence by no more than 1 pixel). Therefore, a multi-scale strategy is applied to ensure ∥p2(s)−p1(s)∥1<1\left\lVert p_{2}^{(s)}-p_{1}^{(s)}\right\rVert_{1}<1 on at least one scale ss. In additional, to avoid conflicts with the depth loss, the photometric loss is evaluated only on pixels without direct depth supervision. The final photometric loss is

where SS is the set of all scaling factors, and (⋅)(s)(\cdot)^{(s)} represents image resizing (with average pooling) by a factor of ss. Losses at lower resolutions are weighted down by ss.

The photometric loss only measures the sum of all individual errors (i.e., color differences computed on each pixel independently) without any neighboring constraints. Consequently, minimizing the photometric loss alone usually results in an undesirable local optimum, where the depth pixels have incorrect values (despite having a low photometric error) and high discontinuity. To alleviate this issue, we add a third term to the loss functions in order to encourage smoothness of the depth predictions. Inspired by , we penalize ∥∇2pred1∥1\left\lVert\nabla^{2}\texttt{pred}_{1}\right\rVert_{1}, the L1\mathcal{L}_{1} loss of the second-order derivatives of the depth predictions, to encourage piecewise-linear depth signal.

In summary, the final loss function for the entire self-supervised framework consists of 3 terms:

where β1,β2\beta_{1},\beta_{2} are relative weightings. Empirically we set β1=0.1\beta_{1}=0.1 and β2=0.1\beta_{2}=0.1.

Implementation

For the sake of benchmarking against state-of-the-art methods, we use the KITTI depth completion dataset for both training and testing. The dataset is created by aggregating LiDAR scans from 11 consecutive frames into one, producing a semi-dense ground truth with roughly 30% annotated pixels. The dataset consists of 85,898 training data, 1,000 selected validation data, and 1,000 test data without ground truth.

For the PnP pose estimation, we dialate the sparse depth images d1\texttt{d}_{1} with a 4×44\times 4 kernel, since the extracted features points might not have spot-on depth measurements. In each epoch, we iterate through the entire training dataset for the current frame 1, and choose a neighbor frame 2 randomly from the 6 nearest frames in time (excluding the current frame itself). In presence of PnP pose estimation failure, T1→2T_{1\to 2} is set to be an identity matrix and the neighbor RGB2\texttt{RGB}_{2} image is overwritten by the current RGB1\texttt{RGB}_{1}. Consequently, the photometric loss is made to be 0, and does not affect the training.

The training framework is implemented in PyTorch . Zero-mean Gaussian random initialization is used for the network weights. We use a batch size of 8 for the RGBd-network, and 16 for the simpler d-network. Adam with a starting learning rate of 10−510^{-5} is used for network optimization. The learning rate is reduced to half every 5 epochs. We use 8 Tesla V100 GPUs with 16G of RAM for training, and 12 epochs takes roughly 12 hours for the RGBd-network and 4 hours for the d-network.

Results

In this section, we present experimental results to demonstrate the performance of our approach. We first compare our network architecture, trained in a purely supervised fashion, against state-of-the-art published methods. Secondly, we conduct an ablation study on the proposed network architecture to gain insight into which components contribute to the prediction accuracy. Lastly, we showcase training results using our self-supervised framework, and present an empirical study on how the algorithm performs under different level of sparsity in the input depth signals.

In this section, we train our best network in a purely supervised fashion to benchmark against other published results. We use the official error metrics for the KITTI depth completion benchmark , including rmse, mae, irmse, and imae. Specifically, rmse and mae stand for the root-mean-square error and the mean absolute error, respectively; irmse and imae stand for the root-mean-square error and the mean absolute error in the inverse depth representation. The results are listed in Table 1 and visualized in Figure 4.

Our d-network leads prior work with a large margin in almost all metrics. The RGBd-network attains even higher accuracy, leading all submissions to the benchmark. Our predicted depth images also have cleaner and sharper object boundaries (e.g., see trees, cars and road signs), which can be attributed to the fact that our network is quite deep (and thus might be able to learn more complex semantic representations) and has large skip connections (and thus preserves image details). Note that all these supervised methods produce poor predictions at the top of the image, because of 2 reasons: (a) the LiDAR returns no measurements, and thus the input to the network is all zero at the top; (b) the 30% semi-dense annotations do not contain labels in these top regions.

2 Ablation Studies

To examine the impact of network components on performance, we conduct a systematic ablation study and list the results column-wise in Table 2.

The most effective components in improving final accuracy includes using RGBd for input and L2\mathcal{L}_{2} loss for training. This is in contrary to the findings that L1\mathcal{L}_{1} is more effective , implying that the optimal loss functions might be dataset- and architecture-dependent. Adding skip connections, training from scratch (without ImageNet-pretraining), and not using max pooling also result in substantial improvement. Increasing network depth (from 18 to 34) and encoders-decoders pairs (from 3 to 5), as well as a proper split of filters allocated to the RGB and the d branches (16/48 split), also create small positive impact on the results. However, additional regularization, including dropout combined with a weight decay, leads to degraded performance.

It is worth noting that alternative encoding of the input depth image (such as the nearest neighbor interpolation or the bilinear interpolation of the sparse depth measurements) does not improve the prediction accuracy. This implies that the proposed network is able to deal with highly sparse input image.

3 Evaluation of the Self-supervised Framework

In this section, we evaluate the self-supervised training framework described in Section 4 on the KITTI validation dataset. We compare 3 different training methods: using only photometric loss without sparse depth supervision, the complete self-supervised framework (i.e., photometric loss with sparse depth supervision), and the pure supervised method using the semi-dense annotations. The quantitative results are listed in Table 3. The self-supervised result produces rmse=1384\text{{rmse}}=1384, which already outperforms some of the prior methods that were trained with semi-dense annotations, such as SparseConvs .

However, note that the true quality of depth predictions trained in a self-supervised fashion is probably underestimated by such evaluation metrics, since the “ground truth” itself is biased. Specifically, the evaluation ground truth is characterized by the same limitations as the training annotations: low-density, as well as absence at the top region. As a result, predictions at the top, where the self-supervised framework provides supervision but semi-dense annotations do not, are not reflected in the error metrics, as illustrated in Figure 5.

The self-supervised framework is effective for not only 64-line lidar measurements, but also lower-resolution lidars and more sparse depth input. In Figure 6(b), we show the validation errors of the networks trained with the self-supervised framework with different levels of sparsity in the depth. When the number of input measurements is too small, the validation error is high. This is expected due to failure in PnP pose estimation. However, with sufficiently many measurements (e.g., at least 4 scanlines, or the equivalent number of samples to at least 2 scanlines when input is uniformaly sampled), the validation error starts to decrease as a power function of the input, similar to training with semi-dense annotations.

4 On Input Sparsity

In many robotic applications, engineers need to address the following question: what’s the LiDAR resolution (which translates to financial cost) required to achieve certain performance? In this section, we try to answer this question by evaluating the accuracy of our LiDAR depth completion technique under different input sparsity and spatial patterns. To this end, we provide an empirical analysis on the depth completion accuracy for different depth input with varying levels of sparsity and spatial patterns. In particular, we downsample the raw LiDAR input in two different manners: reducing the number of laser scans (to simulate a LiDAR with fewer scan lines), and uniformly sub-sampling from all LiDAR measurements available. The results are illustrated in Figure 6, for both of these spatial patterns and both input modalities of d and RGBd.

In Figure 6(a) we show the validation errors when trained with semi-dense annotations. The rmse errors form a straight line in the log-log plot, implying that the depth completion error decreases as a power function cxpcx^{p} of the number of input depth measurements, for some positive cc and negative pp. This also implies diminishing returns on increasing LiDAR resolutions. Comparing the two spatial patterns, uniform random sub-sampling produces significantly higher accuracy than having a reduced number of scan lines, since the input depth samples are more disperse in the pixel space with uniform random sampling. Furthermore, using RGBd substantially reduces prediction error, compared to using only d, when trained with semi-dense annotations. The performance gap is especially significant when the number of depth measurements is low. Note that there is a significant drop of RMSE from 32-line to 64-line LiDAR. This accuracy gain may be attributed to the fact that our network architecture is optimized for 64-line LiDAR.

In Figure 6(b), we show results when trained with our self-supervised framework. As has been discussed in Section 6.3, the validation error starts to decrease steadily as a power function, similar to training with semi-dense annotations, when there are sufficiently many input measurements. However, with the self-supervised framework, using both RGB and sparse depth yields the same level of accuracy as using sparse depth only, which is different from training with semi-dense annotations. The underlying cause of this difference remains to be further investigatedIn the self-supervised framework, the training process is more iterative than training with semi-dense annotations. In particular, it takes many more iterations for the predictions to converge to the correct value. Consequently, the network weights for the RGB input, which has substantially lower correlation with the depth prediction than the sparse depth input, might have dropped to negligible levels during early iterations, resulting in similar performance for using d and RGBd as input. However, this conjecture remains to be verified..

Conclusions

In this paper, we have developed a deep regression model for depth completion of sparse LiDAR measurements. Our model achieves state-of-the-art performance on the KITTI depth completion benchmark, and outperforms existing published work by a significant margin at the time of submission. We also propose a highly scalable, model-based self-supervised training framework for depth completion networks. This framework requires only sequences of RGB and sparse depth images, and outperforms a number of existing solutions trained with semi-dense annotations. Additionally, we present empirical results demonstrating that depth completion errors decrease as a power function with the number of input depth measurements. In the future, we will investigate techniques for improving the self-supervised framework, including better loss functions and taking dynamic objects into account.

This work was supported in part by the Office of Naval Research (ONR) grant N00014-17-1-2670 and the NVIDIA Corporation. In particular, we gratefully acknowledge the support of NVIDIA Corporation with the donation of the DGX-1 used for this research. Finally, we thank Jonas Uhrig and Nick Schneider for providing information on how data is generated for the KITTI dataset .

References