Learning Guided Convolutional Network for Depth Completion

Jie Tang, Fei-Peng Tian, Wei Feng, Jian Li, Ping Tan

I Introduction

Dense depth perception is critical for many robotics applications, such as autonomous driving or other mobile robots. Accurate dense depth perception of the observed image is the prerequisite for solving the following tasks such as obstacle avoidance, object detection or recognition and 3D scene reconstruction. While depth cameras can be easily adopted in indoor scenes, outdoor dense depth perception mainly relies on stereo vision or LiDAR sensors. Stereo vision algorithms still have many difficulties in reconstructing thin and discontinuous objects. So far, LiDAR sensors provide the most reliable and most accurate depth sensing and have been widely integrated into many robots and autonomous vehicles. However, current LiDAR sensors only obtain sparse depth measurements, e.g. 64 scan lines in the vertical direction. Such a sparse depth sensing is insufficient for real applications like robotic navigation. Thus, estimating dense depth map from the sparse LiDAR input is of great value for both academic research and industrial applications.

Many recent works on this topic take deep learning as approach and exploit an additional synchronized RGB image for depth completion. These methods have achieved significantly improvements over conventional approaches . For example, Qiu et al. train a network to estimate surface normal from both the RGB image and LiDAR data and further use the recovered surface normal to guide depth completion. Ma et al. exploit photo-consistency between neighboring video frames for depth completion. Jaritz et at. adopt a depth loss as well as a semantic loss for supervision. Despite the different methods proposed by these works, they basically share the same scheme in multi-modal feature fusion. Specifically, these works adopt the operation like concatenation or element-wise addition to fuse the feature vectors from sparse depth and RGB image together directly for further processing. However, the commonly used concatenation or element-wise addition operation is not such appropriate when considering the heterogenous data and the complex environments. The potentiality of RGB image as guidance is difficult to be fully exploited by applying in such a simple way. In contrast, we suggest a more sophisticated fusion module to improve the performance of the depth completion task.

Our work is inspired by the guided image filtering . In guided image filtering, the output at a pixel is a weighted average of nearby pixels, where the weights are functions of the guidance image. This strategy has been adopted for generic completion/super-resolution of RGB and range images . Inspired by the success of guided image filtering, we seek to learn a guided network to automatically generate spatially-variant convolution kernels according to the input image and then apply them to extract features from sparse depth image by our guided convolution module. Compared with the hand-crafted function for kernel generation in guided image filtering , our end-to-end learned network structure has a potential to produce more powerful kernels with agreement of scene context for depth completion. Compared with standard convolutional module, where the kernel is spatially-invariant and pixels at all the positions share the same kernels, our guided convolutional module has spatially-variant kernels that are automatically generated according to the content. Thus, our network is more powerful to handle various challenging situations in depth completion task.

An obvious drawback of using spatially-variant kernels is the large GPU memory consumption, which is also the original motivation of parameter sharing in the convolutional neural network. Especially when applying the spatially-variant convolution module in the multi-stage fusion for depth completion, the massive GPU memory consumption is even unaffordable for computational platforms (See subsection III-C for memory and computation discussion). Thus, it’s non-trivial to look for a practical way to make the network available. Inspired by recent network compression technique , we factorize the convolution operation in our guided convolution module to two stages, a spatially-variant channel-wise convolution stage and a spatially-invariant cross-channel convolution stage. By using such a novel factorization, we get an enormous reduction of GPU memories such that the guided convolution module can be integrated with the powerful encoder-decoder network in multi-stages in a modern GPU device.

The proposed method is evaluated on both outdoor and indoor datasets, from real-world and synthetic scenes. It outperforms the state-of-the-art methods on KITTI depth completion benchmark and rank 1st at the time of paper submission. Comprehensive ablation studies demonstrate the effectiveness of each component and the fusion strategy used in our method. Compared with other depth completion methods, our method also achieves the best performance on the indoor NYUv2 datset. Last but not least, our model presents strong generalization capability under different depth point densities, various lighting and weather conditions as well as cross-dataset evaluations. Our code will be released at https://github.com/kakaxi314/GuideNet.

II Related Work

Depending on whether there is an RGB image to guide the depth completion, previous methods can be roughly divided into two categories: depth-only methods and image-guided methods. We briefly review these techniques and other literatures relevant to our network design.

Depth-only Methods These methods use a sparse or low-resolution depth image as input to generate a full-resolution depth map. Some early methods reconstruct dense disparity maps or depth maps based on the compressive sensing theory or a combined wavelet-contourlet dictionary . Ku et al. use a series of hand-crafted conventional operators like dilation, hole closure, hole filling, and blurring, etc., to transform sparse depth maps into dense. More recently, deep learning based approaches demonstrate promising results. Uhrig et al. propose a sparsity invariant CNN to deal with sparse data or features by using an observation mask. Eldesokey et al. solve depth completion via generating a full depth as well as a confidence map with normalized convolution. Chodosh et al. combine compressive sensing with deep learning for depth prediction. The main focus of these methods is to design appropriate operators, e.g. sparsity invariant CNN , to deal with sparse inputs and propagate these spare information to the whole image.

In terms of depth super-resolution, some methods exploit a database of paired low-resolution and high-resolution depth image patches or self-similarity searching to generate a high resolution depth image. Some methods further propose to solve depth super-resolution by dictionary learning. Riegler et al. use a deep network to produce a high-resolution depth map as well as depth discontinuities and feed them into a variational model to refine the depth. Unlike these depth super-resolution methods, which take the dense and regular depth image as input. Instead, the depth input in our method is sparse and irregular, and also we train our model end-to-end without any further optimization or post-processing.

Image-guided Methods These methods usually achieve better results, since they utilize an additional RGB image, which provides strong cues on semantic information, edge information, or surface information, etc. Earlier works mainly address depth super-resolution with bilateral filtering , or global energy minimization , where the depth completion is guided by image , semantic segmentation or edge information .

Recently, Zhang et al. propose to predict surface normal and occlusion boundary from a deep network and further utilize them to help depth completion in indoor scenes. Qiu et al. extend a similar surface normal as guidance idea to the outdoor environment and recover dense depth from sparse LiDAR data. Ma et al. propose a self-supervised network to explore photo-consistency among neighboring video frames for depth completion. Huang et al. propose three sparsity-invariant operations to deal with sparse inputs. Eldesokey et al. combine their confidence propagation with RGB information to solve this problem. Gansbeke et al. use two parallel networks to predict depth and learn an uncertainty to fuse two results. Cheng et al. use CNN to learn the affinity among neighboring pixels to help depth estimation.

Although various approaches have been proposed for depth completion with a reference RGB image, they almost share the same strategy in fusing depth and image features, which is simple concatenation or element-wise addition operation. In this paper, inspired by guided image filtering , we propose a novel guided convolution module for feature fusion, to better utilize the guidance information from the RGB image.

Joint Filtering and Guided Filtering Our method is also relevant to joint bilateral filtering and guided image filtering . Joint/guided image filtering utilizes a reference or guidance image as prior and aims to transfer the structures from the reference image to the target image for color/depth image super-resolution , image restoration , etc.

Early joint filtering methods explore common structures between target and reference images and formulate the problem as iterative energy minimization. Recently, Li et al. propose a CNNs based joint filtering for image noise reduction, depth upsampling etc., but the joint filtering is implemented as a simple feature concatenation. Gharbi et al. generate affine parameters by a deep network to perform color transforms for image enhancement. Lee et al. adopt a similar bilateral learning scheme of but generate bilateral weights and apply them once on a pre-obtained depth map for depth refinement. In contrast, our guided convolution module works on image features and serves as a flexibly pluggable component in multiple stages of an encoder-decoder network.

In , Wu et al. propose a guided filtering layer to perform joint upsampling, which is close to our work. It directly reformulates the conventional guided filter and make it differentiable as a neural network layer. As a result, the kernel weights are generated by the same close-form equation of guided filter to filter the input image. This kind of operator is inapplicable to fill-in sparse LiDAR points, as commented by the authors of guided filter in their conference paper . Our method is also inspired by guided filter . Rather than generating guided filter kernels from a specific close-form equation, we consider to learns more general and powerful kernels from the guidance image and applies the kernels to fuse multi-modal features for depth completion task.

Dynamic Filtering On the other hand, in convolutional neural networks, Dynamic Filtering Network (DFN) is a broad category of methods where the network generates filter kernels dynamically based on the input image to enable operations like local spatial transformation on the input features. The general concept first proposed in is mainly evaluated on video (and stereo) prediction with previous frames as input.

Recently, several applications and extensions of DFN have been developed. ‘Deformable convolution’ dynamically generates the offsets to the fixed geometric structure which can be seen as an extension of DFN by focusing on the sampling locations. Simonovsky et al. extends DFN into the graph signals in spatial domain, where the filter weights are dynamically generated for each specific input sample and conditioned on the edge labels. Wu et al. propose an extension of DFN by using multiple sampled neighbor regions to dynamically generate weights with larger receptive fields.

Our kernel generating approach shares the same philosophy with DFN and can be considered as a variant and extension, focusing on multi-stage feature fusion of multi-modal data. The spatially-variant kernels generated by DFN consume large GPU memories and thus are only applied once on low resolution images or features. However, multi-stage feature fusion is critical for feature extraction from sparse depth and color image on the depth completion task, but has not been studied by previous DFN papers. To address it, we design a novel network structure with convolution factorization and further discuss the impact of fusion strategies on depth completion results.

III The Proposed Method

Given a sparse depth map S\mathbf{S} generated by projecting the LiDAR points to the image plane with calibration parameters and a RGB image I\mathbf{I} as guidance reference, depth completion aims to produce a dense depth map D\mathbf{D} of the whole image. The RGB image can provide extremely useful information for depth completion task, as it depicts object boundaries and scene contents.

To explain our guided convolutional network to upgrade S\mathbf{S} to D\mathbf{D} with the guidance of I\mathbf{I}, we first briefly review the guided image filtering which inspires our guided convolution module in subsection III-A. Then we elaborate the design of the guided convolution module in subsection III-B and introduce a novel convolution factorization in subsection III-C. In the next, we explain how this module can be used in a common encoder-decoder network, and the multi-stage fusion scheme used in our method in subsection III-D. Finally, we give implementation details including hyperparameter settings in subsection III-E.

The guided image filtering generates spatially-variant filters according to a guidance image. In our setting of depth completion task, this method would compute the value at a pixel ii in D\mathbf{D} as a weighted average of nearby pixels from S\mathbf{S}, i.e.

Here, ii, jj are pixels indexes and N(i)\mathcal{N}(i) is a local neighborhood of the pixel ii. The kernel weights Wij\mathbf{W}_{ij} are computed according to the guidance image I\mathbf{I} and a hand-crafted closed-form equation similar to the matting Laplacian from . Unless specifically indicating, we omit the index of the image or feature channel for simplifying notations.

This guided image filtering might be applied to image super-resolution like in . However, our input LiDAR points are sparse and irregular. As pointed by the authors of , the guided image filtering cannot work well on sparse inputs. This motivates us to learn more general and powerful filter kernels from the guidance image I\mathbf{I}, rather than using the hand-crafted function for kernel generation. And then we apply the kernels to fuse the multi-modal features, not directly filtering on the input images.

III-B Guided Convolution Module

Here, we elaborate the design of our guided convolution module that generates content-dependent and spatially-variant kernels for depth completion.

As shown in Figure 1, our guided convolution module would server as a flexibly pluggable component to fuse the features from RGB and depth image in multiple stages. It would generate convolutional kernels automatically from the guidance image feature I\mathcal{I} and apply them to the sparse depth map feature S\mathcal{S}. Here, I\mathcal{I} and S\mathcal{S} are features extracted from the guidance image I\mathbf{I} and sparse depth map S\mathbf{S} respectively. We denote the output from this guided convolution module as D\mathcal{D}, which is the extracted feature of depth image. Formally,

The advantages of content-dependent and spatially-variant kernels are two-folds. Firstly, this kind of kernels allows the network to apply different filters to different objects (and different image regions). It is useful because, for example, the depth distribution on a car would be different from that on the road (also, nearby and faraway cars own different depth distributions). Thus, generating the kernels dynamically according to the image content and spatial position would be helpful. Secondly, during training, the gradient of a spatially-invariant kernel is computed as the average over all image pixels from the next layer. Such an average is more likely leading to gradient closing to zero, even thought the learned kernel is far from optimal for every position, which could generate sub-optimal results as pointed by . In comparison, spatially variant kernels can alleviate this problem and make the training better behaved, which towards to stronger results.

III-C Convolution Factorization

However, generating and applying these spatially-variant kernels naïvely would consume a large amount of GPU memory and computation resources. The enormous GPU memory consumption is unaffordable for modern GPU device, when integrating the guided convolution module into multi-stage fusion of an encoder-decoder network. To address this challenge, inspired by recent network compression techniques, e.g. MobileNets , we design a novel factorization as well as a matched network structure to split the guided convolution module into two stages for better memory and computation efficiency. This step is critical to make the network practical.

Memory & Computation Efficiency Analysis. Now, we analyze the improvement of this two-stage strategy in terms of memory and computation efficiency. If the convolution operations ⊗\otimes in Equqation (2) is implemented naïvly, the target depth feature Dp,n\mathcal{D}_{p,n} at a pixel pp and the channel nn can be formalized explicitly as

where kk is the offset in a K×KK\times K filter kernel window centered at pp and mm is the channel index of S\mathcal{S}. Suppose the height and width of the input depth feature S\mathcal{S} are HH and BB respectively. It is easy to figure out that the size of the generated kernel is (M×N×K2×H×B)(M\times N\times K^{2}\times H\times B). In an encoder-decoder network, HH and BB are usually very large in the initial scales of the encoder or the end scales of the decoder. MM and NN usually go up to hundreds or even thousands in the latent space. Hence, the memory consumption is high and unaffordable even for modern GPUs.

By our convolution factorization, we split convolution in Equation (5) into a channel-wise convolution in Equation (3) and a cross-channel convolution in Equation (4). We can explicitly re-formulate these two equations in detail as

The computation complexities of Equation (6) and Equation (7) are O(K2)O(K^{2}) and O(M)O(M) respectively. Therefore, by this novel convolution factorization, we reduce the computational complexity of Dp,n\mathcal{D}_{p,n} from O(M×K2)O(M\times K^{2}) to O(M+K2)O(M+K^{2}).

Moreover, the proposed factorization can reduce GPU memory consumption enormously. This is extremely important for networks with multi-stage fusions. Suppose the memory consumption by the proposed factorization and naïve convolution are MfM_{f} and MsM_{s} respectively, then

As an example, when using 4-byte floating precision and taking M=N=128M=N=128, H=64H=64, B=304B=304, and K=3K=3, which is the setting of the second fusion stage of our network, the proposed two-stage convolution reduces GPU memory from 10.7GB to 0.08GB, nearly 128 times lower for just a single layer. In this way, our guided convolution module can be applied on multiple scales of a network, e.g. in an encoder-decoder network.

III-D Network Architecture

Figure 1 illustrates the overall structure of the proposed network, which is based on two encoder-decoder networks with skip layers. Here, we refer the two networks taking the RGB image I\mathbf{I} and sparse LiDAR depth image S\mathbf{S} as GuideNet and DepthNet respectively. The GuidedNet aims to learn hierarchical feature representations with both low-level and high-level information from RGB image. Such image features are used to generate spatially-variant and content-dependent kernels automatically for depth feature extractions. The DepthNet takes the LiDAR depth image as input and progressively fuse hierarchical image features by the guided convolution module in encoder stage. It then regresses dense depth image at the decoder stage. Both encoders of GuidedNet and DepthNet consist of a trail of ResNet blocks . Convolution layer with stride is used to aggregate feature to low resolution in encoder stage, and deconvolution layer in decoder stage upsamples the feature map to high resolution. We also add standard convolution layers at the beginning of both GuideNet and DepthNet as well as the end of DepthNet.

Please note that during feature fusion, instead of the early or late fusion scheme widely used in the existing methods , we utilize a novel fusion scheme which fuse the decoder features of the GuidedNet to the encoder features of the DepthNet. In our network, image features act as guidance for the generation of depth feature representations. Thus, compared with encoder features, features from the decoder stage in the GuideNet are preferable, as they own more high-level context information. In addition, in contrast to fuse only once, we fuse the two sources in multi-stage, which shows stronger and more reliable results. More comparisons and analyses can be found in subsection IV-D.

III-E Implementation Details

During training, we adopt the mean squared error (MSE) to compute the loss between ground truth and predicted depth. For real-world data, the ground truth depth is often semi-dense, because it is difficult to collect ground truth depth for every pixel. Therefore, we only consider valid pixels in the reference ground truth depth map when computing the training loss. The final loss function is

III-E2 Training Setting

We use ADAM as the optimizer with a starting learning rate of 10−310^{-3} and weight decay of 10−610^{-6}. The learning rate drops by half every 50k50k iterations. We utilize 2 GTX 1080Ti GPUs for training with batch size of 8. Synchronized Cross-GPU Batch Normalization is used in the network training stage. Our method is trained end-to-end FROM SCRATCH. In contrast, some state-of-the-art methods employ extra datasets for training, e.g. DeepLiDAR utilizes synthetic data to train the network for obtaining scene surface normal, and the authors of use a pretrained model on Cityscapes https://www.cityscapes-dataset.com as network initialization.

IV Experiments

We conduct comprehensive experiments to verify our method on both outdoor and indoor datasets, captured in real-world and synthetic scenes. We first introduce all the datasets and evaluation metrics used in our experiments in subsection IV-A and IV-B respectively. Then, as autonomous driving is the major application of depth completion, we compare our method with the state-of-the-art methods on the outdoor scene KITTI dataset in subsection IV-C. It follows by extensive ablation studies on the KITTI validation set in subsection IV-D to investigate the impact of each network component and the fusion scheme used in our method. In subsection IV-E, we verify the performance of proposed method on the indoor scene NYUv2 dataset. Finally, in subsection IV-F, we perform experiments under various settings including input depth with different densities, RGB images captured under various lighting and weather conditions and cross-dataset evaluations to prove generalization capability of our method.

KITTI Dataset The KITTI depth completion dataset contains 86,89886,898 frames for training, 1,0001,000 frames for validation, and another 1,0001,000 frames for testing. It provides public leaderboard http://www.cvlibs.net/datasets/kitti/eval_depth.php?benchmark for ranking submissions. The ground truth depth is generated by registering LiDAR scans temporally. These registered points are further verified with the stereo image pairs to get rid of noisy points. As there are rare LiDAR points at the top of an image, following , input images are cropped to 256×1216256\times 1216 for both training and testing.

Virtual KITTI Dataset Virtual KITTI dataset is a synthetic dataset, where the virtual scenes are cloned from the real world KITTI video sequences. Besides the 5 virtual image sequences cloned from KITTI sequence, it also generates the corresponding image sequences under various lighting conditions (like morning, sunset) and weather conditions (like fog, rain), totally 17,000 image frames. To generate sparse LiDAR points, instead of random sampling from the dense depth map, we use the sparse depth of the corresponding image frame in KITTI dataset as a mask to obtain sparse samples from dense ground truth depth, which makes the distribution of sparse depth on image is close to real-world situation. We split the whole Virtual KITTI dataset to train and test set to fine-tune and evaluate our model respectively. Since the destination is to verify the robustness of our model under various lighting and weather condition, we only fine-tune our model under the original ‘clone’ condition whose weather is good, using sequence of ‘0001’, ‘0002’, ‘0006’ and ‘0018’ for training. And the sequence ‘0020’ with various weather and lighting conditions is used for evaluation. In summary, we have 1289 frames for fine-tuning and 837 frames for each condition to evaluate.

NYUv2 Dataset NYUv2 dataset consists of RGB images and depth images captured by Microsoft Kinect in 464 indoor scenes. Following the similar setting of previous depth completion methods , our method is trained on 50k50k images uniformly sampled from the training set, and tested on the 654 official labeled test set for evaluation. As a preprocessing, the depth values are in-painted using the official toolbox, which adopts the colorization scheme to fill-in missing values. For both train and test set, the original frames of size 640×480640\times 480 are half down-sampled with bilinear interpolation, and then center-cropped to 304×228304\times 228. The sparse input depth is generated by random sampling from the dense ground truth. Due to the input resolution for our network must be a multiple of 3232, we futher pad the images to 320×256320\times 256 as input for our method but evaluate only the valid region of size 304×228304\times 228 to keep fair comparison with other methods.

SUN RGBD Dataset The SUN RGBD dataset is an indoor dataset containing RGB-D images from many other datasets . We only use SUN RGBD dataset for cross-dataset evaluation. Since NYUv2 dataset is a subset of SUN RGBD dataset, we exclude them in evaluation to avoid repetition. We keep all images with the same resolution of NYUv2 dataset as 640×480640\times 480, captured under different scenes. Totally, we evaluate our model on 3944 image frames, with 555 frames captured by Kinect V1 and 3389 captured by Asus Xtion camera. The same pre-processing method for NYUv2 dataset is used to fill depth map. Note, frames captured by Asus Xtion camera are more challenging, because the data comes from a different device.

IV-B Evaluation Metrics

Following the KITTI benchmark and exiting depth completion methods , for outdoor scene, we use these four standard metrics for evaluation: root mean squared error (RMSE), mean absolute error (MAE), root mean squared error of the inverse depth (iRMSE) and mean absolute error of the inverse depth (iMAE). Among them, RMSE and MAE directly measure depth accuracy, while RMSE is more sensitive and chosen as the dominant metric to rank submissions on the KITTI leaderboard. iRMSE and iMAE compute the mean error of inverse depth, which gives less weight for far-away points.

For indoor scene, to be consistent with comparative depth completion methods , the evaluation metrics are selected as root mean squared error (RMSE), mean absolute relative error (REL) and δi\delta_{i} which means the percentage of predicted pixels where the relative error is less a threshold ii. Specifically, ii is chosen as 1.251.25, 1.2521.25^{2} and 1.2531.25^{3} separately for evaluation. Here, a higher ii indicates a softer constraint and a higher δi\delta_{i} represents a better prediction. RMSE is chosen as the primary metric for all the experiment evaluations as it is sensitive to large errors on distant regions.

IV-C Experiments on KITTI Dataset

We first evaluate our method on the KITTI depth completion dataset . Our method is trained end-to-end from scratch on the train set and compared the performance with state-of-the-art methods on test set. Table I lists the quantitative comparison of our method and other top-ranking published methods on the KITTI leaderboard. Our method ranks 1st and exceed all other methods under the primary RMSE metric at the time of paper submission, and presents comparable performance on other evaluation metrics.

Figure 3 shows some visual comparison results with several state-of-the-art methods on the KITTI test set. Our results are shown in the last row. While all methods provide visually plausible results in general, our estimated depth maps reveal more details and are more accurate around object boundaries. For example, our method can better recover depth of background between the arms of a person as highlighted by the magenta circle in Figure 3. The predicted depth of our method owns the most accurate contour in the black car region.

IV-D Ablation Studies

To investigate the impact of each network component and fusion scheme on the final performance, we conduct ablation studies on the KITTI validation dataset. Specifically, we evaluate several different variations of our network. The quantitative comparisons are summarized in Table II.

We can see that the results of ‘Add.’ is a slightly worse than that of ‘Concat.’. This is also reasonable, because image and depth features are heterogeneous data from different sources. By applying addition, we implicitly treat these two different features in the same way, which leads to performance drops. Indeed, most of state-of-the-art methods adopt concatenation to fuse the heterogeneous depth and image features while apply addition to fuse homogeneous depth features from different stages.

IV-D2 Fusion Scheme of GuideNet and DepthNet

As described in subsection III-D, instead of using early or late feature fusion like existing methods , our approach fuses the decoder features of the GuideNet to the encoder features of the DepthNet. To verify the effectiveness of such a fusion scheme, we train and evaluate the performance of fusing the decoder features of the GuideNet to the decoder features of the DepthNet (referred as ‘D-D Fusion’) and fusing the encoder features of the GuideNet to the encoder features of the DepthNet (referred as ‘E-E Fusion’). In the later one, the decoder structure of the GuideNet is removed since it is not used anymore. In this way, our method can be seen as ‘D-E Fusion’.

Table II compares the results of ‘E-E Fusion’ and ‘D-D Fusion’ with our method. The performance drop of the ‘E-E Fusion’ verifies our earlier analysis that the decoder image features own more high-level context information thus can better guide depth feature extraction. The ‘D-D Fusion’, fusing image and depth features in the decoder stage, suffers from even larger performance drop. Comparing the ‘D-D Fusion’ and our final model, we conclude that the image guidance is more effective at encoder stage of depth feature extraction. It’s also reasonable and easy to understand, as feature extracted in early stage can influence the following feature extraction, especially for sparse depth image.

On the other hand, even the weaker fusion strategy in the ‘E-E Fusion’ outperforms conventional feature addition or concatenation. This attributes to our guided convolution module that can generate content-dependent and spatially-variant kernels to promote the depth completion. This observation further proves the effectiveness of the proposed guided convolution module.

IV-D3 Fusion Scheme of Multi-stage Guidance

We also design two other variants to verify the effectiveness of multi-stage guidance scheme. For comparison, based on our guided network, we replace all the guided modules with concatenation except the one in the first fusion stage, and refer it as ‘First Guide’. From the same view, we use ‘Last Guide’ to refer the condition only the guided module in the last fusion stage is remained. Using concatenation for the feature fusion of other stages is from the result, that concatenation can perform a little better than addition operation as shown in Table II.

We can see that both the results of ‘First Guide’ and ‘Last Guide’ are worse than our multi-stage guidance scheme. This demonstrates the effectiveness of our multi-stage guidance design. Also, the ‘First Guide’ performs a little bit better than ‘Last Guide’. It also consists with our early analysis that image guidance is more effective at early stage, since feature extracted in early stage can influence the following feature extraction. Moreover, both the results of ‘First Guide’ and ‘Last Guide’ perform better than the ‘Concat.’. It once more verifies that the designed Guided Convolution Module is a much powerful fusion scheme for depth completion.

IV-E Experiments on NYUv2 Dataset

To verify the performance of our method on indoor scene, we directly train and evaluate our guided network on the NYUv2 dataset , without any specific modification.

Following existing methods, we train and evaluate our method with the settings of 200 and 500 sparse LiDAR samples separately. The quantitative comparisons with other methods are shown in Table III. The results of ‘Bilateral’ , and ‘CSPN’ come from the CSPN . The results of ‘TGV’ , ‘Zhang et al.’ and ‘DeepLiDAR’ are obtained from DeepLiDAR . By using the released implementations, we get the results of ‘Ma et al.’ with 500 samples and ‘NConv-CNN’ with 200 samples. We can see from the results, our method outperforms all other methods in both settings of 500 samples and 200 samples. Without specific modification, our method ranks top under all these 5 evaluation metrics.

We also show some qualitative comparisons on the test set in Figure 5. Our method is compared with ‘NConv-CNN’ and ‘Ma et al.’ on the settings of 200 samples and 500 samples. The most notable regions are selected with cyan rectangles for easy comparisons. From the predicted depth, we can see the results of ‘Ma et al.’ over-smooth the whole image and blur small objects. Even though ‘NConv-CNN’ shows much clear depth predictions, it also suffers obvious detail loss at object structures, especially the thin object boundaries. Our method show sharp transitions aligning to local details and generate the best results.

IV-F Generalization Capability

To prove the generalization capability of our method, we test its performance under different point densities, various lighting and weather conditions as well as cross-dataset evaluations.

We test the performance of our method under different point densities. Our model is the same one trained from scratch only on the KITTI train set without any fine-tuning, to faithfully reflect its generalization capability. For a comparison, we also evaluate another two state-of-the-art methods, ‘NConv-CNN’ and ‘Sparse-to-Dense’ , using their open-source code and the best performed model trained by their authors.

Firstly, we vary the LiDAR input with 5 different levels of density on the KITTI validation set. The KITTI dataset is captured with a 64-line Velodyne LiDAR. However, real industrial applications may only adopt a 32-line or even 16-line LiDAR considering the high sensor cost. To analyze the impact of the sparsity level on the final result, we test with 5 different levels of LiDAR density on the KITTI validation dataset, where the input LiDAR points are randomly sampled according to a given ratio. Specifically, the density ratios of 0.20.2, 0.40.4, 0.60.6, 0.80.8 and 1.01.0 are adopted in our evaluation.

Figure 6 shows the RMSE of our network, ‘NConv-CNN’ and ‘Sparse-to-Dense’ under various LiDAR point densities. With the density decreasing, the ‘NConv-CNN’ shows significant performance drop and its RMSE increases quickly. In comparison, our method and the ‘Sparse-to-Dense’ , on the other hand, degrade gradually and are consistently better than the ‘NConv-CNN’ . The results demonstrate the strong generalization capability of our method under various LiDAR points density ratios.

IV-F2 Various Lighting and Weather Conditions

KITTI dataset is collected in the similar lighting condition and in good weather condition. However, varied weather and lighting conditions always occur in practice and may bring the potential impact on the performance of depth completion. To verify whether our guided network can still work well in these kinds of challenging situations, we conduct evaluation experiments on Virtual KITTI dataset with various lighting (e.g., sunset) and weather (e.g., fog) conditions, and compare our method with other two variants of ‘Add.’ and ‘Concat’ introduced in subsection IV-D. Based on the trained model on KITTI dataset, we fine-tune our method under good ‘clone’ condition, then test its performance under various lighting and weather condition in a different sequence.

We evaluate our methods and two variants under the ‘clone’, ‘fog’, ‘morning’, ‘overcast’, ‘rain’ and ‘sunset’ conditions separately. Figure 7 depicts the results of three methods under various conditions. We can easily find, compared with ‘Add.’ and ‘Concat’, our method achieves the best RMSE among all the conditions. Also, the RMSE results of our method keep stable across all the situations, which can verify the generalization capability of our method under various lighting and weather conditions.

IV-F3 Cross-dataset Evaluation

In order to show the generalization of our method, we also conduct cross-dataset evaluations by using the models trained on NYUv2 dataset to directly test on SUN RGBD dataset .

The comparison results are listed in Table IV and Table V for dataset captured by Kinect V1 and Asus Xtion camera respectively. Both settings of 500 samples and 200 samples are evaluated by using the comparison models trained on NYUv2 dataset. We can see our method still outperforms other methods with the best RMSE and reports close results with NYUv2 dataset. The results demonstrate the strong cross-dataset generalization capability of our method. We also present some quantitative results in Figure 8. The first three rows selected in red rectangle are results on images captured by Kinect V1, and the last three rows in green rectangle are results from Xtion. The priority of our method can be found easily from the predicted depth, especially the selected regions.

By comparing the results in Table III, Table IV and Table V, we can find that all these three methods yield a little worse results on the dataset collected by Xtion, which may be caused by different camera intrinsic parameters and the extrinsic parameters between image sensor and depth sensor. How to design method with better generalization capability between different devices is an interesting direction for the future study.

V Conclusion

We propose a guided convolutional network to recover dense depth from sparse and irregular LiDAR points with an RGB image as guidance. Our novel guided network can dynamically predict content-dependent and spatially-variant kernel weights according to the guidance image to facilitate depth completion. We further design a convolution factorization to reduce GPU memory consumption such that our guided convolution module can be applied in powerful encoder-decoder network with multi-stage fusion scheme. Extensive experiments and ablation studies verify the superior performance of our guided convolutional network and the effectiveness of the feature fusion strategy on depth completion. Our method not only shows strong results on both indoor and outdoor scenes, but also presents strong generalization capability under different point densities, various lighting and weather conditions as well as cross-dataset evaluations. While this paper specifically focuses on the problem of depth completion, we believe that other tasks in computer vision involving multi-sources as input can also benefit from the design of our guided convolution module and the fusion scheme in our method.

References