High Quality Monocular Depth Estimation via Transfer Learning

Ibraheem Alhashim, Peter Wonka

Introduction

Depth estimation from 2D images is a fundamental task in many applications including scene understanding and reconstruction Lee2011; moreno2007active; Hazirbas2016FuseNetID. Having a dense depth map of the real-world can be very useful in applications including navigation and scene understanding, augmented reality Lee2011, image refocusing moreno2007active, and segmentation Hazirbas2016FuseNetID. Recent developments in depth estimation are focusing on using convolutional neural networks (CNNs) to perform 2D to 3D reconstruction. While the performance of these methods has been steadily increasing, there are still major problems in both the quality and the resolution of these estimated depth maps. Recent applications in augmented reality, synthetic depth-of-field, and other image effects Hedman2018; Cao2018; Wang2018 require fast computation of high resolution 3D reconstructions in order to be applicable. For such applications, it is critical to faithfully reconstruct discontinuity in the depth maps and avoid the large perturbations that are often present in depth estimations computed using current CNNs.

Based on our experimental analysis of existing architectures and training strategies Eigen2014; Li2015; Laina2016; Xu2017; Fu2018DeepOR we set out with the design goal to develop a simpler architecture that makes training and future modifications easier. Despite, or maybe even due to its simplicity, our architecture produces depth map estimates of higher accuracy and significantly higher visual quality than those generated by existing methods (see Fig. 1). To achieve this, we rely on transfer learning were we repurpose high performing pre-trained networks that are originally designed for image classification as our deep features encoder. A key advantage of such a transfer learning-based approach is that it allows for a more modular architecture where future advances in one domain are easily transferred to the depth estimation problem.

Our contributions are threefold. First, we propose a simple transfer learning-based network architecture that produces depth estimations of higher accuracy and quality. The resulting depth maps capture object boundaries more faithfully than those generated by existing methods with fewer parameters and less training iterations. Second, we define a corresponding loss function, learning strategy, and simple data augmentation policy that enable faster learning. Third, we propose a new testing dataset of photo-realistic synthetic indoor scenes, with perfect ground truth, to better evaluate the generalization performance of depth estimating CNNs.

We perform different experiments on several datasets to evaluate the performance and quality of our depth estimating network. The results show that our approach not only outperforms the state-of-the-art and produces high quality depth maps on standard depth estimation datasets, but it also results in the best generalization performance when applied to a novel dataset.

Related Work

The problem of 3D scene reconstruction from RGB images is an ill-posed problem. Issues such as lack of scene coverage, scale ambiguities, translucent or reflective materials all contribute to ambiguous cases where geometry cannot be derived from appearance. In practice, the more successful approaches for capturing a scene’s depth rely on hardware assistance, e.g. using laser or IR-based sensors, or require a large number of views captured using high quality cameras followed by a long and expensive offline reconstruction process. Recently, methods that rely on CNNs are able to produce reasonable depth maps from a single or couple of RGB input images at real-time speeds. In the following, we look into some of the works that are relevant to the problem of depth estimation and 3D reconstruction from RGB input images. More specifically, we look into recent solutions that depend on deep neural networks.

has been considered by many CNN methods where they formulate the problem as a regression of the depth map from a single RGB image Eigen2014; Laina2016; Xu2017; Hao2018DetailPD; Xu2018StructuredAG; Fu2018DeepOR. While the performance of these methods have been increasing steadily, general problems in both the quality and resolution of the estimated depth maps leave a lot of room for improvement. Our main focus in this paper is to push towards generating higher quality depth maps with more accurate boundaries using standard neural network architectures. Our preliminary results do indicate that improvements on the state-of-the-art are possible to achieve by leveraging existing simple architectures that perform well on other computer vision tasks.

Multi-view

stereo reconstruction using CNN algorithms have been recently proposed Huang2018DeepMVSLM. Prior work considered the subproblem that looks at image pairs Ummenhofer2017, or three consecutive frames Godard2018DiggingIS. Joint key-frame based dense camera tracking and depth map estimation was presented by Zhou2018DeepTAMDT. In this work, we seek to push the performance for single image depth estimation. We suspect that the features extracted by monocular depth estimators could also help derive better multi-view stereo reconstruction methods.

Transfer learning

approaches have been shown to be very helpful in many different contexts. In recent work, Zamir et al. investigated the efficiency of transfer learning between different tasks Zamir2018TaskonomyDT, many of which were are related to 3D reconstruction. Our method is heavily based on the idea of transfer learning where we make use of image encoders originally designed for the problem of image classification huang2017densely. We found that using such encoders that do not aggressively downsample the spatial resolution of the input tend to produce sharper depth estimations especially with the presence of skip connections.

Encoder-decoder

networks have made significant contributions in many vision related problems such as image segmentation Ronneberger2015u, optical flow estimation Dosovitskiy2015, and image restoration LehtinenMHLKAA18. In recent years, the use of such architectures have shown great success both in the supervised and the unsupervised setting of the depth estimation problem Godard2017; Ummenhofer2017; Huang2018DeepMVSLM; Zhou2018DeepTAMDT. Such methods typically use one or more encoder-decoder network as a sub part of their larger network. In this work, we employ a single straightforward encoder-decoder architecture with skip connections (see Fig. 2). Our results indicate that it is possible to achieve state-of-the-art high quality depth maps using a simple encoder-decoder architecture.

Proposed Method

In this section, we describe our method for estimating a depth map from a single RGB image. We first describe the employed encoder-decoder architecture. We then discuss our observations on the complexity of both encoder and decoder and its relation to performance. Next, we propose an appropriate loss function for the given task. Finally, we describe efficient augmentation policies that help the training process significantly.

Fig. 2 shows an overview of our encoder-decoder network for depth estimation. For our encoder, the input RGB image is encoded into a feature vector using the DenseNet-169 huang2017densely network pretrained on ImageNet Deng2009. This vector is then fed to a successive series of up-sampling layers LehtinenMHLKAA18, in order to construct the final depth map at half the input resolution. These upsampling layers and their associated skip-connections form our decoder. Our decoder does not contain any Batch Normalization Ioffe2015BNA or other advanced layers recommended in recent state-of-the-art methods Fu2018DeepOR; Hao2018DetailPD. Further details about the architecture and its layers along with their exact shapes are described in the appendix.

Complexity and performance.

The high performance of our surprisingly simple architecture gives rise to questions about which components contribute the most towards achieving these quality depth maps. We have experimented with different state-of-the-art encoders Bianco2018, of more or less complexity than that of DenseNet-169, and we also looked at different decoder types Laina2016; Wojna2017TheDI. What we experimentally found is that, in the setting of an encoder-decoder architecture for depth estimation, recent trends of having convolutional blocks exhibiting more complexity do not necessarily help the performance. This leads us to advocate for a more thorough investigation when adopting such complex components and architectures. Our experiments show that a simple decoder made of a 2×2\times bilinear upsampling step followed by two standard convolutional layers performs very well.

2 Learning and Inference

A standard loss function for depth regression problems considers the difference between the ground-truth depth map yy and the prediction of the depth regression network y^\hat{y} Eigen2014. Different considerations regarding the loss function can have a significant effect on the training speed and the overall depth estimation performance. Many variations on the loss function employed for optimizing the neural network can be found in the depth estimation literature Eigen2014; Laina2016; Ummenhofer2017; Fu2018DeepOR. In our method, we seek to define a loss function that balances between reconstructing depth images by minimizing the difference of the depth values while also penalizing distortions of high frequency details in the image domain of the depth map. These details typically correspond to the boundaries of objects in the scene.

For training our network, we define the loss LL between yy and y^\hat{y} as the weighted sum of three loss functions:

The first loss term LdepthL_{depth} is the point-wise L1 loss defined on the depth values:

The second loss term LgradL_{grad} is the L1 loss defined over the image gradient g\boldsymbol{g} of the depth image:

Lastly, LSSIML_{SSIM} uses the Structural Similarity (SSIM) Wang2004SSIM term which is a commonly-used metric for image reconstruction tasks. It has been recently shown to be a good loss term for depth estimating CNNs Godard2017. Since SSIM has an upper bound of one, we define it as a loss LSSIML_{SSIM} as follows:

Note that we only define one weight parameter λ\lambda for the loss term LdepthL_{depth}. We empirically found and set λ=0.1\lambda=0.1 as a reasonable weight for this term.

An inherit problem with such loss terms is that they tend to be larger when the ground-truth depth values are bigger. In order to compensate for this issue, we consider the reciprocal of the depth Ummenhofer2017; Huang2018DeepMVSLM where for the original depth map yorigy_{orig} we define the target depth map yy as y=m/yorigy=m/y_{orig} where mm is the maximum depth in the scene (e.g. m=10m=10 meters for the NYU Depth v2 dataset). Other methods consider transforming the depth values and computing the loss in the log space Eigen2014; Ummenhofer2017.

Augmentation Policy.

Data augmentation, by geometric and photo-metric transformations, is a standard practice to reduce over-fitting leading to better generalization performance krizhevsky2012imagenet. Since our network is designed to estimate depth maps of an entire image, not all geometric transformations would be appropriate since distortions in the image domain do not always have meaningful geometric interpretations on the ground-truth depth. Applying a vertical flip to an image capturing an indoor scene may not contribute to the learning of expected statistical properties (e.g. geometry of the floors and ceilings). Therefore, we only consider horizontal flipping (i.e. mirroring) of images at a probability of 0.50.5. Image rotation is another useful augmentation strategy, however, since it introduces invalid data for the corresponding ground-truth depth we do not include it. For photo-metric transformations we found that applying different color channel permutations, e.g. swapping the red and green channels on the input, results in increased performance while also being extremely efficient. We set the probability for this color channel augmentation to 0.250.25. Finding improved data augmentation policies and their probability values for the problem of depth estimation is an interesting topic for future work Cubuk2018AutoAugmentLA.

Experimental Results

In this section we describe our experimental results and compare the performance of our network to existing state-of-the-art methods. Furthermore, we perform ablation studies to analyze the influence of the different parts of our proposed method. Finally, we compare our results on a newly proposed dataset of high quality depth maps in order to better test the generalization and robustness of our trained model.

is a dataset that provides images and depth maps for different indoor scenes captured at a resolution of 640×480640\times 480 Silberman2012. The dataset contains 120K training samples and 654 testing samples Eigen2014. We train our method on a 50K subset. Missing depth values are filled using the inpainting method of Levin2004. The depth maps have an upper bound of 10 meters. Our network produces predictions at half the input resolution, i.e. a resolution of 320×240320\times 240. For training, we take the input images at their original resolution and downsample the ground truth depths to 320×240320\times 240. Note that we do not crop any of the input image-depth map pairs even though they contain missing pixels due to a distortion correction preprocessing. During test time, we compute the depth map prediction of the full test image and then upsample it by 2×2\times to match the ground truth resolution and evaluate on the pre-defined center cropping by Eigen et al. Eigen2014. At test time, we compute the final output by taking the average of an image’s prediction and the prediction of its mirror image.

KITTI

is a dataset that provides stereo images and corresponding 3D laser scans of outdoor scenes captured using equipment mounted on a moving vehicle geiger2013vision. The RGB images have a resolution of around 1241×3761241\times 376 while the corresponding depth maps are of very low density with lots of missing data. We train our method on a subset of around 26K images, from the left view, corresponding to scenes not included in the 697 test set specified by Eigen2014. Missing depth values are filled using the inpainting method mentioned earlier. The depth maps have an upper bound of 80 meters. Our encoder’s architecture expects image dimensions to be divisible by 32 huang2017densely, therefore, we upsample images bilinearly to 1280×3841280\times 384 during training. During testing, we first scale the input image to the expected resolution and then upsample the output depth image from 624×192624\times 192 to the original input resolution. The final output is computed by taking the average of an image’s prediction and the prediction of its mirror image.

2 Implementation Details

We implemented our proposed depth estimation network using TensorFlow tensorflow2015-whitepaper and trained on four NVIDIA TITAN Xp GPUs with 12GB memory. Our encoder is a DenseNet-169 huang2017densely pretrained on ImageNet Deng2009. The weights for the decoder are randomly initialized following glorot2010understanding. In all experiments, we used the ADAM jlb2015adam optimizer with learning rate 0.00010.0001 and parameter values β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999. The batch size is set to 8. The total number of trainable parameters for the entire network is approximately 42.6M parameters. Training is performed for 1M iterations for NYU Depth v2, needing 20 hours to finish. Training for the KITTI dataset is performed for 300K iterations, needing 9 hours to train.

3 Evaluation

We quantitatively compare our method against state-of-the-art using the standard six metrics used in prior work Eigen2014. These error metrics are defined as:

average relative error (rel): 1n∑pn∣yp−y^p∣y\frac{1}{n}\sum_{p}^{n}\frac{\lvert y_{p}-\hat{y}_{p}\rvert}{y};

root mean squared error (rms): 1n∑pn(yp−y^p)2)\sqrt{\frac{1}{n}\sum_{p}^{n}(y_{p}-\hat{y}_{p})^{2})};

average (log⁡10\log_{10}) error: 1n∑pn∣log⁡10(yp)−log⁡10(y^p)∣\frac{1}{n}\sum_{p}^{n}\lvert\log_{10}(y_{p})-\log_{10}(\hat{y}_{p})\rvert;

threshold accuracy (δi\delta_{i}): %\% of ypy_{p} s.t. max(ypy^p,y^pyp)=δ<thr\text{max}(\frac{y_{p}}{\hat{y}_{p}},\frac{\hat{y}_{p}}{y_{p}})=\delta<thr for thr=1.25,1.252,1.253thr=1.25,1.25^{2},1.25^{3};

where ypy_{p} is a pixel in depth image yy, y^p\hat{y}_{p} is a pixel in the predicted depth image y^\hat{y}, and nn is the total number of pixels for each depth image.

Qualitative results.

We conduct three experiments to approximately evaluate the quality of the results using three measures on the NYU Depth v2 test set. The first measure is a perception-based qualitative metric that measures the quality of the results by looking at the similarity of the resulting depth maps in image space. We do so by rendering a gray scale visualization of the ground truth and that of the predicted depth map and then we compute the mean structural similarity term (mSSIM) of the entire test dataset 1T∑iTSSIM(yi,y^i)\frac{1}{T}\sum_{i}^{T}{SSIM}(y_{i},\hat{y}_{i}). The second measure considers the edges formed in the depth map. For each sample, we compute the gradient magnitude image of both the ground truth and the predicted depth image, using the Sobel gradient operator Sobel1968, and then threshold this image at values greater than 0.5 and compute the F1 score averaged across the set. The third measure is the mean cosine distance between normal maps extracted from the depth images of the ground truth and the predicted depths also averaged across the set. Fig. 4 shows visualizations of some of these measures.

Fig. 6 shows a gallery of depth estimation results that are predicated using our method along with a comparison to those generated by the state-of-the-art. As can be seen, our approach produces depth estimations at higher quality where depth edges better match those of the ground truth and with significantly fewer artifacts.

4 Comparing Performance

In Tab. 1, the performance of our depth estimating network is compared to the state-of-the-art on the NYU Depth v2 dataset. As can be seen, our model achieves state-of-the-art on all but two quantitative metrics. Our model is able to outperform the existing state-of-the-art Fu2018DeepOR while requiring fewer parameters, 42.6M vs 110M, fewer number of training iterations, 1M vs 3M, and with fewer input training data, 50K samples vs 120K samples. A typical source of error for single image depth estimation networks is the estimated absolute scale of the scene. The last row in Tab. 1 shows that when accounting for this error, by multiplying the predicted depths by a scalar that matches the median with the ground truth Zhou2017, we are able to achieve with a good margin state-of-the-art for the NYU Depth v2 dataset on all metrics. The results in Tab. 3 show that for the same dataset our method outperforms state-of-the-art on our defined quality approximating measures. We conduct these experiments for methods with published pre-trained models and code.

In Tab. 2, the performance of our network is compared to the state-of-the-art on the KITTI dataset. Our method is the second best on all the standard metrics. We suspect that one reason our method does not outperform the state-of-the-art on this particular dataset is due to the nature of the provided depth maps. Since our loss function is designed to not only consider point-wise differences but also optimize for edges and appearance preservation by looking at regions around each point, the learning process does not converge well for very sparse depth images. Fig. 3 clearly shows that while quantitatively our method might not be the best, the quality of the produced depth maps is much better than those produced by the state-of-the-art.

5 Ablation Studies

We perform ablation studies to analyze the details of our proposed architecture. Fig. 5 shows a representative look into the testing performance, in terms of validation loss, when changing some parts of our standard model or modifying our training strategy. Note that we performed these tests on a smaller subset of the NYU Depth v2 dataset.

In this experiment we substitute the pretrained DenseNet-169 with a denser encoder, namely the DenseNet-201. In Fig. 5 (red), we can see the validation loss is lower than that of our standard model. The big caveat, though, is that the number of parameters in the network grows by more than 2×2\times. When considering using DenseNet-201 as our encoder, we found that the gains in performance did not justify the slow learning time and the extra GPU memory required.

Decoder depth.

In this experiment we apply a depth reducing convolution such that the features feeding into the decoder are half what they are in the standard DenseNet-169. In Fig. 5 (blue), we see a reduction in the performance and overall instability. Since these experiments are not representative of a full training session the performance difference in halving the features might not be as visible as we have observed when running full training session.

Color Augmentation.

In this experiment, we turn off our color channel swapping-based data augmentation. In Fig. 5 (green), we can see a significant reduction as the model tends to quickly falls into overfitting to the training data. We think this simple data augmentation and its significant effect on the neural network is an interesting topic for future work.

6 Generalizing to Other Datasets

To illustrate how well our method generalizes to other datasets, we propose a new dataset of photo-realistic indoor scenes with nearly perfect ground truth depths. These scenes are collected from the Unreal marketplace community UnrealMarket2018. We refer to this dataset as Unreal-1k. It is a random sampling of 1000 images with their corresponding depth maps selected from renderings of 32 virtual scenes using the Unreal Engine. Further details about this dataset can be found in the appendix. We compare our NYU Depth v2 trained model to two supervised methods that are also trained on the same dataset. For inference, we use the public implementations for each method. The hope of this experiment is to demonstrate how well do models trained on one dataset perform when presented with data sampled from a different distribution (i.e. synthetic vs. real, perfect depth capturing vs. a Kinect, etc.).

Tab. 1 shows quantitative comparisons in terms of the average errors over the entire Unreal-1k dataset. As can be seen, our method outperforms the other two methods. We also compute the qualitative measure mSSIM described earlier. Fig. 7 presents a visual comparison of the different predicted depth maps against the ground truth.

Conclusion

In this work, we proposed a convolutional neural network for depth map estimation for single RGB images by leveraging recent advances in network architecture and the availability of high performance pre-trained models. We show that having a well constructed encoder, that is initialized with meaningful weights, can outperform state-of-the-art methods that rely on either expensive multistage depth estimation networks or require designing and combining multiple feature encoding layers. Our method achieves state-of-the-art performance on the NYU Depth v2 dataset and our proposed Unreal-1K dataset. Our aim in this work is to push towards generating higher quality depth maps that capture object boundaries more faithfully, and we have shown that this is indeed possible using an existing architectures. Following our simple architecture, one avenue for future work is to substitute the proposed encoder with a more compact one in order to enable quality depth map estimation on embedded devices. We believe their are still many possible cases of leveraging standard encoder-decoder models alongside transfer learning for high quality depth estimation. Many questions on the limits of our proposed network and identifying more clearly the effect on performance and contribution of different encoders, augmentations, and learning strategies are all interesting to purse for future work.

References

Appendix A Appendix

Tab. 5 shows the structure of our encoder-decoder with skip connections network. Our encoder is based on the DenseNet-169 huang2017densely network where we remove the top layers that are related to the original ImageNet classification task. For our decoder, we start with a 1×11\times 1 convolutional layer with the same number of output channels as the output of our truncated encoder. We then successively add upsampling blocks each composed of a 2×2\times bilinear upsampling followed by two 3×33\times 3 convolutional layers with output filters set to half the number of inputs filters, and were the first convolutional layer of the two is applied on the concatenation of the output of the previous layer and the pooling layer from the encoder having the same spatial dimension. Each upsampling block, except for the last one, is followed by a leaky ReLU activation function Maas13LeRELU with parameter α=0.2\alpha=0.2. The input images are represented by their original colors in the range $withoutanyinputdatanormalization.Targetdepthmapsareclippedtotherangewithout any input data normalization. Target depth maps are clipped to the range[0.4,10]$ in meters.

A.2 The Unreal-1K Dataset

We propose a new dataset of photo-realistic synthetic indoor scenes having near perfect ground truth depth maps. The scenes cover categories including living areas, kitchens, and offices all of which have realistic material and different lighting scenarios. These scenes, 32 scenes in total, are collected from the Unreal marketplace community UnrealMarket2018. For each scene we select around 40 objects of interest and we fly a virtual camera around the object and capture images and their corresponding depth maps of resolution 640×480640\times 480. In all, we collected more than 20K images from which we randomly choose 1K images as our testing dataset Unreal-1k. Fig. 8 shows example images from this dataset along with depth estimations using various methods.

A.3 Additional Ablation Studies

We perform additional ablation studies to analyze more details of our proposed architecture. Fig. 9 shows a representative look into the testing performance, in terms of validation loss, when changing some parts of our standard model. The training in these experiments is performed on the NYU Depth v2 dataset Silberman2012 for 750K iterations (15 epochs).

In this experiment, we examine the effect of using an encoder that is initialized using random weights as opposed to being pre-trained on ImageNet which is what we use in our proposed standard model. In Fig. 9 (purple), we can see the validation loss is greatly increased when training from scratch. This further validates that the performance of our depth estimation is positively impacted by transfer learning.

Skip connections.

In this experiment, we examine the effect of removing the skip connections between layers of the encoder and decoder. In Fig. 9 (green), we can see the validation loss is decreased, compared to our proposed standard model, resulting in worse depth estimation performance.

Batch size.

In this experiment, we look at different values for the batch size and its effect on performance. In Fig. 9 (red and blue), we can see the validation loss for batch sizes 2 and 16 compared to our standard model (orange) with batch size 8. Setting the batch size to 8 results in the best performance out of the three values while also training for a reasonable amount of time.