Monocular Depth Estimation using Multi-Scale Continuous CRFs as Sequential Deep Networks
Dan Xu, Elisa Ricci, Wanli Ouyang, Xiaogang Wang, Nicu Sebe
I Introduction
While estimating the depth of a scene from a single image is a natural ability for humans, devising computational models for accurately predicting depth information from RGB data is a challenging task. Many attempts have been made to address this problem in the past. In particular, recent works have achieved remarkable performance thanks to powerful deep learning models . Assuming the availability of a large training set of RGB-depth pairs, monocular depth prediction from single images can be regarded as a pixel-level continuous regression problem and Convolutional Neural Network (CNN) architectures are typically employed.
In the last few years significant efforts have been made in the research community to improve the performance of CNN models for pixel-level prediction tasks (e.g. semantic segmentation, contour detection). Previous works have shown that, for depth estimation as well as for other pixel-level classification or regression problems, more accurate estimates can be obtained by combining information from multiple scales . This can be achieved in different ways, e.g. fusing feature maps corresponding to different network layers or designing an architecture with multiple inputs corresponding to images at different resolutions. Other works have demonstrated that, by adding a Conditional Random Field (CRF) in cascade to a convolutional neural architecture, the performance can be greatly enhanced and the CRF can be fully integrated within the deep model enabling end-to-end training with back-propagation . However, these works mainly focus on pixel-level prediction problems in the discrete domain (e.g. semantic segmentation). While complementary, so far these strategies have been only considered in isolation and no previous works have exploited multi-scale information within a CRF inference framework.
In this paper we argue that, benefiting from the flexibility and the representational power of graphical models, we can optimally fuse representations derived from multiple CNN side-output layers using structured constraints, improving performance over traditional multi-scale strategies. By exploiting this idea, we introduce a novel framework to estimate depth maps from single still images. Opposite to previous work fusing multi-scale features by weighted averaging or concatenation, we propose to integrate multi-layer side-output information by designing a novel approach based on continuous CRFs. Specifically, we present two different methods. The first approach is based on a single multi-scale unified CRF model, while the other considers a cascade of scale-specific CRFs. We also show that, by introducing a common CNN implementation for mean-fields updates in continuous CRFs, both models are equivalent to sequential deep networks and an end-to-end approach can be devised for training. Through extensive experimental evaluation we demonstrate that the proposed CRF-based approach produces more accurate depth maps than traditional multi-scale approaches for pixel-level prediction tasks . Moreover, by performing experiments on the publicly available NYU Depth V2 , Make3D and KITTI datasets, we show that our approach is able to robustly reconstruct depth with good visual quality (Fig.1) and outperforms state of the art methods for the monocular depth estimation task.
This paper extends our earlier work through proposing and investigating different multi-scale connection structures for message passing, further enriching the related works, providing more approach details, and significantly expanding experimental results and analysis. To summarize, the contribution of this paper is threefold:
Firstly, we propose a novel approach for predicting depth maps from RGB inputs which exploits multi-scale estimations derived from CNN inner semantic layers by structurally fusing them within a unified CNN-CRF framework.
Secondly, as the task of pixel-level depth prediction implies inferring a set of continuous values, we show how mean field (MF) updates can be implemented as sequential deep models, enabling end-to-end training of the whole network. We believe that our MF implementation will be useful not only to researchers working on depth prediction, but also to those interested in other problems involving continuous variables. Therefore, our code is made publicly available at https://github.com/danxuhk/ContinuousCRF-CNN.git.
Thirdly, our experiments demonstrate that the proposed multi-scale CRF framework is superior to previous methods integrating information from different semantic network layers by combining multiple losses or by adopting feature concatenations . We also show that our approach outperforms state of the state of the art monocular depth estimation methods on public benchmarks and that the proposed CRF-based models can be employed in combination with different pre-trained CNN architectures, consistently enhancing their performance.
The remainder of this paper is organised as follows. We first introduce related work in Section II, and then the proposed multi-scale CRF models for monocular depth estimation is presented in Section III. We further elaborate how the proposed models can be implemented as sequential neural network for end-to-end joint optimization in Section IV. The experimental results and analysis are elaborated in Section V, and we conclude the paper in Section VI.
II Related work
Our approach is built upon recent successes of deep CNN architectures for image classification and fully convolutional networks for dense semantic image segmentation . We briefly introduce the most related works by organizing them into three main aspects, i.e. monocular depth estimation, multi-scale CNN and dense pixel-level prediction via combination of CNN and CRFs.
Monocular depth estimation. Previous approaches for depth estimation from single images can be grouped into three main categories: (i) methods operating on hand crafted features, (ii) methods based on graphical models and (iii) methods adopting deep convolutional neural networks.
Earlier works addressing the depth prediction task belong to the first category. Hoiem et al. proposed photo pop-up, a fully automatic method for creating a basic 3D model from a single photograph by introducing an assumption of ‘ground-vertical’ geometric structure. Karsch et al. developed Depth Transfer, a non parametric approach based on SIFT Flow, where the depth of an input image is reconstructed by transferring the depth of multiple similar images and then applying some warping and optimizing procedures. Instead of directly recovering depth from appearance features, Liu et al. explored using semantic scene segmentation results to guide the 3-D depth reconstruction. Similarly, Ladicky et al. also demonstrated the benefit of combining semantic object labels with depth features. However, the hand-crafted representations are not robust enough for this challenging problem.
In the second category, some works exploited the flexibility of graphical models to reconstruct depth information. For instance, Delage et al. proposed a dynamic Bayesian framework for recovering 3D information from indoor scenes. A discriminatively-trained multiscale Markov Random Fields (MRFs) were introduced in , in order to optimally fuse local and global features. Depth estimation was treated as an inference problem in a discrete-continuous CRF model in . However, these works did not employ deep networks.
More recent approaches for depth estimation are based on CNNs . For instance, Eigen et al. proposed a multi-scale approach for depth prediction, considering two deep networks, one performing a coarse global prediction based on the entire image, and the other refining predictions locally. This approach was extended in to handle multiple tasks (e.g. semantic segmentation, surface normal estimation). Wang et al. introduced a CNN for joint depth estimation and semantic segmentation. The obtained estimates were further refined with Hierarchical CRFs. The most similar work to ours is , where the representational power of deep CNN and continuous CRFs is jointly exploited for depth prediction. However, the method proposed in is based on superpixels and the information associated to multiple scales is not exploited in their graphical model.
Multi-Scale CNNs. The problem of combining information from multiple scales has recently received considerable interest in various computer vision tasks. In a deeply supervised fully convolutional neural network was proposed for edge detection by weighted combination of multiple side outputs. Skip-layer networks, where the feature maps derived from different semantic layers of a primary front-end network are jointly considered in an output layer, have also become very popular . Other works considered multi-stream architectures, where multiple parallel networks receiving inputs at different scale are fused . Cai et al. proposed a multi-scale method via combining the predictions obtained from feature maps with different resolution for object detection. Dilated convolutions (e.g. dilation or à trous) have been also employed in different deep network models in order to aggregate multi-scale contextual information . However, in these works, the multi-scale representations or estimations are typically combined by using simple concatenation or weighted averaging operation. We are not aware of previous works exploring fusing deep multi-scale information within a CRF framework.
Dense pixel-level prediction via combination of CNN and CRFs. The combination of CNN and CRFs has shown great usefulness for dense pixel-level structured prediction . Some existing works utilize CRFs as a post processing module for further refining the predictions from the CNN . To benefit from end-to-end learning, Zhang et al. proposed a CRF-RNN model which jointly optimizes a front-end deep network with a discrete CRF for semantic image segmentation. Xu et al. proposed an attention-gated deep CRF framework for pixel-level contour prediction. However, as far as we know, this work is a first attempt to combine multi-scale continuous CRFs with deep convolutional neural network for constructing a unified model for end-to-end monocular depth estimation.
III Multi-Scale CRF Models for Monocular Depth Estimation
In this section we introduce our deep model with the designed multi-scale continuous CRFs for monocular depth estimation from RGB images. We first formalize the problem of depth prediction and give a brief overview of the proposed approach. Then, we describe two different variants of the proposed multi-scale model, one based on a cascade of CRFs and the other on a single multi-scale unified CRFs.
Following previous works we formulate the task of depth prediction from monocular RGB input as the problem of learning a non-linear mapping from the image space to the output depth space . More formally, let be a training set of pairs, where denotes an input RGB image with pixels and represents its corresponding real-valued depth map.
For learning we consider a deep model made of two main building blocks (Fig. 2). The first component is a CNN architecture with a set of intermediate side outputs , , produced from different layers with a mapping function . For simplicity, we denote with the set of front-end network layer parameters and with the parameters of the network branch producing the side output associated to the -th layer (see Section V-A for details of our implementation). In the following we denote this network as the front-end CNN.
The second component of our model is a fusion block. As shown in previous works , features generated from different CNN layers capture complementary information. The main idea behind the proposed fusion block is to use CRFs to effectively integrate the side output maps of our front-end CNN for robust depth prediction. Our approach develops from the intuition that these representations can be combined within a sequential framework, i.e. performing depth estimation at a certain scale and then refining the obtained estimates in the subsequent level. Specifically, we introduce and compare two different multi-scale models, both based on CRFs, and corresponding to two different versions of the fusion block. The first model is based on a single multi-scale unified CRFs, which integrates information available from different scales and simultaneously enforces smoothness constraints between the estimated depth values of neighboring pixels and neighboring scales. The second model implements a cascade of scale-specific CRFs: at each scale a CRF is employed to recover the depth information from side output maps and the outputs of each CRF model are used as additional observations for the subsequent model. In Section III-B1 we describe the two models in details, while in Section IV we show how they can be implemented as sequential deep networks by stacking several elementary blocks. We call these blocks C-MF blocks, as they implement Mean Field updates for Continuous CRFs (Fig. 2).
III-B Multi-scale Fusion with Continuous CRFs
We now elaborate the proposed CRF-based models for fusing multi-scale side-outputs derived from different semantic layers of the front-end deep convolutional neural networks.
Given a vector with a dimension of obtained by concatenating the side output score maps and a vector with a dimension of expressing real-valued output variables, we define a CRF modeling the following conditional distribution:
where is the partition function acting as a normalization factor for probabilities. The energy function is defined as:
and indicates the hidden variable associated to scale and pixel . The first term is the sum of quadratic unary terms defined as:
where is the regressed depth value at pixel and scale obtained with . The second term is the sum of pairwise potentials describing the relationship between pairs of hidden variables and and is defined as follows:
where is a weight which specifies the relationship between the estimated depth of the pixels and at scale and , respectively; is the number of kernels.
To perform inference we rely on the mean-field theory to approximate with another distribution , where , expressing a product of independent marginals. By minimizing the Kullback-Leibler divergence between the distribution of and , we obtain the solution of . As the log distribution has a quadratic form w.r.t. and can be represented as Gaussian distribution, the following mean-field updates can be derived:
Here and are the variance and mean of the distribution , respectively.
To define the weights we introduce the following assumptions. First, we assume that the estimated depth at scale only depends on the depth estimated at previous scale. Second, for relating pixels at the same and at previous scale, we set weights depending on kernel functions , which consists of Gaussian kernels with form of . Here, and indicate some features derived from the input image for pixels and . are user-defined bandwidth parameters . Following previous works , we use pixel positions and color values as features, leading to two kernel functions, i.e. a bilateral appearance kernel using both the pixel positions and the color value features and a spatial smoothness kernel using only the pixel positions features, for modeling dependencies of pixels at scale and other two for relating pixels at neighboring scales. Under these assumptions, the mean-field updates (5) and (6) can be rewritten as:
III-B2 Multi-Scale Cascade CRF Model
The unary and pairwise terms can be defined analogously to the above-introduced unified multi-scale model. In particular the unary term, reflecting the similarity between the observation and the hidden depth value , is:
where we consider Gaussian kernels, one for appearance features, and the other accounting for pixel positions. Similar to the multi-scale CRF model, under mean-field approximation, the following updates can be derived:
At the test time, we use the estimated depth variables corresponding to the cascade CRF model of the finest scale as our final predicted depth map .
IV Multi-Scale Models as Sequential Deep Networks
In this section, we describe how the two proposed CRFs-based models can be implemented as sequential deep networks, enabling end-to-end training of our whole deep network model (the front-end CNN and the fusion module). We first show how the mean-field iterations derived for the multi-scale and the cascade models can be implemented by designing a common structure, the continuous mean-field updating (C-MF) block, consisting into stack of a series of CNN operations. Then, we present the resulting sequential network structures and details of the training phase for optimizing the whole deep network.
By analyzing the two proposed CRF models, we can observe that the mean-field updates derived for the cascade and for the multi-scale models share common terms. As stated above, the main difference between the two is the way the estimated depth at previous scale is handled at the current scale. In the multi-scale CRFs, the relationship among neighboring scales is modeled in the hidden variable space, while in the cascade CRFs the depth estimated at previous scale acts as an observed variable.
In the cascade CRF model, differently from the multi-scale unified CRF model, acts as an observed variable. To design a common C-MF block among the two models, we introduce two gate functions G1 and G2 (Fig. 4) controlling the computing flow and allowing to easily switch between the two approaches. Both gate functions accept a user-defined boolean parameter. In our setting, the value 1 corresponds to the multi-scale CRF and the value 0 corresponds to the cascade model. Specifically, if G1 is equal to 1, the gate function G1 passes to the Gaussian filtering block, otherwise passes it to the element-wise addition block with the computed message. Similarly, G2 controls the computation of the normalization terms and switches between the computation of Eq. (7) and Eq. (12). In other words, if G2 equals to 0, then the Gaussian filtering and weighting operations for and are disabled. Importantly, for each step in the C-MF block we implement the calculation of error differentials for the back-propogation as in .
There are two different types of CRF parameters to be learned, i.e. the bandwidth parameters and the Gaussian-kernel weights . For optimizing these CRF parameters, similar to , the bandwidth values are pre-defined for simplifying the calculation, and we implement the backward differential computation for the weights of Gaussian kernels . In this way are learned automatically with back-propagation.
IV-B From Mean-Field Updates to Sequential Deep Networks
Fig. 4 illustrates the implementation of the proposed two CRF-based models using the designed C-MF block described above. In the figure, each blue-dashed box is associated to a mean-field iteration. The cascade model as shown in Fig. 5(b) consists of single-scale CRFs. At the -th scale, mean-field iterations are performed and then the estimated depth outputs are passed to another CRF model of the subsequent scale after a Rectified Linear Unit (ReLU) operation. The ReLU used here has two aspects of consideration: first the depth predictions should be always positive, and second we want to increase the nonlinearity of the sequential network for better mapping. To implement a single-scale CRF, we stack C-MF blocks and make them share the parameters, while we learn different parameters for different CRFs. For the multi-scale model, one full mean-field update involves scales simultaneously, obtained by combining C-MF blocks. We further stack iterations for learning and inference. The parameters corresponding to different scales and different mean-field iterations are shared. In this way, by using the common C-MF layer, we implement the two proposed multi-scale continuous CRFs models as deep sequential networks enabling end-to-end training with the front-end network.
IV-C Multi-Scale Message Passing Structures
The proposed work aims at multi-scale structured fusion and prediction, the connection structure between the different multi-scale predictions for message passing plays an important role in the performance. In this section, we thus propose and investigate different message passing structures. Fig. 3 illustrates several structures include top down structure, skip-connection structure and all to one structure. The top down structure is similar to the bottom up structure depicted in Fig. 2, which gradually refines the score maps from coarse to fine. The skip connection structure aims at utilizing more complementary information via skipping scales. The all to one structure uses all the other scales to refine the finest scale. Since all the message passing structures involve two scales at each time, we are able to build all these proposed connection structures by using the proposed aforementioned neural-network implemented C-MF block. The experimental investigation of these structures is illustrated in the experimental part.
IV-D Optimization of The Whole Network
We train the whole network using a two phase scheme. In the first phase (pretraining), the parameters of the base front-end network and the parameters of the side-output generation sub-branch networks are learned by minimizing the sum of distinct side losses as in , corresponding to side outputs. We define the optimization objective using a square loss over training samples as follows:
When the whole network optimization is finished, the test can be performed end-to-end, i.e. given a test RGB image as input the network directly outputs an estimated depth map.
V Experiments
To demonstrate the effectiveness of the proposed multi-scale CRF models for monocular depth prediction, we performed experiments on three publicly available datasets: the NYU Depth V2 , the Make3D and the KITTI datasets. In the following we first describe the experimental setup and the implementation details, and then present the experimental results and analysis.
The NYU Depth V2 dataset contains 120K unique pairs of RGB and depth images captured with a Microsoft Kinect. The datasets consists of 249 scenes for training and 215 scenes for testing. The images have a resolution of . To speed up the training phase, following previous works we consider only a small subset of images. This subset has 1449 aligned RGB-depth pairs: 795 pairs are used for training, 654 for testing. Following , we perform data augmentation for the training samples. The RGB and depth images are scaled with a ratio and the depths are divided by . Additionally, we horizontally flip all the samples and randomly crop them to pixels. The data augmentation phase produces 4770 training pairs in total.
The Make3D dataset contains 534 RGB-depth pairs, split into 400 pairs for training and 134 for testing. We resize all the images to a resolution of as done in to preserve the aspect ratio of the original images. We adopted the same data augmentation scheme used for NYU Depth V2 dataset but, for we randomly generate two samples each via cropping, obtaining 4K training samples.
The KITTI dataset is built for various computer vision tasks within the context of autonomous driving, which contains depth videos captured through a LiDAR sensor deployed on a driving vehicle. For the training and testing split, we follow the protocol made by Eigen et al. for a better comparison with existing works. Specifically, 61 scenes are selected from the raw data. Total 22,600 images from 32 scenes are used for training, and 697 images from the other 29 scenes are used for testing. Following , the ground-truth depth maps are generated by reprojecting the 3D points collected from velodyne laser into the left monocular camera. The resolution of RGB images are reduced half from original for training and testing.
V-A2 Evaluation Metrics
Following previous works , we adopt the following evaluation metrics to quantitatively assess the performance of our depth prediction model. Specifically, we consider:
scale invariant rms log error as used in , rms(sc-inv.);
V-B Implementation Details
We implemented the proposed deep model using the popular Caffe framework on a single Nvidia Tesla K80 GPU with 12 GB memory. More details on the front-end CNN architectures, the generation of multi-scale side outputs and the parameter settings are elaborated as follows.
To study the influence of the frond-end CNN, we consider several network architectures including: (i) AlexNet , (ii) VGG-16 , (iii) a fully convolutional encoder-decoder network derived from VGG-16, referred as VGG-ED , (iv) a Convolution-Deconvolution network based on VGG-16, referred as VGG-CD , and (v) ResNet-50 . For AlexNet, VGG-16 and ResNet-50, we obtain the side outputs from the last semantic convolutional layer of different convolutional blocks, in which each the layer produces feature maps with the same shape. The scheme utilized for the generation will be introduced in the next section. The number of side outputs considered in our experiments is 5, 5 and 4 for AlexNet, VGG-16 and ResNet-50, respectively. As VGG-ED and VGG-CD have been widely used for dense pixel-level prediction tasks, we also investigate them in the experimental analysis. Both VGG-ED and VGG-CD have a symmetric network structure, and five side outputs are then generated from the different blocks of the decoder or the deconvolutional network part.
V-B2 Generation of multi-scale CNN side-outputs
Our approach can be applied with any multi-scale front-end CNN models including those with skip-connections. We here briefly describe the scheme we adopt to build CNN side outputs from the front-end CNN for the multi-scale fusion with CRFs. In a convolutional layer is first used to generate a score map from the feature map and then a deconvolutional (deconv) layer is adopted as a bilateral upsampling operator to enlarge the score map such as to obtain the same size of the input image. However, we noticed that by adopting the approach in the generated side outputs associated to the feature maps with smaller size are very coarse, causing a lot scene details missing. To address this problem, after the convolutional layer, we stack several deconv layers, each of them enlarging the output map by two times. A Rectified Linear Unit (ReLU) is applied after each deconv layer. After the last deconv layer we use a crop layer to cut the extra margin and obtain a side output with the same resolution of the ground-truth image. We employ this scheme to obtain side outputs for AlexNet, VGG-16 and ResNet-50, while for VGG-CD and VGG-ED, we use the same setting as in , as their decoder or deconvolutional part is able to obtain more fine-grained side outputs. Table I shows detailed network parameters used to obtain the side output from the last convolutional block of ResNet-50 (i.e. from the layer res5c).
V-B3 Parameters settings
As described in Section IV-D, training consists of a pretraining and a fine tuning phase. In the first phase, we train the front-end CNN with parameters initialized with the corresponding ImageNet pretrained models. For AlexNet, VGG-16, VGG-ED and VGG-CD, the batch size is set to 12 and for ResNet-50 to 8. The learning rate is initialized at and decreases by 10 times around every 50 epochs. 80 epochs are performed for pretraining in total. The momentum and the weight decay are set to 0.9 and 0.0005, respectively. When the pretraining is finished, we connect all the side outputs of the front-end CNN to our CRFs-based multi-scale deep models for end-to-end training of the whole network. In this phase, the batch size is reduced to 6 and a fixed learning rate of is used. The same parameters of the pre-training phase are used for momentum and weight decay. The bandwidth weights for the Gaussian kernels are obtained through cross validation. The number of mean-field iterations is set to 5 for efficient training for both the cascade CRFs and multi-scale CRFs. We do not observe significant improvement using more than 5 iterations. Training the whole network takes around hours on the Make3D dataset, hours on the KITTI dataset and hours on the NYU v2 dataset.
V-C Experimental Results
To present the experimental results, we start from an ablation study for investigating the performance impact of different front-end network architectures, the effectiveness of the proposed CRF-based multi-scale fusion models and the influence of the stacking orders for making the sequential neural network. Then we compare the overall performance with the state of the art methods, and finally the qualitative results and running time are analyzed.
As discussed above, the proposed multi-scale CRF-based fusion models are general and different deep architectures can be used for the front-end network. In this section we evaluate the impact of this choice on the depth estimation performance. We consider both the case of the pretrained front-end models (i.e. only side losses are employed but the multi-scale CRF models are not plugged), indicated with ‘pretrain’, and the case of the fine-tuned models, including the front-end network with the multi-scale cascade CRFs (cascade-CRFs). The results of the experiments are shown in Table II. As expected, in both cases deeper CNN architectures produced more accurate predictions, and ResNet-50 achieves the best performance among all the front-end networks. Moreover, VGG-CD is slightly better than VGG-ED, and both these models outperforms VGG-16, showing that the symmetric network structure is beneficial for the dense pixel-level prediction problems. Importantly, for all considered front-end networks there is a significant increase in performance when applying the proposed CRF-based models.
Figure 6 depicts some examples of predicted depth maps using different front-end networks on the NYU Depth V2 test dataset. As we can see from the figure, the qualitative results confirm that the deeper architecture leads to better depth recovery. By comparing the reconstructed depth maps obtained with pretrained models (e.g. using only the front-end networks VGG-CD and ResNet-50) with those generated with our multi-scale models, it is clear that our approach remarkably improves prediction accuracy and visual quality.
V-C2 Evaluation of different multi-scale CRF fusion models
To evaluate the effectiveness of the proposed CRF-based multi-scale fusion models, we conduct experiments on the NYU Depth V2 dataset and consider the following baselines:
(i) the ‘HED’ method in , where multiple side outputs are fused with a weighted averaging scheme and the sum of multiple side output losses is jointly minimized as deep supervision with a cross-entropy loss, while we use the square loss as our problem involves continuous variables;
(ii) the ‘Hypercolumn’ method , where multi-scale feature maps generated from different semantic network layers are concatenated and fused;
(iii) a continuous CRF (‘C-CRF’) applied on the prediction of the front-end network, i.e. plugging after the last output layer as a post-processing module without end-to-end training.
For the first two baselines, we want to compare our models with other popular methods for fusing multi-scale CNN information, while the third one aims at demonstrating the effectiveness of the continous CRF itself. In these experiments we consider VGG-CD as the front-end CNN architecture. The results of the comparison are shown in Table III. It is evident that with our CRF-based fusion models (both the cascade CRFs and the unified CRFs) more accurate depth maps can be obtained, demonstrating that our idea of integrating complementary information derived from CNN side output maps within a graphical model framework is more effective than traditional fusion schemes. Table III also compares the proposed cascade and unified models. As expected, the unified model produces more accurate depth maps, at the price of an increased computational cost. This can also be observed from Table II. The C-CRF (in Table III) improves the depth estimation at all metrics over the VGG-CD (pretrain) (in Table II) with a clear gap, showing the CRF model is very useful for refining the deeply predicted map. By jointly learning with the front-end (i.e. end-to-end training), ours (single-scale) further boosts the performance. Finally, we analyze the impact of adopting multiple scales and compare our complete models (5 scales) with their version when only a single and three side output layers are used. It is evident that the performance can be improved by increasing the number of scales.
V-C3 Evaluation of multi-scale message passing structures
We evaluate the influence of different multi-scale message passing structures using the cascade CRF model. Four connection structures as depicted in Fig. 3 are compared. Table IV shows the monocular depth estimation results on NYUD-v2 dataset. The comparison results confirm that the message passing structure indeed has an impact on the final performance. The bottom up and top down structures have similar performance, while the skip-connection structure slightly outperform these two. The all to one structure performs the best, producing around 2.0% gain in terms of the rel metric than the top down structure, which means that directly passing message to the finest prediction scale from the rest scales can absorb more complementary information than the gradual passing fashions used in the first three structures.
V-C4 Comparison with state of the art
We also compare our approach with state of the art methods on all the datasets. For previous works we directly report results taken from the original papers. Table V shows the results of the comparison on the NYU Depth V2 dataset. For our approach we consider the cascade model and use two different training sets for pretraining: the small set of 4.7K pairs employed in all our experiments and a larger set of 95K images as in . Note that for fine tuning we only use the small set. As shown in the table, our approach outperforms all competing methods and it is the second best model when we use only 4.7K images. This is remarkable considering that, for instance, in 120K image pairs are used for training. Our model achieves the best results on all the metrics via using 95K pretraining samples and using the proposed all to one message passing structure.
We also perform a comparison with several state of the art methods on the Make3D dataset (Table VI). Following , the error metrics are computed in two different settings, i.e. considering (C1) only the regions with ground-truth depth less than 70 and (C2) the entire image. It is clear that the proposed approach is significantly better than previous methods. In particular, comparing with Laina et al. , the best performing method in the literature, it is evident that our approach, both in case of the cascade and the multi-scale models, outperforms by a significant margin when Laina et al. also adopt a square loss. It is worth noting that in a training set of 15K image pairs is considered, while we employ much less training samples. By increasing our training data (i.e. K in the pretraining phase), our multi-scale CRF model also outperforms with Huber loss (log10 and rms metrics). The final performance is further boosted by considering the all to one structure similar to NYUD v2 dataset. Finally, it is very interesting to compare the proposed method with the approach in Liu et al. , since in a CRF model is also employed within a deep network trained end-to-end. Our method significantly outperforms in terms of accuracy. Moreover, in a time of 1.1sec is reported for performing inference on a test image but the time required by superpixels calculations is not taken into account. Oppositely, with our method computing the depth map for a single image takes about 1 sec in total.
The state of the art comparison on KITTI dataset is shown in Table VII. The competitors include Saxena et al. , Eigen et al. , Liu et al. , Zhou et al. , Garg et al. , Godard et al. and Kuznietsov et al. . As the same setting of ours, the first four methods use single monocular images in the training phase, while the last two considered two monocular images with a stereo setting for training. Among the first four competitors, Eigen et al. significantly outperforms the others in terms of the metric of the mean relative error (rel), due to the usage of large-scale training data (more than 1 million samples). While our model achieves much better performance than Eigen et al. in all metrics with much less data (22.6K samples). Although the training of the last two methods (requiring two monocular images) is not equal to our setting, the proposed approach with both the bottom-up and the all to one structures still produces better results than them with clear performance gap in all metrics. Kuznietsov et al. reports results for both the stereo training and the monocular supervised training. It is not directly comparable with the stereo training setting, which is significantly different as it requires both left and right images from a binocular camera. Ours focuses on monocular depth estimation and achieves lower error performance comparing with theirs using the same monocular setting. Fig. 8 also shows some qualitative comparison results with these methods, further demonstrating the advantageous performance of our approach.
V-C5 Qualitative depth estimation results
Fig. 6, 7 and 9 show some examples of the qualitative depth estimation results and the comparison with the competing methods on the NYUD-V2, Make3D and KITTI dataset respectively. It is clear that the proposed approach is able to produce sharper depth estimation with better visual quality compared with the classic CNN structures, which demonstrates the importance of the prediction aided by the CRFs with appearance and smoothness constraints. Fig. 9 also shows a qualitative comparison between the pretrained front-end CNN and the fine-tuned whole model. It can be observed that our approach can recover more scene structures and details. We believe that this is probably because the effective structured fusion of the coarse-to-fine multi-scale predictions of the deep network with the proposed CRF models. For the influence of the variance in the CRF model on the prediction errors, as the variance term is actually acted as a normalization factor after the message passing. It may have influence but the main influence is dominated by the predictions of deep front-end CNN based on our observation from the experimental results.
V-C6 Empirical run-time analysis
Computational run-time complexity is an important aspect for deep structured prediction models. In this paragraph we provide a short discussion about the computational cost of the proposed CRFs-based models. As shown in the paper, the multi-scale CRF model achieves better accuracy and lower error than the cascade model for both the NYU Depth V2 and the Make3D experiments. However, as expected, the cascade model is more advantageous in terms of the running time. For instance, considering ResNet-50 as the front-end CNN, the time required at test phase for one image is seconds w.r.t. the cascade model and seconds w.r.t. the multi-scale model, and the image resolution is pixels. Higher resolution of the network input usually brings more computational overhead. We also test the running time given the input resolution of and it costs around seconds for processing one image. We believe that if we reduce the receptive field of the CRF model from fully connected to partially connected, the computing time could be significantly reduced.
VI Conclusion
In this paper, we introduced a novel approach for predicting depth maps from a single RGB image. The core of the method is a novel framework based on continuous CRFs for fusing multi-scale score-level side-outputs derived from different semantic CNN layers. We demonstrated that this framework can be used in combination with several common CNN architectures and can be implemented for end-to-end training. The extensive experiments confirmed the validity of the proposed multi-scale fusion approach. While this paper specifically addresses the problem of depth prediction, we believe that other tasks in computer vision involving pixel-level predictions of continuous variables, can also benefit from our implementation of the mean-field updating within the CNN framework.
Currently, the multi-scale fusion is performed on the score level. Further research direction will investigate the integration of both the feature- and the score-level multi-scale information within a unified graphical model. Moreover, the study of strategies for further improving the training and testing efficiency of the CNN-CRF models will also be an interesting aspect in the future work. The monocular depth estimation is particularly useful for various cross-modal recognition and detection tasks. A straightforward follow-up of this work would be designing a joint multi-task deep model to transfer the learned depth model for aiding other similar dense prediction problems such as contour detection and semantic segmentation.