Learning monocular depth estimation with unsupervised trinocular assumptions
Matteo Poggi, Fabio Tosi, Stefano Mattoccia
Introduction
Depth plays a crucial role in many computer vision applications and active 3D sensors are becoming very popular. Nonetheless, such sensors may have severe shortcomings. For instance, the Kinect 1 is not suited at all for outdoor environments flooded by sunlight. Moreover, such sensor allows only for close range depth measurements. On the other hand, a popular active depth sensor perfectly suited for outdoor environments is LIDAR (e.g., Velodyne). However, sensors based on such technology are typically expensive and often cumbersome for some practical applications. Thus, inferring depth with passive sensors based on standard imaging technology would be highly desirable being cheap, lightweight and suited for indoor and outdoor environments. In this context, acquiring images from different viewpoints allows inferring depth exploiting geometry constraints. On the other hand, estimating depth from a single image is indeed an ill-posed problem. Nonetheless, this latter approach would overcome some major constraints such as the need for simultaneous acquisition in binocular stereo or handling dynamic objects in depth-from-motion approaches. Although a geometrically ambiguous problem, Convolutional Neural Networks (CNNs) achieved outstanding results in monocular depth estimation by casting it as a learning task in both supervised and unsupervised manner. In particular, the latter paradigm addresses the hunger for data typical of deep learning tasks by training networks to produce a depth representation minimizing the warping error between images acquired from multiple points of view rather than the error with respect to difficult to source ground-truth depth labels. In this field, the work of Godard et al. represents state-of-the-art for unsupervised monocular depth estimation. Deploying stereo imagery for training, a CNN learns to infer disparity from a single reference image and warps the target image accordingly to minimize the appearance error between the warped and the reference image. This strategy yields state-of-the-art performance . The CNN is trained to infer disparity from a single reference image and the target image is warped accordingly minimizing the appearance error between warped and reference image. This way, the depth representation learned by the network is affected by artifacts in specific image regions inherited from the stereo setup (e.g., the left border using the left image as the reference) and in occluded areas. The post-processing step proposed in partially compensates for these artifacts. However, it requires a double forward of the input image and its horizontally flipped version thus obtaining two predictions with artifacts, respectively, on the left and right side of depth discontinuities. Such issues are softened in the final map at the cost of doubling processing time and memory footprint.
In this paper, we propose to explicitly take into account these artifacts training our network on imagery acquired by a trinocular setup. By assuming the availability of three horizontally aligned images at training time, our network learns to process the frame in the middle and produce inverse depth (i.e., disparity) maps according to all the available viewpoints. By doing so, we can attenuate the aforementioned occlusion artifacts because they occur in different regions of the estimated outputs. However, since trinocular setups are generally uncommon and hence datasets seldom available, we will show how to rely on popular stereo datasets such as CityScapes and KITTI to enforce our trinocular training assumption. Experimental results clearly prove that, deploying stereo pairs with a smart strategy aimed at emulating a trinocular setup, our Three-view Network (3Net) is able anyway to learn a three-view representation of the scene as shown intuitively in Figure 1 and how it leads to more robust monocular depth estimation compared to state-of-the-art methods trained on the same binocular stereo pair with a conventional paradigm. Figure 1 highlights the behavior of 3Net: we can see how disparity maps (b) and (d), from the point of view of two frames respectively on the left and right side of the reference image, show mirrored artifacts in occluded regions. Combining the two opposite views enables to compensate for these issues and produces a more accurate map (c) centered on the reference frame. Please note that KITTI does not explicitly contain trinocular views as those shown in Figure 1 and that this behavior is learned by 3Net trained only on standard binocular data. Indeed, images and depth maps in (b) and (d) are inferred by our network. Exhaustive experimental results on the KITTI 2015 stereo dataset and the Eigen split of the KITTI dataset clearly show that 3Net, trained on standard binocular stereo pairs, improves state-of-the-art methods for unsupervised monocular depth estimation, regardless of the cues deployed for training.
Related Work
In this section, we review the literature concerning single view depth estimation in both supervised and unsupervised manner. Moreover, we also consider early works on multi-baseline stereo setup being these approaches relevant to our proposal.
Supervised depth-from-mono. The following techniques share the need for difficult to source ground-truth depth measurements for training, thus posing a substantial limitation to their practical deployment. Saxena et al. estimated depth and local planes using a MRF framework. Ladicky et al. proved that semantic can help depth estimation using a boosting classifier. More recently, CNN has emerged as mainstream strategy to estimate depth from a single image . Ummenhofer et al. proposed DeMoN, a deep model to infer both depth and ego-motion from a pair of subsequent frames acquired by a single camera. Fu et al. introduced a novel strategy to discretize depth and cast the learning process as an ordinal regression problem, while Xu et al. integrated CRF models into deep architectures to improve depth prediction. Luo et al. formulated the monocular depth estimation problem as a view synthesis procedure followed by a deep stereo matching approach. Kumar et al. introduced a Recurrent Neural Network (RNN) aimed at predicting depth from monocular video sequences. Lastly, Atapour et al. exploited image style transfer and adversarial training to predict depth from real images training the network on a large amount of synthetic data.
Unsupervised depth-from-mono. Rethinking depth estimation as an image reconstruction task allowed to avoid the need for ground-truth depth labels and some works concerned with view synthesis paved the way for this purpose. Flynn et al. proposed DeepStereo to generate new points of view training on images acquired by multiple cameras. Xie et al. trained their Deep3D framework to create, from a single image, a target frame paired with the input according to a stereo setup by learning a disparity representation.
Unsupervised monocular depth estimation methods can be broadly categorized into two main categories according to the cues used to replace ground-truth labels. The first one leverages images with known relative camera pose, typically acquired by a calibrated stereo rig, following the strategy outlined by Deep3D . A seminal work using this methodology was proposed by Garg et al. . Godard et al. deploying spatial transformer networks and left-right consistency were able to notably improve depth accuracy. More compact models can be trained the same way and deployed on embedded systems as well.
The second category concerns the use of imagery acquired by an unconstrained moving camera . Differently, from the previous methodology, temporally adjacent frames acquired by a single moving camera may contain dynamic objects that need to be explicitly handled during re-projection. Moreover, camera pose is unknown and needs to be estimated together with depth. On the other hand, such a strategy does not require a stereo camera to collect training samples. On this track, Zhou et al. proposed a model to infer depth from unconstrained video sequences by computing a reconstruction loss between subsequent frames and predicting, at the same time, the relative pose between them. This strategy was improved by Mahjourian et al. thanks to a 3D point-cloud alignment loss and by Wang et al. including a differentiable implementation of Direct Visual Odometry (DVO) with a novel depth normalization strategy. Yin et al. proposed GeoNet, a framework for depth and optical flow estimation from monocular sequences. Finally, we mention the work of Zhan et al. which combined both strategies (i.e., training on stereo sequences) and the semi-supervised works of Kuznietsov et al. and Kumar et al. .
Multi-baseline stereo. It is generally recognized that using more than two views has the potential to improve the quality of depth estimation. An early work concerning multi-camera stereo was proposed by Minoru and Akira deploying a triangular rig, while Okutomi and Kanade achieved accurate depth measurements combining stereo from multiple baseline cameras horizontally aligned. Kang et al. proposed a method to handle the increasing number of occlusions occurring in multi-view stereo setup, while Ayache and Lustman designed a three cameras rig for robotic applications and Garcia et al. proposed a pose detection algorithm based on a trinocular stereo system. In the last decade, along with the availability of off-the-shelf stereo cameras (e.g., Intel RealSense) some multi-baseline stereo systems too were commercially made available. For instance, the Bumblebee XB3 was used to acquire the RobotCar dataset , counting millions of images acquired driving for about 1000 Km. Honneger et al. developed a multi-baseline camera with on-board FPGA, enabling real-time processing of dense disparity maps. Therefore, a trinocular stereo configuration for training, like the one we advocate in our work, would be undoubtedly feasible. Nonetheless, our strategy is feasible and useful even with conventional binocular datasets.
Method overview
In this section, we propose a framework aimed at enforcing a trinocular assumption for training in an unsupervised manner a network for monocular depth estimation. We will outline the rationale behind this choice and the differences with known techniques in the literature. Then, deploying a conventional binocular stereo dataset, we will show how our strategy allows advancing state-of-the-art.
While traditional depth-from-mono frameworks learn to estimate from an input image by minimizing the prediction error with respect to a ground-truth map whose pixels are labelled with real depth measurements, the introduction of image-reconstruction based losses moved this task to an unsupervised learning paradigm. In particular, estimated depth is used to project across different points of view exploiting 3D geometry and camera pose thus obtaining supervision signals through the minimization of the re-projection error. According to the literature reviewed in Section 2, the training methodology based on images acquired with a stereo camera, as in , removes the need to infer pose estimation required when gathering data with a single unconstrained camera.
Figure 2 also highlights a further main difference between the two frameworks. While a traditional UNet architecture is used by previous works in literature , we build two separate decoders respectively in charge of estimating and separately. This strategy adds a negligible overhead regarding memory and runtime requirements, being the encoder the most computationally expensive module of the framework (i.e., the decoder mostly applies upsampling operations). According to our experiments, training a single decoder to infer a disparity representation for both points of view yields slightly worse results.
2 Interleaved training for binocular images
To effectively learn mirrored representation and compensate for occlusions/borders, the framework outlined so far relies on a set of three horizontally aligned images at training time. Although sensors designed to acquire such imagery are currently available, for instance the aforementioned Bumblebee XB3, it is still quite uncommon to find publicly available images obtained in such configuration. Indeed, in this sense, the Oxford RobotCar dataset represents an exception providing a large amount of street scenes captured with the trinocular XB3 sensor. Unfortunately, the provided calibration parameters only allow obtaining aligned views between left-right and center-right cameras, hence not permitting to align the three views as we desire. Nonetheless, we describe in this section how to train our framework leveraging the proposed trinocular assumption with a much more common binocular setup (e.g., KITTI dataset). Given a stereo pair made of images and , Figure 3 depicts how to enforce the trinocular assumption by scheduling an interleaved training of the network. We update the parameters of the network by optimizing its four outputs and in two steps:
It is worth to note that, following this protocol, every time we run a training iteration on a stereo pair the network learns all the depth representations output of our framework. Moreover, the two learned disparity pairs from the two views are obtained according to the same baseline (i.e., the same of the training stereo pairs), making them consistent and hence easy to combine in . Therefore, the network learns a trinocular representation even if it actually never sees the scene with such setup. Indeed, this strategy is very effective as supported by experimental evidence in Section 5.
Implementation details
In this section, we provide a detailed description of our framework, designed with the TensorFlow APIs. The source code is available at https://github.com/mattpoggi/3net.
For our 3Net we follow a quite established design strategy adopted by other methods in this field , based on an encoder-decoder architecture. The peculiarity of our approach consists in two different decoders, as depicted in Figure 3, in charge of learning disparity representations w.r.t. two points of view located respectively on the left and right side of the input image. In our network, depicted in Figure 3, each decoder generates outputs at four different scales, respectively: full, half, quarter and resolution. As encoder, we tested VGG and ResNet50 to obtain the most fair and complete comparison w.r.t. , being it our baseline and state-of-the-art. To obtain the final map , we merge the contribution of and using the same post-processing procedure applied in , thus keeping 5% left-most pixels from , 5% right-most from and averaging the remaining ones.
2 Training losses
We train 3Net to minimize a multi-component loss made of appearance, smoothness and consistency-check terms similarly to , namely and .
The first term uses a weighted sum of SSIM and L1 between all four warped pairs and real images as shown on top of Figure 3. The second applies an edge aware smoothness constraint to estimated disparities and as described in . Finally, the consistency-check term includes left-right losses between pairs .
For a detailed description of and please refer to or our supplementary material.
Thus, according to the interleaved training schedule described in Section 3.2, we optimize 3Net splitting the function 5 into two sub-losses deployed in the two different phases:
We also evaluated an additional loss term to enforce consistency between depth representation centered on , being the baseline equal on both directions. However, this term propagates occlusions artifacts between the two depth maps leading to worse results. We point out that despite the interleaved training protocol outlined, in any phase the outcome of 3Net always consists of four depth maps and . Of course, this happens at testing/inference time as well, when 3Net is fed with a single image. Considering that decoders outputs depth maps at four scales, all losses are computed for each of them as in .
3 Training protocol
We assume as baseline the framework proposed by Godard et al. using a binocular setup for unsupervised training. For a fair comparison, we train our models following the same guidelines reported in . In particular, we use CityScapes (CS) and KITTI raw sequences datasets for training, this latter sub-sampled according to two training splits of data to be able to compare our results with any recent works in this field using unsupervised learning. We refer to these two subsets as KITTI split (K) and Eigen split (E) . The three training sets count respectively about 23k, 29k and 22.6k stereo pairs. As pointed out by first works using image reconstruction losses , training on different datasets helps the network to achieve higher-quality results. Therefore, to better assess the performance of each method, we report experimental results training the networks on K or E. Moreover, we also report results training on CityScapes and then fine tuning on K or E (respectively, referred to as CS+K and CS+E in the tables). Consistently with , we run 50 epochs of training on each single dataset using a batch size of 8 and input resized to . We use Adam optimizer with and , setting an initial learning rate of halved after 30 and 40 epochs. We maintain the same hyperparameters configuration for and defined in and the same data augmentation procedure as well.
Experimental results
In this section, we assess the performance of our 3Net framework with respect to state-of-the-art. In all our tests, we report 7 main metrics measuring the average depth error (Abs Rel, Sq Rel, RMSE and RMSE log, the lower the better) and three accuracy scores ( and , the higher the better), First, we report experiments on the K split assuming Godard et al. as baseline. Then, we exhaustively compare 3Net with top performing unsupervised frameworks for depth-from-mono estimation, highlighting how our proposal is state-of-the-art. It is worth stressing that the proposed interleaved training procedure of 3Net, described in Section 3.2, allows for a fair comparison with any other method included in our evaluation being all trained exactly on the same (binocular) datasets. Finally, we report qualitative results concerning the trinocular representation learned by 3Net.
Table 1 reports experimental results on the KITTI 2015 stereo dataset . The evaluation was carried out on 200 stereo pairs with available high quality ground-truth disparity annotations. Additionally, being the outputs of and 3Net disparity maps, in our evaluation we include the D1-all score representing the percentage of pixels having a disparity error larger than 3.
We compare the raw output of our network with the map predicted by Godard et al. with and without post-processing (namely “+pp” in the table) running, respectively, a single or two forwards of the network. Moreover, since 3Net can benefit from the same refinement technique by running two predictions on and its horizontally flipped version, we also provide results for our network applying the same post-processing. Therefore, we estimate post-processed and before combining them into . Anyway, we report for clarity in the last column of the table, the number of forwards required by each entry.
As reported on the first two rows of Table 1, training the networks on KITTI data only, our method outperforms the competitor on all metrics except D1-all when running a single forward and it performs very similar to the post-processed version of reported in the third row of the table. Rows 3 and 4 highlight that, performing two forwards and post-processing, 3Net + pp outperforms Godard et al. + pp again on all metrics except D1-all.
Previous works in literature proved that transfer learning from CityScape dataset to KITTI is beneficial and leads to more accurate depth estimation, thus we follow this guideline training on CS+K as well. The last four rows of Table 1 compare both frameworks with and without post-processing. Without post-processing, we can notice how pre-training on CityScapes dataset allows 3Net to outperform on all metrics including D1-all. In the last two rows, applying post-processing to the output of both models, 3Net outperforms the competitor on all metrics tying on and . Figure 4 qualitatively shows depth maps predicted by (b) and 3Net (c) without applying any post-processing to better perceive the improvements lead by our framework.
Summarizing, experiments on the KITTI split highlighted how enforcing the trinocular assumption is more effective than leveraging a conventional stereo paradigm for training. Moreover, these results prove that stereo pairs can be used in a smarter way following our interleaving strategy.
2 Eigen split
Table 3 reports evaluation with the split of data of Eigen et al. , made of 697 images and relative depth measurements acquired with a Velodyne sensor. The table collects results concerning most recent works addressing unsupervised monocular depth estimation. For each method, we indicate the kind of supervision it leverages on: monocular sequences (Temporal), stereo pairs (Stereo) or stereo sequences (Stereo+Temp.). We report results either training on E only or on CS+E, allowing to compare our scores with state-of-the-art approaches. We point out that all methods, including our proposal, are trained exactly on the same images and all of them see the same scenes Zhan et al. report scores training on E only or after pre-training on NYU dataset . For fairness we report the first setup only.. On top, we report results for models trained on the Eigen split of data. For GeoNet , Godard et al. and our method we report results for both VGG and ResNet50 encoders. We can notice that, in general, methods trained using stereo data usually outperform those trained on monocular video sequences, as evident from recent literature . Zhan et al. leveraging temporally adjacent stereo frames outperform, on most metrics, . Nevertheless, 3Net achieves better scores except for still without exploiting temporal supervision. This proves that a smarter deployment of binocular training samples, i.e. by applying our interleaved training to fulfill trinocular hypothesis, is an effective alternative to sequence supervision. It is worth to note that Wang et al. obtain better scores on most metrics (RMSE, RMSE log and metrics) w.r.t. and 3Net with the VGG encoder. However, by switching to the ResNet50 encoder, Godard et al. overtakes most recent works that use Time supervision with and without post-processing. Systematically, 3Net always outmatches and consequently all its competitors. In particular, we point out that 3Net ResNet50 without post-processing already achieves some better scores compared to Godard et al. ResNet50 + pp performing, respectively, a single and a double forward.
On the bottom of Table 3, we resume results achieved by models trained on CS+E. We observe the same trend highlighted in the previous experiments, being and our proposal the most effective solutions for this task thanks to stereo supervision. In equal conditions, i.e. same encoder and number of forwards, 3Net always outperforms the framework of Godard et al. exploiting the trinocular assumption. Moreover, the proposed technique leads to major improvements such that 3Net VGG outperforms ResNet50 model by Godard et al. (rows 20th and 21st), 3Net ResNet50 without post-processing achieves more accurate results than the best configuration ResNet50 + pp (rows 22nd and 23rd) and finally 3Net ResNet50 + pp outmatches all known frameworks for unsupervised depth-from-mono estimation. These facts clearly highlight that our proposal is state-of-the-art.
It is important to underline that the availability of a real trinocular rig would most probably allow training a more accurate model, given the larger amount of images w.r.t. a binocular stereo rig. The interleaved training proposed in this paper allows to overcome the lack of trinocular training samples using binocular pairs and also allows for a more fair comparison with other techniques leveraging this latter configuration only. This fact proves that the effectiveness of our strategy is due to the rationale behind it and not driven by a more extensive availability of data.
View synthesis
Conclusions
In this paper, we proposed a novel methodology for unsupervised training of a depth-from-mono CNN. By enforcing a trinocular assumption, we overcome some limitations due to binocular stereo images used as supervision and obtain a more accurate depth estimation with our 3Net architecture. Although three horizontally aligned views are seldom available, we proposed an interleaved training protocol allowing to leverage on traditional binocular datasets. This latter technique also ensures for a fair comparison w.r.t. all previous works and allows us to prove that 3Net outperforms all unsupervised techniques known in the literature, establishing itself as state-of-the-art. Moreover, 3Net learns a trinocular representation of the world, making it suitable for image synthesis purposes and other interesting future developments.
Acknowledgements
We gratefully acknowledge the support of NVIDIA Corporation with the donation of the Titan X GPU used for this research.
References
Supplementary material
This document provides additional details and experimental results concerned with paper ”Learning monocular depth estimation under unsupervised trinocular assumption”. The supplementary material is organized as follows: Section 8 reports detailed explanation of the loss functions used at training time, Section 9 describes how we obtain with 3Net and how we post-process it, Section 10 comments additional experiments on the Eigen split assuming as maximum depth 50 meters and Section 11 collects additional qualitative results, Finally, Section 12 reports run time analysis for 3Net and .
Training losses
In the paper, all loss functions are computed at four scales, ranging from full image resolution to . The global loss function is defined as:
where , and represent, respectively, the appearance, smoothness and consistency terms, while , and are hyper-parameters. In particular, we set and .
Appearance Loss. It measures the reconstruction error between a warped image and the original one. It is obtained by a weighted sum of a SSIM based score and a L1 distance over pixel intensities.
Smoothness Loss. This term favours the propagation of similar disparity values in low-textured regions, thus enforcing smoothness. It is obtained computing horizontal and vertical gradients on both disparity image and reference image, discouraging disparity smoothness in presence of strong image gradients.
Left-Right Disparity Consistency Loss. It enforces consistency between reference-to-target and target-to-reference disparity maps. It relies on the L1 distance between reference-to-target map and warped, according to the former, target-to-reference map.
Depth computation and post-processing
For the sake of clarity, we describe in detail how we combine and to obtain the final output map . In the authors obtained and by processing, respectively, both and its horizontally flipped version . The two maps were combined as follows:
Following this principle, we combine our and maps as follows:
Running two forwards, we can post-process both intermediate maps and
being and obtained as:
Depth estimation: additional experiments with 50m cap
We report additional experimental results on the Eigen split , evaluating depth maps up to a maximum distance of 50 meters as reported in some recent works . Table 3 contains a comparison between all previous works reporting this experiment as well and our best model, i.e. 3Net ResNet50 + pp. This further evaluation confirms, once again, the superiority of our technique with respect to all competitors.
View synthesis and multi-baseline stereo
Finally, deploying 3Net ResNet50 + pp trained on CS+E, we provide additional qualitative results for depth-from-mono estimation and view synthesis. Figure 6 and 7 reports six examples taken from the evaluation set of the Eigen split . In particular, we show in the leftmost column the generated left view (a), the single input image fed to our network (b) and the generated right view (c). In the mid column, the three output maps of 3Net, respectively, (d), (e) and (f). Finally, in the rightmost column, we report disparity maps obtained processing with SGM the three stereo pairs obtainable with 3Net from the three views (one real, two synthetic) depicted in the leftmost column. In particular, the disparity maps computed by SGM are concerned with three stereo pairs: left-to-center (g), center-to-right (h) and left-to-right (i). It is worth to note that the left-to-right stereo pair (i) is made of two completely novel views synthesized by our network. The other two stereo pairs contain the input image and a a novel image synthesized by 3Net.
Observing (a), (b) and (c) we can easily notice three different view points: the two virtual cameras are located at the left and right side of the real camera (i.e., the central one). The three maps in the middle column clearly show artifacts occurring near depth discontinuities and occlusions in (d) and (f) and how they are greatly dampen in the final output of our network (e). Finally, we can perceive how (g) and (i) share the same reference image (synthetic left) and how they compute different disparity values according to different baselines, narrow and wide, made available by the three-view virtual rig enabled by 3Net.
A video showing the performance of 3Net on the KITTI sequence 2011_10_03_drive_0047_sync not part of the Eigen split imagery used for training is available at this link: https://www.youtube.com/watch?v=uMA5YWJME4M. Finally, the source code is available at this link: https://github.com/mattpoggi/3net
Runtime analysis
In this section, we briefly compare the runtime of 3Net compared to the models by Godard et al. . On high-end GPUs (e.g., Titan X Pascal), the difference between the two models either running single or double forward is negligible, taking between 0.09 and 0.11 seconds both. Nevertheless, in case of applications deploying different architectures the margin rises.
In particular, Table 4 compares the execution times of the considered models using ResNet50 encoder on a CPU Intel Core i7-7700K. Times are averaged on the entire Eigen split testing set. We report numbers at resolution (i.e., the dimensions used by at inference time), as well as at full KITTI resolution, to stress how the difference between them increases with the image size. We can see how the second encoder in 3Net adds about 50% overhead, while forwards usually doubles it. However, by recalling results reported in the main paper (Table 2, last 3 rows on bottom), 3Net ResNet50 running a single forward is more accurate and faster than ResNet50 running two forwards.
Acknowledgements
We gratefully acknowledge the support of NVIDIA Corporation with the donation of the Titan X GPU used for this research.