Depth Prediction Without the Sensors: Leveraging Structure for Unsupervised Learning from Monocular Videos

Vincent Casser, Soeren Pirk, Reza Mahjourian, Anelia Angelova

Previous Work

Scene depth estimation has been a long standing problem in vision and robotics. Numerous approaches, involving stereo or multi-view depth estimation exist. Recently a learning-based concept for image-to-depth estimation has emerged fueled by availability of rich feature representations, learned from raw data [\citeauthoryearEigen, Puhrsch, and Fergus2014, \citeauthoryearLaina et al.2016]. These approaches have shown compelling results as compared to traditional methods [\citeauthoryearKarsch, Liu, and Kang2014b]. Pioneering work in unsupervised image-to-depth learning has been proposed by [\citeauthoryearZhou et al.2017, \citeauthoryearGarg, Carneiro, and Reid2016] where no depth or ego-motion is needed as supervision. Many subsequent works have improved the initial results in both the monocular setting [\citeauthoryearYang et al.2017, \citeauthoryearYin2018] and when using stereo during training [\citeauthoryearGodard, Aodha, and Brostow2017, \citeauthoryearUmmenhofer et al.2017, \citeauthoryearZhan et al.2018, \citeauthoryearYang et al.2018a].

However, these methods still fall short in practice because object movements in dynamic scenes are not handled. In these highly dynamic scenes, the abovementioned methods tend to fail as they can not explain object motion. To that end, optical flow models, trained separately, have been used with moderate improvements [\citeauthoryearYin2018, \citeauthoryearYang et al.2018b, \citeauthoryearYang et al.2018a]. Our motion model is most aligned to these methods as we similarly use a pre-trained model, but propose to use the geometric structure of the scene and model all objects’ motion including camera ego-motion. The refinement method is related to prior work [\citeauthoryearBloesch et al.2018] who use lower dimensional representations to fuse subsequent frames; our work shows that this can be done in the original space to a very good quality.

Main Method

The main learning setup is unsupervised learning of depth and ego-motion from monocular video [\citeauthoryearZhou et al.2017], where the only source of supervision is obtained from the video itself. We here propose a novel approach which is able to model dynamic scenes by modeling object motion, and that can optionally adapt its learning strategy with an online refinement technique. Note that both ideas are tangential and can be used either separately or jointly. We describe them individually, and demonstrate their individual and joint effectiveness in various experiments.

Using a warping operation of one image to an adjacent one in the sequence, we are able to imagine how a scene would look like from a different camera viewpoint. Since the depth of the scene is available through θ(Ii)\theta(I_{i}), the ego-motion to the next frame ψE\psi_{E} can translate the scene to the next frame and obtain the next image by projection. More specifically, with a differentiable image warping operator ϕ(Ii,Dj,Ei→j)→I^i→j\phi(I_{i},D_{j},E_{i\rightarrow j})\rightarrow\hat{I}_{i\rightarrow j}, where I^i→j\hat{I}_{i\rightarrow j} is the reconstructed jj-th image, we can warp any source RGB-image IiI_{i} into IjI_{j} given corresponding depth estimate DjD_{j} and an ego-motion estimate Ei→jE_{i\rightarrow j}. In practice, ϕ\phi performs the warping by reading from transformed image pixel coordinates, setting I^i→jxy=Iix^y^\hat{I}_{i\rightarrow j}^{xy}=I_{i}^{\hat{x}\hat{y}}, where [x^,y^,1]T=KEi→j(Djxy⋅K−1[x,y,1]T)[\hat{x},\hat{y},1]^{T}=KE_{i\rightarrow j}(D_{j}^{xy}\cdot K^{-1}[x,y,1]^{T}) are the projected coordinates. The supervisory signal is then established using a photometric loss comparing the projected scene onto the next frame I^i→j\hat{I}_{i\rightarrow j} with the actual next frame IjI_{j} image in RGB space, for example using a reconstruction loss: Lrec=min⁡(∥I^1→2−I2∥L_{rec}=\min(\|\hat{I}_{1\rightarrow 2}-I_{2}\|.

Algorithm Baseline

We establish a strong baseline for our algorithm by following best practices from recent work [\citeauthoryearZhou et al.2017, \citeauthoryearGodard, Aodha, and Brostow2018]. The reconstruction loss is computed as the the minimum reconstruction loss between warping from either the previous frame or the next frame into the middle one:

proposed by [\citeauthoryearGodard, Aodha, and Brostow2018] to avoid penalization due to significant occlusion/disocclusion effects. In addition to the reconstruction loss, the baseline uses an SSIM [\citeauthoryearWang et al.2004] loss, a depth smoothness loss and applies depth normalization during training, which demonstrated success in prior works [\citeauthoryearZhou et al.2017, \citeauthoryearGodard, Aodha, and Brostow2017, \citeauthoryearWang et al.2018]. The total loss is applied on 44 scales (αj\alpha_{j} are hyperparameters):

Motion Model

We introduce an object motion model ψM\psi_{M} which shares the same architecture as the ego-motion network ψE\psi_{E}, but is specialized to predicting motions of individual objects in 3D (Figure 2). Similar to the ego-motion model, it takes an RGB image sequence as input, but this time complemented by pre-computed instance segmentation masks. The motion model is then tasked to learn to predict the transformation vectors per object in 3D space, which creates the observed object appearance in the respective target frame. Thus, computing warped image frames is now not only a single projection based on ego-motion as in prior work [\citeauthoryearZhou et al.2017], but a sequence of projections that are then combined appropriately. The static background is generated by a single warp based on ψE\psi_{E}, whereas all segmented objects are then added by their appearance being warped first according to ψE\psi_{E} and then ψM\psi_{M}. Our approach is conceptually different from prior works which used optical flow for motion in 2D image space [\citeauthoryearYin2018] or 3D optical flow [\citeauthoryearYang et al.2018a] in that the object motions are explicitly learned in 3D and are available at inference. Our approach not only models objects in 3D but also learns their motion on the fly. This is a principled way of modeling depth independently for the scene and for each individual object.

To model object motion, we first apply the ego-motion estimate to obtain the warped sequences (I^1→2,I2,I^3→2)(\hat{I}_{1\rightarrow 2},I_{2},\hat{I}_{3\rightarrow 2}) and (S^1→2,S2,S^3→2)(\hat{S}_{1\rightarrow 2},S_{2},\hat{S}_{3\rightarrow 2}), where the effect of ego-motion has been removed. Assuming that depth and ego-motion estimates are correct, misalignments within the image sequence are caused only by moving objects. Outlines of potentially moving objects are provided by an off-the-shelf algorithm [\citeauthoryearHe et al.2017] (similar to prior work that use optical flow [\citeauthoryearYang et al.2018a] that is not trained on either of the datasets of interest). For every object instance in the image, the object motion estimate M(i)M^{(i)} of the ii-th object is computed as:

and the equivalent for I^3→2(F)\hat{I}^{(F)}_{3\rightarrow 2}. In the above, we denote the gradients per each term. Note that the employed masking ensures that no pixel in the final warping result gets occupied more than once. While there can be regions which are not filled, these are handled implicitly by the minimum loss computation. Our algorithm will automatically learn individual 3D motion per object which can be used at inference.

Imposing Object Size Constraints

effectively prevents all segmented objects to degenerate into infinite depth, and forces the network to produce not only a reasonable depth but also matching object motion estimates. We scale by D‾\overline{D}, which is the mean estimated depth of the middle frame, to reduce a potential issue of trivial loss reduction by jointly shrinking priors and the depth prediction range. To our knowledge this is the first method to address common degenerative cases in a fully monocular training setup in 3D. Since this constraint is an integral part of the modeling formulation, the motion models are trained with LscL_{sc} from the beginning. However, we observed that this additional loss can successfully correct wrong depth estimates when applying it to already trained models, in which case it works by correcting depth for moving objects.

Test Time Refinement Model

One advantage of having a single-frame depth estimator is its wide applicability. However, this comes at a cost when running continuous depth estimation on image sequences as consecutive predictions are often misaligned or discontinuous. These are caused by two major issues 1) scaling inconsistencies between neighboring frames, since both our and related models have no sense of global scale, and 2) low temporal consistency of depth predictions. In this work we contend that fixing the model weights during inference is not required or needed and being able to adapt the model in an online fashion is advantageous, especially for practical autonomous systems. More specifically, we propose to keep the model training while performing inference, addressing these concerns by effectively performing online optimization. In doing that, we also show that even with very limited temporal resolution (i.e., three-frame sequences), we can significantly increase the quality of depth predictions both qualitatively and quantitatively. Having this low temporal resolution allows our method to still run on-line in real-time, with a typically negligible delay of a single frame. The online refinement is run for NN steps (N=20N=20 for all experiments) which are effectively fine-tuning the model on-the-fly; NN determines a good compromise between exploiting the online tuning sufficiently and preventing over-training which can cause artifacts. The online refinement approach can be seamlessly applied to any model including the motion model described above.

Experimental Results

Extensive experiments have been conducted on depth estimation, ego-motion estimation and on transfer learning to new environments. We use common metrics and protocols for evaluation adopted by prior methods. With the same standards as in related work, if depth measurements in the groundtruth are invalid or unavailable, they are masked out in the metric computation. We use the following datasets:

KITTI dataset (K). The KITTI dataset [\citeauthoryearGeiger et al.2013] is the main benchmark for evaluating depth and ego-motion prediction. It has LIDAR sensor readings, used for evaluation only. We use standard splits into training, validation and testing, commonly referred to as the ‘Eigen’ split [\citeauthoryearEigen, Puhrsch, and Fergus2014], and evaluate depth predictions up to a fixed range (80 meters).

Cityscapes dataset (C). The Cityscapes dataset [\citeauthoryearCordts et al.2016] is another popular and also challenging dataset for autonomous driving. It contains 3250 training and 1250 testing examples which are used in our setup. Of note is that this dataset contains many dynamic scenes with multiple moving objects. We use it for training and for evaluating transfer learning, without fine-tuning.

Fetch Indoor Navigation dataset. This dataset is produced by our Fetch robot [\citeauthoryearWise et al.2016] collected for the purposes of indoor navigation. We test an even more challenging transfer learning scenario when training on an outdoor navigation dataset, Cityscapes, and testing on the indoor one without fine-tuning. The dataset contains 1,6261,626 images from a single video sequence, recorded at 8fps.

Figure 3 visualizes the results of our method compared to state-of-the-art methods and Table 1 shows quantitative results. Both show a notable improvement over the baseline and over previous methods in the literature. With an absolute relative error of 0.10870.1087, our method is outperforming competitive models that use motion, 0.1310.131 [\citeauthoryearYang et al.2018a] and 0.1550.155 [\citeauthoryearYin2018]. Furthermore, our results, although monocular, are approaching methods which use stereo or a combination of stereo and monocular, e.g. [\citeauthoryearGodard, Aodha, and Brostow2017, \citeauthoryearKuznietsov, Stuckler, and Leibe2017, \citeauthoryearYang et al.2018a, \citeauthoryearGodard, Aodha, and Brostow2018].

The main contributions of the motion model are that it is able to learn proper depth for moving objects and it learns better ego-motion. Figure 4 shows several examples of dynamic scenes from the Cityscapes dataset, which contain many moving objects. We note that our baseline, which is by itself a top performer on KITTI, is failing on moving objects. Our method makes a notable difference both qualitatively (Figure 4) and quantitatively (see Table 2). Another benefit provided by our motion model is that it learns to predict individual object motions. Figure 6 visualizes the learned motion for individual objects. See the project webpage for a video which demonstrates depth prediction as well as relative speed estimation which is well aligned with the apparent ego-motion of the video.

Refinement model.

We observe improvements obtained by the refinement model on both KITTI and Cityscapes datasets. Figure 5 shows results of the refinement method only as compared to the baseline. As seen for both evaluating on KITTI or Cityscapes dataset the refinement is helpful in recovering the geometry structure better. In our results we observe that the refinement model is most helpful when testing across datasets, i.e. in data transfer.

Experimental Results on the Cityscapes Dataset

In this section we evaluate our method on the Cityscapes dataset, where a lot of object motion is present in the training set. Table 2 shows our experimental results when training on the Cityscapes data, and then evaluating on KITTI (without further fine-tuning on KITTI training data). This experiment clearly demonstrates the benefit of our method as we see significant improvements from 0.205 to 0.153 absolute relative error for the proposed approach, which is particularly impressive in the context of state-of-the-art error of 0.233. It is also seen that improvements are accomplished by both the motion and the refinement model individually and jointly. We note that the significant improvement of the combined model stems from both the appropriate depth learning of many moving objects (Figure 4) enabled by the motion component, and the refinement component that actively refines geometry in the scene (Figure 5).

Visual Odometry Results

Table 3 summarizes our ego-motion results, which are conducted by a standard protocol adopted by prior work [\citeauthoryearZhou et al.2017, \citeauthoryearGodard, Aodha, and Brostow2018] on parts of the KITTI odometry dataset. The total driving sequence lengths tested are 1,702 meters and 918 meters, respectively. As seen our algorithm performance is the best among the state-of-the-art methods, even compared to ones that use more temporal information, or established methods such as ORB-SLAM. Proper handling of motion is the biggest contributor to improving our ego-motion estimation.

Experiments on Fetch Indoor Navigation Dataset

Finally, we verify the approach in an indoor environment setting, by testing on data collected by the Fetch robot [\citeauthoryearWise et al.2016]. This is a particularly challenging transfer learning scenario as training is done on Cityscapes (outdoors) and testing is done on a dataset collected indoors by a different robot platform, representing a significant domain shift between these datasets. Figure 7 visualizes the results on the Fetch data. Our algorithm produces better and more realistic depth estimates and is able to notably improve the baseline method and successfully adapt to new environments. Notably, the algorithm is able to capture well large transparent glass doors and windows and reflective surfaces. We observe that transfer works best if the amount of motion in between frames is somewhat similar. Also, to have additional information available and not lead to degenerate evolution, camera motion should be present. Thus, in a static state, online refinement should not be applied.

Implementation details. The code is implemented in TensorFlow and publicly available. The input images are resized to 416×128416\times 128 (with center cropping for Cityscapes). The experiments are run with: learning rate 0.00020.0002, L1 reconstruction weight 0.850.85, SSIM weight 0.150.15, smoothing weight 0.040.04, object-motion constraint weight 0.00050.0005 (although 0.00020.0002 seems to work better for KITTI), batch size of 44, L2 weight regularization of 0.050.05. We perform on-the-fly augmentation by horizontal flipping during testing.

Conclusions and Future Work

The method presented in this paper addresses the monocular depth and ego-motion problem by modeling individual objects’ motion in 3D. We also propose an online refinement technique which adapts learning on the fly and can transfer to new datasets or environments. The algorithm achieves new state-of-the-art performance on well established benchmarks, and produces higher quality results for dynamic scenes. In the future, we plan to apply the refinement method over longer sequences so as to incorporate more temporal information. Future work will also focus on full 3D scene reconstruction which is enabled by the proposed depth and ego-motion estimation methods.

Acknowledgements. We would like to thank Ayzaan Wahid for helping us with data collection.

References