Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic Scenes

Fabian Brickwedde, Steffen Abraham, Rudolf Mester

Introduction

In applications such as mobile robots or autonomous vehicles a representation of the surrounding environment is utilized, e.g. to fulfill a navigation task. From a computer vision point of view, the 3D position and motion of a pixel in the image is denoted as 3D scene flow , which is traditionally estimated based on a temporal series of stereo images . In this work, we propose a novel scene flow estimation method, Mono-SF, for a monocular camera setup focusing on dynamic traffic scenes. Monocular camera systems are often preferred over stereo cameras due to being more cost efficient and to avoid the effort of calibrating the stereo rig. However, 3D scene flow estimation is an ill-posed problem in a monocular camera setup. To solve the ambiguity, previous monocular methods assumed that the moving objects are in contact with the surrounding environment or that the scene follows a smoothness prior regarding surface and motion . These assumptions might be violated and the methods still require a relative translational motion of the camera to the scene. In contrast to the multi-view geometry-based approaches, methods were proposed (e.g. ) that provide depth estimates from a single image at a reasonable level of quality. However, single-view depth estimation and multi-view geometry are mostly tackled as two individual tasks or fused in a way that is only applicable for static scenes . Our proposed Mono-SF method combines multi-view geometry with single-view depth information in a probabilistic optimization framework to provide consistent 3D scene flow estimates. Thereby, both kinds of information are exploited and the single-view depth serves to solve the multi-view geometry-based ambiguity.

Previous methods showed that a suitable representation of particularly traffic scenes is the decomposition into 3D planar surface elements, each one assigned to a rigid body. A rigid body is either the background or a potentially moving object. Following this model, Mono-SF jointly estimates the 3D geometry of each plane and 6D motion of each rigid body considering a) the multi-view geometry by warping the reference image into the consecutive image, b) probabilistic single-view depth estimates, and c) scene model smoothness priors (see Fig. 1). Additionally, an instance segmentation is exploited to detect the set of potentially moving objects.

As an additional contribution, we propose ProbDepthNet, a convolutional neural network (CNN) that estimates pixel-wise probability depth distributions from a single image rather than just single depth values such as . Whereas the problem of overconfident estimates is a well-known problem in classification , it is typically ignored in probabilistic approaches for regression . Therefore, we propose a novel recalibration technique: CalibNet, a small subsequent part of ProbDepthNet, is trained on a hold-out split of the training data to compensate for overfitting effects and to provide well-calibrated distributions.

Our Mono-SF approach is evaluated with respect to several state-of-the-art monocular baselines and an ablation study confirms the importance of the individual components of the proposed optimization framework. Furthermore, ProbDepthNet is validated to provide well-calibrated depth distributions. Our experiments show that several previous probabilistic approaches suffer from overconfident estimates – an effect that could be compensated by adding our proposed CalibNet for recalibration. The suitability of ProbDepthNet for integrating single-view depth information in Mono-SF is confirmed, especially due to the importance of providing single-view depth information in a probabilistic and well-calibrated form.

Related work

The works related to the approach presented here are divided into three categories: In the first category are the stereo-based scene flow methods which inspired our Mono-SF scene model and optimization framework. The second category provides an overview of methods for monocular scene reconstruction comprising the baseline methods. Finally, the category of probabilistic deep learning represents works related to the probabilistic design of ProbDepthNet.

Stereo Scene Flow: Scene flow estimation was introduced by Vedula et al. as a joint optimization of 3D geometry and motion of the scene based on a sequence of stereo images. Mostly variational approaches were used subsequently to extend the scene flow concept . However, Vogel et al. were the first that significantly outperformed individual stereo and optical flow methods on their respective tasks for dynamic traffic scenes. They represented the dynamic scene as a collection of rigid moving planar surface elements and jointly optimized the geometry and the motion of each plane considering scene model priors. Menze et al. formulated the problem by a set of rigid moving objects and jointly optimized their motion with the geometry of each plane. This representation is particularly beneficial if the association of planes to objects is supported by an instance segmentation as proposed in . Our Mono-SF model corresponds to these approaches, called object or instance scene flow , but Mono-SF uses only monocular images.

Monocular Scene Reconstruction: Traditionally, monocular scene reconstruction is based on the structure from motion (SfM) principle. The SfM-based approaches can be divided into several categories: First, rigid SfM-based methods estimate the 3D geometry of a rigid scene based on its relative motion to the camera, e.g. a static scene and a moving camera . Second, the non-rigid SfM principle is typically used to derive the deformation of a single object . Third, multi-body SfM is the concept of reconstructing individual moving parts of the scene separately . However, the absolute and relative scales of the reconstructions are unknown in general. Scene model assumptions are needed to solve this scale ambiguity, e.g. that moving objects are in contact with the surrounding environment or that the scene follows a smoothness prior regarding surface and motion .

Even though the idea of single-view depth estimation is by far not new , the real breakthrough was achieved by usage of deep learning methods. Pioneering, Eigen et al. proposed a CNN that is trained in a supervised manner and estimates the depth in a coarse to fine scheme. Afterward, various self-supervised and unsupervised approaches were proposed using either an image reconstruction loss in a stereo setup or in a monocular image sequence . Fu et al. formulated the depth estimation as an ordinal regression problem, which led to the currently leading approach in the KITTI depth prediction benchmark as reported by . Multi-task CNNs that estimate optical flow alongside the depth were proposed . Thereby, both tasks benefit from each other by a combined training loss. DeMoN could also exploit multi-view information for depth estimation during inference. However, it is focused and applied only to static scenes as it just estimates a single camera motion for the whole scene.

Whereas single-view depth estimation and multi-view geometry are mostly taken as individual tasks, a few works combine both. The single-view depth estimation can be useful for scale estimation in monocular visual odometry or fused with SfM-based depth estimates in static environments . Kumar et al. used single-view depth estimation for depth initialization in a multi-body or non-rigid SfM-based approach similar to . Brickwedde et al. proposed a fusion of single-view depth estimates and optical flow to provide a column-wise segmentation in stick-like rigid elements of particularly traffic scenes. In contrast to these methods, Mono-SF is formulated as a scene flow estimation problem and integrates probabilistic single-view depth distributions instead of single depth values.

Probabilistic Deep Learning: The methods of single-view depth estimation mentioned in the previous section do not provide an uncertainty measure or probabilistic distribution of the depth estimates. Kendall and Gal distinguished two kind of uncertainties, epistemic and aleatoric uncertainty. Epistemic uncertainty corresponds to the uncertainty of the model parameters or the ignorance which model generates the training data, whereas aleatoric uncertainty refers to noise in the input data . Malinin et al. extended this definition by introducing the distributional uncertainty to represent out-of-distribution data. To estimate the extent of aleatoric uncertainty in a regression problem, different strategies have been proposed. First, a probability distribution can be learned by minimizing the negative log-likelihood on the training data . Second, Ilg et al. proposed a single network that is pushed to estimate a complementary set of hypotheses. Thereby, the aleatoric uncertainty is encoded by the empirical distribution of these hypotheses. Third, Gast and Roth replaced each layer with a probabilistic layer to propagate an input uncertainty through the network. The ProbDepthNet method presented here falls under the category of estimating the aleatoric uncertainty with a single network and single inference such as . For classification problems, Guo et al. showed that modern neural networks tend to overfit on the training data, which results in highly overconfident estimates. Recalibration techniques were proposed to compensate for this effect .

Method

The monocular scene flow estimation method, Mono-SF, is designed to combine multi-view geometry with probabilistic single-view depth information in a probabilistic optimization framework. First, a CNN, called ProbDepthNet, providing single-view depth information in a probabilistic and well-calibrated form is described. Second, the Mono-SF model and optimization framework are presented.

To integrate the single-view depth estimates in Mono-SF in a statistical manner, ProbDepthNet is designed to represent the uncertainty of each estimate. Thus, the main objective of ProbDepthNet is not to provide a single depth estimate, but to provide a probability density function of the depth for each pixel p\mathbf{p} given an input image II. The depth is encoded by its inverse form d=Z−1d=Z^{-1}, where ZZ is the z-coordinate of the 3D-position in camera coordinates. ProbDepthNet estimates a pixel-wise probability density function pp(d∣I)p_{\mathbf{p}}(d\mid I) parameterized as a mixture of Gaussians:

KK represents the number of components, λi\lambda_{i} are the weights, μi\mu_{i} are the mean values, and σi\sigma_{i} are the variances of the ii-th component. Compared to a single Gaussian distribution, a mixture model is able to capture more general distributions, e.g. a multimodal distribution. But, the mixture of Gaussians is more an exemplary choice and other parameterizations of a probability distribution can be used as well.

u,v∈ΩGTu,v\in\Omega_{GT} are all pixels in the image with valid ground truth depth values dGTd_{GT} and μi,λi,σi\mu_{i},\lambda_{i},\sigma_{i} are the outputs of the trained network.

To overcome the limitations of lidar data in terms of density, range, and field of view, an intermediate fusion based on stereo images is used for ground truth depth generation. First, the lidar point cloud is projected to the image and inconsistent measurements are removed to handle occlusion problems. Second, these sparse depth maps are completed considering a photometric distance between the two stereo images by using an SGM-based approach .

ProbDepthNet learns to estimate a pixel-wise depth distribution by observing the depth distribution during the training process. Thereby, the depth distribution captures the aleatoric uncertainty regarding the theory of Kendall and Gal . The aleatoric uncertainty is considered to be the most dominant uncertainty in many vision applications . Our experiments show that CalibNet for recalibration is also applicable to different probabilistic approaches similar to .

2 Monocular Scene Flow

This section presents the Mono-SF optimization framework, structured as follows: First, the decomposition of the scene into piecewise planar surface elements and rigid bodies is described. Second, the optimization is formulated as an energy minimization problem combining a) multi-view geometry-based photometric distance, b) the probabilistic single-view depth estimates of ProbDepthNet and c) scene model smoothness priors. Finally, the inference and initialization of the optimization problem are presented.

Energy Minimization Problem: The main idea of Mono-SF is that the scene geometry and motion should be consistent in terms of warping the reference image I0I_{0} in the consecutive image I1I_{1} and consistent to the depth distributions p(d∣I0)p(d\mid I_{0}) and p(d∣I1)p(d\mid I_{1}) provided by ProbDepthNet. Formally, Mono-SF jointly optimizes the 6D motion of each rigid body Tj\mathbf{T}_{j} and 3D normal of each plane ni\mathbf{n}_{i} as an energy minimization problem. The energy term EE consists of unary data terms Φ(p0,ni,Tj)\Phi(\mathbf{p}_{0},\mathbf{n}_{i},\mathbf{T}_{j}) for each pixel p0\mathbf{p}_{0} and pairwise smoothness terms Ψ(ni,nj)\Psi(\mathbf{n}_{i},\mathbf{n}_{j}) for each two planes nk\mathbf{n}_{k} and nl\mathbf{n}_{l} adjacent in the image k,l∈Nk,l\in\mathcal{N}:

Tj\mathbf{T}_{j} is the rigid body corresponding to the plane ni\mathbf{n}_{i}.

The unary terms Φ(p0,ni,Tj)\Phi(\mathbf{p}_{0},\mathbf{n}_{i},\mathbf{T}_{j}) consist of two parts. First, Φpho(p0,ni,Tj)\Phi^{pho}(\mathbf{p}_{0},\mathbf{n}_{i},\mathbf{T}_{j}) minimizes an appearance-based photometric distance between pixel p0\mathbf{p}_{0} and its projected position in the consecutive image. Second, Φtsvd(p0,ni,Tj)\Phi^{svd}_{t}(\mathbf{p}_{0},\mathbf{n}_{i},\mathbf{T}_{j}) prefers a 3D position consistent to the estimated depth probabilities of ProbDepthNet at time t=0t=0 and t=1t=1:

The terms are weighted by Θ0\Theta_{0} or Θ1\Theta_{1}, respectively. The photometric distance Φpho(p0,ni,Tj)\Phi^{pho}(\mathbf{p}_{0},\mathbf{n}_{i},\mathbf{T}_{j}) rates the similarity of the two corresponding image positions p0\mathbf{p}_{0} and p1\mathbf{p}_{1} as the hamming distance of their respective 5×55\times 5 Census descriptors truncated at τ0\tau_{0}. The corresponding image coordinates p1\mathbf{p}_{1} in the second image I1I_{1} are defined by a homography considering the 3D normal ni\mathbf{n}_{i} and the motion of the corresponding rigid body Tj\mathbf{T}_{j}:

Rj\mathbf{R}_{j} and tj\mathbf{t}_{j} is the decomposition of Tj\mathbf{T}_{j} into rotation matrix and translation vector. K\mathbf{K} is the intrinsic camera matrix.

The term Φtsvd(p0,ni,Tj)\Phi_{t}^{svd}(\mathbf{p}_{0},\mathbf{n}_{i},\mathbf{T}_{j}) rates the consistency of the depth of pixel p0\mathbf{p}_{0} based on the ProbDepthNet estimates. Whereas the depth d0(p0,ni)d_{0}(\mathbf{p}_{0},\mathbf{n}_{i}) at time t=0t=0 is directly defined by the corresponding scaled normal vector ni\mathbf{n}_{i}, the motion of the corresponding rigid body Tj\mathbf{T}_{j} needs to be considered to derive the depth d1(p0,ni,Tj)d_{1}(\mathbf{p}_{0},\mathbf{n}_{i},\mathbf{T}_{j}) at time t=1t=1. Both depth values are rated by the negative log-likelihood of the probability provided by ProbDepthNet for their respective image ItI_{t} and image coordinate pt\mathbf{p}_{t}:

The image coordinates p1\mathbf{p}_{1} are again defined as in Eq. (5).

The previous data terms include the single-view depth information and multi-view geometry-based photometric distance. Additionally, scene model priors are integrated similar to as pairwise smoothness terms Ψ(nk,nl)\Psi(\mathbf{n}_{k},\mathbf{n}_{l}) preferring a smooth structure in terms of depth Ψd(nk,nl)\Psi^{d}(\mathbf{n}_{k},\mathbf{n}_{l}) and orientation Ψori(nk,nl)\Psi^{ori}(\mathbf{n}_{k},\mathbf{n}_{l}), each part weighted by Θ2\Theta_{2} or Θ3\Theta_{3}:

For each shared boundary pixel p0∈Bk,l\mathbf{p}_{0}\in\mathcal{B}_{k,l} of plane nk\mathbf{n}_{k} and nl\mathbf{n}_{l}, a difference in depth is penalized:

Analogously, a smooth orientation of planes adjacent in the image is preferred by measuring the similarity of the normal vectors nk\mathbf{n}_{k} and nl\mathbf{n}_{l}:

Both smoothness terms are truncated by τ1\tau_{1} or τ2\tau_{2} to regard discontinuities in the depth or orientation, for example between different objects. The hyper-parameters Θ\Theta and τ\tau are defined differently according to the rigid body type, background or object, and differently for adjacent planes belonging to different rigid bodies. These dependencies are neglected in the previous equations for ease of reading.

Inference: The scene flow estimation is formulated as the energy minimization problem in Eq. (3). Assuming a suitable initialization, that will be discussed in the next section, an iterative optimization approach can be applied. Following the proposed optimization of the object scene flow methods , particle max-product belief propagation is used for 10 iterations with 5 particles for each 6D rigid body motion and 10 particles for each 3D normal vector.

Initialization: The optimization problem needs a suitable initialization of all variables. In the first step, the set of rigid bodies is initialized including their scale-aware 6D motions. Traditionally, the known camera height or an additional inertial measurement unit is used for scale-aware monocular visual odometry in the automotive domain. However, this only provides scale information for the camera ego-motion. The key idea applied here is to integrate single-view depth information to provide the metric scale. In contrast to , we apply this idea additionally for scale-aware pose estimation of moving objects. First, object instances in the images I0I_{0} and I1I_{1} detected by a Mask R-CNN (implementation of ) are paired based on sparse flow correspondences (p0i,p1i)(\mathbf{p}^{i}_{0},\mathbf{p}^{i}_{1}) using a simple voting scheme. Each object instance, as well as the background, builds a rigid body. Second, the 6D motion Tj∈SE(3)\mathbf{T}_{j}\in SE(3) of each rigid body is optimized jointly with a set of 3D points Xi∈X\mathbf{X}_{i}\in\mathcal{X} (one for each flow correspondence lying in the corresponding instance masks) by minimizing

Φtproj(pti,Xi,Tj)\Phi_{t}^{proj}(\mathbf{p}^{i}_{t},\mathbf{X}_{i},\mathbf{T}_{j}) is the reprojection error of Xi\mathbf{X}_{i} with respect to the flow-based image positions pti\mathbf{p}^{i}_{t} weighted by Θ4\Theta_{4}. Φtsvd(Xi,Tj)\Phi_{t}^{svd}(\mathbf{X}_{i},\mathbf{T}_{j}) rates the consistency of the 3D points Xi\mathbf{X}_{i} to the ProbDepthNet estimates analogously to Eq. (6). The energy term of Eq. (10) is optimized using the Levenberg-Marquardt solver implemented in .

Subsequently, the set of 3D planes is initialized. First, a dense depth map is computed based on a semi-global matching adapted to the monocular case similarly to . Again, the depth estimates are additionally rated by the ProbDepthNet estimates. Second, the superpixels including their 3D normal ni\mathbf{n}_{i} are initialized using the approach in . The pixels of a plane are enforced to be of the same instance to get a unique association with a rigid body.

Experiments

In the first part of this section ProbDepthNet is analyzed: Qualitative results of ProbDepthNet are shown, the generalization capabilities to other datasets are presented and an ablation study confirms the importance of the recalibration technique to provide well-calibrated distributions. In the second part, the Mono-SF optimization framework is evaluated by showing qualitative results and a quantitative evaluation with respect to several state-of-the-art methods. Additionally, two ablation studies confirm the claimed ProbDepthNet design for Mono-SF and support the importance of the individual components of Mono-SF.

The experiments are conducted on a ProbDepthNet model trained for the KITTI scene flow training set . The model is trained on 33 sequences of the KITTI raw dataset that are not part of the scene flow set. Around 75% / 25% of the sequences are used for training DepthNet / CalibNet. It is trained for 15 epochs using Adam optimizer with a learning rate of 10−410^{-4} halved every 5 epochs and a small batch size of 4. The input images are scaled to a size of 512×256512\times 256 and a mixture of Gaussians with 8 components is used.

The following ablation study analyzes the proposed recalibration by adding the CalibNet trained on a hold-out split. Our proposed training by minimizing the negative log-likelihood (NLL) is related to the approach in . But, to provide a comparison of different probabilistic approaches, the DepthNet part is also trained using a multi-hypothesis strategy (’Hypo ’) similar to or transformed to its ’assumed density filtering’-counterpart (’ADF’) as proposed by . Fig. 6 shows the mean NLL on the KITTI scene flow set (which is not part of the training data) every 1000 training steps. In the bottom plot of Fig. 6, the calibration of the final models is evaluated. The frequency of ground truth depth values inside a given interval should be the same as the cumulative probability of the estimated distribution. The impact of overfitting effects varies among the different approaches – but all approaches suffer from such an effect and provide overconfident estimates. Furthermore, CalibNet is validated as an useful recalibration technique applicable to different probabilistic approaches.

For integration in Mono-SF, a model is additionally pre-trained on Cityscapes . Compared to previous non-probabilistic methods for single-view depth estimation such as , the main benefit of ProbDepthNet is providing well-calibrated depth distributions. However, in addition to correct uncertainties, the underlying estimates should have sufficient quality as well. A quantitative evaluation (see supplementary material) shows that the accuracy of the depth estimates represented by the total means of the distributions is comparative to and slightly below .

2 Monocular Scene Flow

Mono-SF estimates the 3D scene flow from monocular images focusing on dynamic traffic scenes, which means providing the 3D position and 3D motion of each pixel. The following results and evaluations are based on the equivalent representation as the depth of each pixel at both times (t=0t=0, t=1t=1) and the optical flow. Thereby, the 3D position and the ability of the approaches to predict a 3D point from t=0t=0 to t=1t=1 based in its 3D motion is evaluated. Exemplary qualitative results of Mono-SF are shown for the KITTI (see Fig. 7) and Cityscapes dataset (see Fig. 8). Please see the supplementary material for further results.

The quantitative evaluation is based on the KITTI scene flow dataset , which reports the frequencies of errors for the depth at time t=0t=0 (D1) and t=1t=1 (D2) and the optical flow (Fl). An estimate is considered as an error if it exceeds a threshold of 3 pixels and 5% in terms of stereo disparity or optical flow endpoint error. Furthermore, an estimate is only defined as a valid scene flow estimate (SF) if it fulfills all the D1, D2, and Fl metrics. All metrics are evaluated separately for moving objects (fg), the static scene (bg) and both combined (all).

We propose four categories of state-of-the-art monocular baseline methods. In the first category are the multi-task networks, GeoNet , DF-Net and EveryPixel . These CNNs are trained in an unsupervised manner and are able to provide single-view depth estimates for both images and optical flow estimates. For the GeoNet and DF-Net, their published code and models are used. The results of the EveryPixel approach are stated in their paper (D2 metric is excluded as it seems to be inconsistent). As a second category, single-view depth estimation (’LRC’ or ’DORN ’) and optical flow estimation (’MirrorFlow ’ or ’HD3-F ’) are combined as individual tasks. Due to the fact that the published models of ’DORN ’ and ’HD3-F ’ used parts of the dataset for training, these methods are disregarded for ranking. The third group comprises the multi-body or non-rigid SfM-based methods DMDE and S.Soup . The fourth category consists of the methods MFA , Mono-Stixel and our Mono-SF approach, which are methods that fuse single-view depth information with multi-view geometry. DMDE, S.Soup, and MFA were only evaluated on its depth estimates capped at 50m using a mean absolute relative error (MRE). For the Mono-Stixel approach, the authors provide us the results on a scene flow metric using MirrorFlow and LRC as inputs. The results of the quantitative evaluation are shown in Table 1. To the best of our knowledge, it is the first time that these methods are evaluated and compared as a scene flow estimation problem. The results show that the methods of the fourth group that combine single-view depth and multi-view geometry outperforms the other methods. Mono-SF shows the best rating on most of the metrics and especially outperforms previous methods on the scene flow (SF) metrics. The approach and implementation of Mono-SF is currently not focused on runtime and needs around 41 seconds per image on a single CPU-core. Mono-SF was also submitted to the KITTI scene flow benchmark (see Table 2). Mono-SF is the first monocular method and would have been ranked at the 13th place with respect to the 21 published stereo scene flow methods.

3 Ablation Studies

To analyze the importance of the proposed ProbDepthNet design, the results of four Mono-SF variants based on different single-view depth estimations are provided in Table 3. The two Mono-SF variants ”Mono-SF (LRC )” and ”Mono-SF (w/o prob. depth)” utilized CNNs that provide only single-view depth values instead of depth distributions. Whereas ”Mono-SF (LRC )” is based on the LRC method for single-view depth estimation, ”Mono-SF (w/o prob. depth)” is based on the non-probabilistic estimates of ProbDepthNet represented by the total means of the distributions. The depth values are integrated by assuming the same Gaussian distribution (determined on a test set) for all pixels. Mono-SF based on the probabilistic ProbDepthNet (”Mono-SF”) outperforms both. This supports the claimed ProbDepthNet design to provide single-view depth estimates in a probabilistic form. Furthermore, the improvements compared to a variant based on the ProbDepthNet excluding CalibNet ”Mono-SF (w/o recalib.)” support that the recalibration technique is an essential component.

In Table 4, the individual components of the Mono-SF optimization framework are analyzed by removing some parts of the proposed energy minimization problem (setting their weights to zero). The initialization of Mono-SF described in Sec. 3.2 is denoted by the row without checkmarks. Compared to this initialization, the scene flow formulation of Mono-SF results in further improvement. Additionally, the ablation study shows that each part of the energy term contributes to the final performance; the multi-view geometry, the single-view depth information and the scene model smoothness priors.

Conclusion

In this paper, we proposed Mono-SF for joint estimation of the 3D geometry and motion of particularly traffic scenes by combining multi-view geometry with single-view depth information. For a sensible statistical integration, we showed the importance of providing single-view depth information in a probabilistic and well-calibrated form, which is made possible by our proposed ProbDepthNet including a novel recalibration technique.

References