Uncertainty-Aware CNNs for Depth Completion: Uncertainty from Beginning to End

Abdelrahman Eldesokey, Michael Felsberg, Karl Holmquist, Mikael Persson

Introduction

The recent surge of deep neural networks (DNNs) has led to remarkable breakthroughs on several computer vision tasks, e.g. object classification and detection , semantic segmentation , and object tracking . However, this was achieved at the cost of increased model complexity, inducing new concerns such as: how do these black-box models infer their predictions? and how certain are they about these predictions? Failing to address these concerns impairs the reliability of DNNs. For instance, Huang et al. showed that it is possible to fool state-of-the-art object detectors to produce false and highly certain predictions using physical and digital manipulations. Therefore, there is a compelling need for investigating interpretability and uncertainty of DNNs to be able to trust them in safety-critical environments.

Recently, a growing attention was given towards untangling the complexity of DNNs to enhance their reliability by analyzing how they make predictions and quantifying the uncertainty of these predictions. Probabilistic approaches such as Bayesian deep learning (BDL) have contributed to this endeavor by modifying DNNs to output the parameters of a probabilistic distribution, e.g. mean and variance, which yields uncertainty information about the predictions . The availability of a reliable uncertainty measure facilitates the understanding of DNNs and applying safety procedures in case of model failure or high uncertainty. Several BDL approaches were proposed for different computer vision tasks such as object classification and segmentation , optical flow , and object detection . All these approaches assume undisturbed dense input images, but to the best of our knowledge, there exist no statistical approach that addresses sparse problems.

An essential task of this type is scene depth completion. Modeling uncertainty for this task is crucial due to the inherent noisy and sparse nature of depth sensors, caused by multi-path interference and depth ambiguities . Previous approaches proposed to learn some intermediate confidence masks to mitigate the impact of disturbed measurements inside their networks . However, none of these approaches has demonstrated the probabilistic validity of the intermediate confidence masks. Moreover, they do not provide an uncertainty measure for the final prediction. Therefore, it is still an open problem to fully model the uncertainty in DNN approaches to scene depth completion.

Gustafsson et al. made an attempt by evaluating two of the existing BDL approaches for dense regression problems, i.e. MC-Dropout and ensembling , on the task of depth completion. They utilized the Sparse-to-Dense network as a baseline and modified it to estimate the parameters of a Gaussian distribution. Experiments on the KITTI-Depth dataset showed that both approaches can produce high-quality uncertainty maps for the final prediction, but with the prediction accuracy severely degraded compared to the baseline model. Besides, both approaches train an ensemble of the baseline model requiring multiple inferences during test time. This leads to computational and memory overhead making these approaches unsuitable for the task of depth completion in practice due to their poor prediction accuracy and computational inefficiency.

Specifically designed for confidence-accompanied and sparse data are the normalized convolutional neural networks (NCNNs) . NCNNs consist of a serialization of confidence-equipped convolution layers that make use of an input confidence map. These layers produce the output of the convolution operation as well as an output confidence that is propagated to the following layer. When applied to the problem of depth completion, input confidences at the first layer are assumed to be binary following , ones at valid input points and zeros otherwise. However, this assumption is problematic since depth data can be disturbed as noted in the KITTI-Depth dataset . Therefore, the use of binary masks for modeling input uncertainty in NCNNs becomes inappropriate, and hinders their use as the true input confidence is unknown. Also, the output confidence from NCNNs according to lacks any probabilistic interpretation that qualifies it as a reliable uncertainty measure.

In this paper, we propose two main contributions. First, we employ the inherent dependency of NCNNs on the input confidence to train an estimator for this confidence in a self-supervised manner. Since disturbed measurements are expected to increase the prediction error, we back-propagate the error gradients to learn the input confidence that minimizes the error. This way, the network learns to assign low confidences to disturbed measurements that increase the error and high confidences to valid measurements. This approach establishes a new methodology for handling sparse and noisy data by suppressing the disturbed measurements before feeding them to the network. As shown empirically, this approach is more interpretable and efficient than utilizing a complex black-box model that is expected to implicitly rectify for the disturbed measurements.

Second, we derive a probabilistic NCNN (pNCNN) framework that produces meaningful uncertainty estimates in the probabilistic sense, whereas the output confidence from the standard NCNNs lacks any probabilistic characteristics. We formulate the training process as a maximum likelihood estimation problem and we derive the loss function for pNCNN training. These reformulations are the necessary extensions for fully Bayesian NCNNs.

By applying our approach to the task of unguided depth completion on the KITTI-Depth dataset , we achieve a remarkably better prediction accuracy at a very low computational cost compared to the existing BDL approaches. Moreover, the quality of the uncertainty measure from our single network is better than BDL approaches with ensembles of 1-32 networks. When compared against non-statistical approaches, we perform on par with state-of-the-art methods with millions of parameters using a significantly smaller network (670k parameters). Besides, and contrarily to state-of-the-art methods, we produce a high-quality prediction uncertainty measure aside with the prediction. Finally, we show that our approach is applicable to other sparse problems by evaluating it on multi-path interference correction and sparse optical flow rectification.

Related Work

The task of scene depth completion is receiving an increasing attention due to the impact of depth information on different computer vision tasks. Typically, it aims to produce a dense and denoised depth map y\mathbf{y} from a noisy sparse input x\mathbf{x}. Several approaches were proposed to learn a mapping y=f(x)\mathbf{y}=f(\mathbf{x}) by exploiting different input modalities, where ff is a DNN. Ma et al. proposed a deep regression model that combines the sparse input depth with the corresponding RGB modality. Jaritz et al. evaluated different fusion schemes to combine the sparse depth with RGB images. Chen et al. proposed a joint network that exploits 2D and 3D representations for the depth data. The key similarity between these approaches is that they all perform very well in terms of prediction accuracy and they implicitly handle disturbed measurements in the network. Nonetheless, none of these methods considered modeling the uncertainty of the data or the prediction.

Recently, several approaches promoted the use of confidences to filter out noisy predictions within the network. Qui et al. learned confidence masks from RGB images to mask out noisy depth measurements at occluded regions. Gansbeke et al. proposed the use of confidences to fuse two network streams utilizing sparse depth and RGB images respectively. Similarly, Xu et al. predict a confidence mask that is used to mitigate the impact of noisy measurements on different components of their network. However, none of these methods provided any prediction uncertainty measure for the final prediction.

This was addressed by another approach that utilizes confidences and provides an output confidence for the final prediction. Normalized convolutional neural networks (NCNNs) take sparse depth x\mathbf{x} and a confidence mask c0\mathbf{c}^{0} as input, propagate the confidence, and produce a dense output y\mathbf{y} as well as an output confidence map cL\mathbf{c}^{L}, i.e., (y,cL)=f(x,c0)(\mathbf{y},\mathbf{c}^{L})=f(\mathbf{x},\mathbf{c}^{0}), for a DNN with LL layers. However, since the input confidence is unknown, a binary input confidence c0\mathbf{c}^{0} is assumed, which is problematic in case of disturbed input as shown in Figure (1a). Further, the output confidence cL\mathbf{c}^{L} has no probabilistic interpretation and shows no significant correlation with the prediction error.

To address these challenges, we look at the problem from a different perspective. We propose to learn the input confidence from the disturbed measurements by employing the confidence propagation property of NCNNs. We attach a network hh to a NCNN and we train them end-to-end to learn the input confidence that minimizes the prediction error, i.e., (y,cL)=f(x,h(x))(\mathbf{y},\mathbf{c^{L}})=f(\mathbf{x},h(\mathbf{x})). Further, to produce accurate uncertainty measure for the final prediction, we derive a probabilistic version of the NCNNs and we formulate the training as a maximum likelihood problem. When our proposed approach is evaluated on the KITTI-Depth dataset , it performs on par with state-of-the-art approaches with millions of parameters using a significantly smaller network, while providing a highly accurate uncertainty measure for the final prediction. In contrast to BDL approaches in , we achieve excellent uncertainty estimation without sacrificing prediction accuracy or computational efficiency.

The rest of the paper is organized as follows. We briefly describe the method of NCNNs in 3.1 and 3.2, and our proposed approach for learning the input confidence in section 3.3. Afterwards, we introduce a probabilistic version of NCNNs, derive the loss for training, and describe our architecture in section 4. Experiments and analysis are given in section 5. Finally, we conclude the paper in section 6.

Self-supervised Input Confidence Learning

The signal/confidence philosophy promotes the separation between the signal and its confidence for efficiently handling noisy and sparse signals. For example, this separation allows differentiating missing signal points with no information from zero-valued valid points. The normalized convolution is one approach that follows the this philosophy to perform the convolution operation.

For confidence-equipped signals, the normalized convolution performs convolution using only the confident points of the signal, while estimating the non-confident ones from their vicinity using some applicability function. This prevents noisy and missing measurements from disturbing the calculations. In this section, we give a brief description of normalized convolution and the trainable normalized convolution layer that can estimate an optimal applicability . Subsequently, we propose a novel approach to learn the input confidence in a self-supervised manner.

Throughout the paper, we assume a global signal Y\mathcal{Y} with a finite size NN that is convolved in a sliding window fashion. At each point in the signal yiy_{i}, a local signal y\mathbf{y} of size nn constitutes the neighborhood at this point. The local signal y\mathbf{y} will be referred to as the signal, and yiy_{i} will be referred to as the signal center.

If we arrange the basis functions into the columns of a matrix B\mathbf{B}, then the image of the signal under the subspace spanned by the basis is obtained as y=Br\mathbf{y}=\mathbf{B}\mathbf{r}, where r\mathbf{r} is a vector of coordinates. These coordinates can be estimated from a weighted least-squares problem (WLS) between the signal y\mathbf{y} and the image of it under the new basis:

where the weights matrix W\mathbf{W} is a product of Wa=diag(a)\mathbf{W}_{\mathbf{a}}=\text{diag}(\mathbf{a}) and Wc=diag(c){\mathbf{W}_{\mathbf{c}}=\text{diag}(\mathbf{c})}. The WLS solution is given as :

Finally, the WLS solution r^WLS\hat{\mathbf{r}}_{\text{WLS}} can be used to approximate the signal under the new basis as:

2 Normalized Convolutional Neural Networks

In normalized convolution, the applicability is chosen manually. Eldesokey et al. proposed a normalized convolutional neural network layer (NCNN) that utilized the standard back-propagation in DNNs to learn the optimal applicability function a\mathbf{a} for a given dataset, while assuming a binary input confidence. This was achieved by using the naïve basis in (2), i.e. B=1n\mathbf{B}=\mathbf{1}_{n}:

where 1n\mathbf{1}_{n} is a vector of ones, ⊙\odot is the Hadamard product, ⟨.∣.⟩\langle.|.\rangle is the scalar product, r^i\hat{r}_{i} is a scalar which is equivalent to the estimated value at the signal center y^i\hat{y}_{i}. They proposed to propagate the confidence from the NCNN layer as:

where the output confidence from one layer is the input confidence to the next layer.

3 Self-Supervised Input Confidence Estimation using NCNNs

The assumption of binary input confidences adopted by can be problematic in real datasets. An example is the KITTI-Depth dataset , where some of the input values do not match the groundtruth due to LiDAR projection errors (shown in Figure 4 top). In this case, a binary input confidence would lead to artifacts in the output as NCNNs are dependent on the input confidence as shown in the calculations of (4). This dependency of the outputs on the input confidences facilitates learning the confidences. The inclusion of the input confidences in the calculations of the output from each layer indicates that the loss of the network would constitute gradients with respect to these confidences. Therefore, we can employ these gradients to learn input confidences that minimize the loss function.

We propose to use an input confidence estimation network that receives the input data and produces an estimate for the input confidence that is fed to the first layer of the NCNN. This network is trained end-to-end with the NCNN and the error gradients from the NCNN are back-propagated to the confidence estimation network, allowing it to learn the input confidence that minimizes the overall prediction error. We use a compact UNet for the confidence estimation network with a Softplus activation at the final layer that will produce valid confidence values in the interval [0,∞[[0,\infty[. The pipeline is illustrated in Figure 2 (upper part).

Probabilistic NCNNs

Figure (1b) shows an example of the output confidence from the last NCNN layer when we estimate the input confidences using our proposed approach from the previous section. The figure shows that the output confidences do not exhibit a proper uncertainty measure that is strongly correlated with the error.

To obtain proper uncertainties from NCNNs, we introduce a probabilistic version of NCNNs by deriving the connection between the normalized convolution and statistical least-squares approaches. Then, we utilize this connection to produce reliable uncertainties with probabilistic characteristics. Finally, we apply the proposed theory to NCNNs and we derive a loss function for training them to produce accurate uncertainties.

In ordinary least-squares (OLS) problems, constant variance is assumed for all observations of the signal. Generalized least-squares (GLS), on the other hand, offers more flexibility to handle individual variance per observation. The weighted-least squares problem in (2) can be viewed as a special case of the GLS, where observations are heteroskedastic with unequal noise levels.

Assume the image of the signal under the subspace B\mathbf{B} is defined as y=Br+e{\mathbf{y}=\mathbf{B}\mathbf{r}+\mathbf{e}}, where e\mathbf{e} is a random noise variable with zero mean and variance σ2V\sigma^{2}\mathbf{V}. This variance models the heteroscedastic uncertainty of the observations in the signal, where σ2\sigma^{2} is global for each signal, and V\mathbf{V} is a positive definite matrix describing the covariance between the observations. The GLS solution to this problem reads :

When comparing the two solutions in (2) and (6), they are only equivalent if V−1\mathbf{V^{-1}} is diagonal, which leads to V=(WaWc)−1\mathbf{V}=(\mathbf{W}_{\mathbf{a}}\mathbf{W}_{\mathbf{c}})^{-1}. The diagonality of the covariance matrix indicates that different samples of the signal are independent and have different variances depending on the confidence and the applicability function.

We utilize the GLS solution r^GLS\hat{\mathbf{r}}_{\text{GLS}} to estimate the signal similar to (3) as y^=Br^GLS\hat{\mathbf{y}}=\mathbf{B}\hat{\mathbf{r}}_{\text{GLS}}. The uncertainty of y^\hat{\mathbf{y}} can be estimated as:

Note that Wa and Wc\mathbf{W}_{\mathbf{a}}\text{ and }\mathbf{W}_{\mathbf{c}} are non-stochastic, where the former is estimated during NCNN training and the latter can be learned using our proposed approach in section 3.3. On the other hand, σ2\sigma^{2} is unknown and needs to be estimated.

2 Output Uncertainty for NCNNs

In case of NCNNs with the naïve basis B=1n{\mathbf{B}=\mathbf{1}_{n}}, the uncertainty measure in (7) simplifies to:

This indicates an equal uncertainty for the whole neighborhood, but since we are only interested in signal center y^i\hat{y}_{i}, (8) reduces to:

It is evident that the output confidence described in (5) disregards the stochastic noise variance σi2\sigma_{i}^{2}. However, to obtain a proper uncertainty measure, this variance needs to be incorporated in the output confidence. We propose to estimate the noise variance σi2\sigma_{i}^{2} from the output confidence of the last NCNN layer by means of a noise variance estimation network as illustrated in Figure 2. To achieve this, we need a loss function that allows training the proposed framework.

3 The Loss Function for Probabilistic NCNNs

We consider each point yiy_{i} in the global signal Y\mathcal{Y}, where the neighborhood at this point is the local signal y\mathbf{y}. This local signal can be represented under some basis as y^=Br^\hat{\mathbf{y}}=\mathbf{B}\hat{\mathbf{r}}, where the estimated coordinates r^\hat{\mathbf{r}} are calculated from (6,2). We assume that the estimate of the signal follows a multivariate normal distribution y^∼Nm(Br^,σ2B(B∗WaWcB)−1B∗){\hat{\mathbf{y}}\sim\mathcal{N}_{m}(\mathbf{B}\hat{\mathbf{r}},\sigma^{2}\mathbf{B}(\mathbf{B}^{*}\mathbf{W}_{\mathbf{a}}\mathbf{W}_{\mathbf{c}}\mathbf{B})^{-1}\mathbf{B}^{*})} where the variance is defined in (7). In case of the naïve basis, we will have a univariate normal distribution y^i∼N(r^i,σi2/⟨a∣c⟩){\hat{y}_{i}\sim\mathcal{N}(\hat{r}_{i},\sigma_{i}^{2}/\langle\mathbf{a}|\mathbf{c}\rangle)}, where the variance is defined in (9). More formally, a NCNN outputs the mean r^iL\hat{r}_{i}^{L} of the normal distribution around y^i\hat{y}_{i}, and the scalar product ⟨a∣c⟩\langle\mathbf{a}|\mathbf{c}\rangle in the denominator of the variance. Yet, the noise variance σ2\sigma^{2} needs to be estimated to comply with the definition in (9).

We denote the variance term as si=σi2/⟨a∣c⟩s_{i}=\sigma_{i}^{2}/\langle\mathbf{a}|\mathbf{c}\rangle, where a and c\mathbf{a}\text{ and }\mathbf{c} are the applicability and the output confidence from the last NCNN layer. The least squares solution in (4) can be formulated as a maximum likelihood problem of a Gaussian error model for the last NCNN layer LL:

where w\mathbf{w} denotes the network parameters, and r^iL\hat{r}_{i}^{L} is calculated based on (4). By taking log likelihood of (10) instead, we obtain:

The first term is a constant and is ignored, and the cost function is defined as minimizing the negative log likelihood:

where the scalar 1/2{1/2} has been discarded. This cost function shares similarity with the aleatoric uncertainty loss proposed in . The difference is that sis_{i} in our case depicts an uncertainty measure that encodes observation noise variance and the output confidence from NCNN, while in , it is the variance of the noise. Note that this cost function can be derived using any error model from the exponential family, e.g. Laplace distribution as in . Next, we show the architecture design that is used for training our proposed probabilistic approach.

4 Probabilistic NCNN Architecture

Given a dataset that contains undisturbed data Y\mathcal{Y} as groundtruth and a disturbed version Y˙\dot{\mathcal{Y}} as input, we aim to train a network that produces the clean data given the disturbed one. An illustration for our full pipeline is shown in Figure 2. The first component estimates the input confidence from the disturbed input and feed both of them to the NCNN network. The output confidence from the last NCNN layer is fed to another compact UNet to estimate the noise parameter σi2\sigma^{2}_{i} and to produce sis_{i} in (1). Finally, the prediction from the NCNN network and the estimated uncertainty sis_{i} are fed to the loss.

Note that the noise variance estimation network takes only the output confidence from the NCNN as input, contrarily to existing approaches that estimate the uncertainty from the final prediction . This indicates that our confidences can efficiently encode the uncertainty information, which is also demonstrated in the experiments section.

Experiments

To demonstrate the capabilities of our proposed approach, we evaluate it on the KITTI-Depth dataset for the task of unguided depth completion (no RGB guidance is used). We first compare against Bayesian Deep Learning approaches, e.g. MC-Dropout and ensembling , in terms of prediction accuracy and the quality of the uncertainty measure. Then, we show comparison against the conventional non-statistical approaches. Afterwards, we perform an ablation study for different components of our pipeline and we experiment with an ensemble of our proposed network. Finally, we demonstrate the generalization capabilities of our approach by evaluating it on multi-path interference correction and optical flow rectification. The source code is available on Github https://github.com/abdo-eldesokey/pncnn.

Our pipeline is illustrated in Figure 2 and more details are given in the supplementary materials. We evaluate three variations of our network: our network where only the input confidence estimation part that is trained using the L1 or the L2 norm (NCNN-Conf), our full network trained with the proposed loss in (1) (pNCNN), and our full network trained with a modified version of the loss in (1), where we apply an exponential function to sis_{i} in the data term (pNCNN-Exp). This modification is to robustify our loss to outliers violating the presumed Gaussian error model for the data term. Training was performed using the Adam optimizer with an initial learning rate of 0.010.01 that is decayed with a factor of 10−110^{-1} every 3 epochs.

Evaluation Metrics We use the following two measures:

Prediction Error We use the error metrics from the KITTI-Depth such as Mean Average Error (MAE), Root Mean Square Error (RMSE) and their inverses.

Quality of Uncertainty We use the sparsitification error plots and the area under sparsification error plots (AUSE) as a measure for the quality of the uncertainty.

2 Results Compared to Statistical Methods

Gustafsson et al. evaluated the MC-Dropout and ensembling by modifying the head of the Sparse-to-Dense (S2D) network to output the parameters of a Gaussian distribution. They evaluated an ensemble of 1-32 instances of S2D with 26M parameters each an taking the mean of these instances for the final prediction. Note that their network utilizes both depth and RGB images, while our approach consist of a single network that is fully unguided and uses only depth data.

Figure 3 shows a two-metric comparison with respect to AUSE and RMSE. Our NCNN-Conf performs best in terms of RMSE, while it performs worst in terms of AUSE. On the other hand, our full network trained with the proposed loss, pNCNN, produces the best uncertainty measure with an AUSE of 0.053 outperforming an ensemble of 32 networks. Moreover, it achieves a significantly lower RMSE than MC-Dropout and ensembling. However, it performs inferior to NCNN-Conf in terms of RMSE with a moderate gap. The variation of our network that is trained with a modified loss, pNCNN-Exp, closes this gap and performs on-par with NCNN-Conf in terms of RMSE with a minor degradation of AUSE compared to pNCNN.

3 Results Compared to Non-Statistical Methods

We also compare our proposed approach against the non-statistical unguided approaches. Table 1 summarizes the results on the test set of the KITTI-Depth dataset. Our NCNN-Conf-L1 outperforms all other methods on three out of four metrics when compared individually, except for Spade, where we are better on two metrics and on-par on one metric. Note the improvement of our approach over the standalone NCNN, where we achieve a performance boost of ∼45%\sim 45\% by providing more accurate input confidences. Our probabilistic model trained using a Gaussian error model and a Laplace error model, pNCNN-Exp trained with the modified loss performs equally good to the NCNN-Conf-L2, but additionally providing proper output uncertainties.

4 Ablation Study

First we show the impact of each component of our proposed network on a qualitative example from the KITTI-Depth dataset. Figure 4 shows an example where the input measurements do not coincide with the groundtruth. The standard NCNN assigns 1-confidences to all measurements, which results in a corrupted prediction (first row). When we apply our input confidence estimation, the disturbed measurements are successfully identified and assigned zero confidence (second row). However, the output confidence is almost identical to the input confidence and shows no strong correlation with the accuracy. When we apply our full pipeline, the disturbed measurements are identified and the output uncertainty becomes highly correlated with the prediction error (third row).

Next, we show in Table 2 the impact of modifying different components of our pipeline. When the confidence estimation is discarded in w/o conf-est and binary input confidence is used, the RMSE is degraded, while the network still manages to achieve good AUSE. Similarly, when the noise variance estimation network is discarded in w/o var-est, the RMSE is severely degraded as the input confidence estimation network tries to make up for the absence of the variance estimation network. When the final prediction from the NCNN is fed along with the output confidence to the noise variance estimation network in w depth-pred, no improvement is gained in terms of AUSE. This demonstrates that our uncertainty measure efficiently encode the uncertainty information in the NCNN confidence stream without looking at the prediction. Finally, when we employ a Laplace error model for the loss in w Laplace-loss, i.e., the L1 norm for residuals, the MAE improves, while AUSE is degraded since it is calculated based on the RMSE.

5 Ensemble of pNCNN

To examine whether our probabilistic approach can be extended to a fully Bayesian approach, we form an ensemble of four pNCNN network that were initialized randomly and trained on random subset of the KITTI-Depth dataset. We evaluate multiple fusion approaches which are summarized in Table 3. Fusion by selecting the most confident pixel from each network, maxConf, achieves the best results, outperforming taking the mean, which is commonly used. Taking a weighted mean using confidences, wMean, or a maximum likelihood estimation, MLE, also gives better results than the standard mean. This demonstrated the potential of using the proposed output confidences in more sophisticated fusion schemes.

6 Mutli-Path Interference (MPI) Correction

To demonstrate the generalization capabilities of our approach on other kinds of noise, we evaluate it on depth data from a Time-of-Flight (ToF) camera, i.e. Kinect2, that suffers from MPI. We use the FLAT dataset for this purpose which provides raw measurements for three different frequencies and phases. We use the libfreenect2 to calculate the depth from the measurements and we compare against applying the bilateral filtering on the noisy depth.

Table 4 summarizes the results, where we outperform the Bilateral filtering with a significant margin in terms of RMSE error when evaluated both on noisy and clean data with no MPI. Bilateral filtering on the other hand performs worse than doing no processing as it assigns zeros to pixels close to edges. When edges are not considered for evaluation, bilateral filtering improves the results slightly, but is outperformed by our approach.

7 Sparse Optical Flow Rectification

We generate the input flow by applying the Lucas-Kanade method to pairs of images from driving sequences. The groundtruth is produced by geometrical verification over several frames under a multiple rigid body assumption . Figure 5 shows an example for rectifying the corrupted measurement and densifying the flow field. More results are given in the supplementary materials.

8 What happens if the input is undisturbed?

An essential question is how our confidence estimation network will perform if the input data is not disturbed? To answer this question, we train our network NCNN-Conf and pNCNN on the NYU dataset , where the input is sampled from the groundtruth depth. We use 1000 depth points sampled uniformly with a sparsity level of 0.6%0.6\%. Figure 7 and Table 7 show that both our methods surprisingly improves the results compared to the standalone NCNN . This is a result of allowing the confidence estimation network to assign proper confidences to points based on their proximity to edges similar to non-linear filtering. This leads to sharper edges and better reconstruction of objects.

Conclusion

We proposed a self-supervised approach for estimating the input confidence for sparse data based on the NCNNs. We also introduced a probabilistic version of NCNNs that enable the to output meaningful uncertainty measures. Experiments on the KITTI dataset for unguided depth completion showed that our small network with 670k parameters achieves state-of-the-art results in terms of prediction accuracy and it provides an accurate uncertainty measure. When compared against the existing probabilistic method for dense problems, our proposed approach outperforms all of them in terms of the prediction accuracy, the quality of the uncertainty measure, and the computational efficiency. Moreover, we showed that our approach can be applied to other sparse problems as well. These results demonstrate the gains from adhering to the signal/uncertainty philosophy compared to conventional black-box models.

Acknowledgments: This work was supported by the Wallenberg AI, Autonomous Systems and Software Program (WASP) and Swedish Research Council grant 2018-04673.

References

Implementation Details

In this section, we give more details on the implementation of our proposed method such as the loss function and the design of the confidence estimation and the noise variance estimation networks.

We drove a loss function to train the proposed probabilistic normalized convolutional neural networks (pNCNN), which reads:

where sis_{i} is the proposed uncertainty measure and it is equal to σi2/⟨a∣c⟩\sigma_{i}^{2}/\langle\mathbf{a}|\mathbf{c}\rangle. For convenience and numerical stability, we modify the regularization term so that sis_{i} becomes consistent with the data term. This leads to:

This can be expanded using the definition of sis_{i}:

where aL,cL\mathbf{a}^{L},\mathbf{c}^{L} are the learned applicability and the output confidence of the last normalized convolution layer LL respectively. This expansion makes it clear that our proposed uncertainty measure depends both on the output confidence from the normalized convolution layer and observations noise variance. A higher noise variance will reduce the output confidence from the NCNN and vice versa. This indicates that our proposed uncertainty measure encodes the single observation noise as well as the confidence with respect to the neighboring pixels.

2 The Architecture

We propose to learn the input confidence using a compact UNet that is trained end-to-end with a normalized convolutional neural network (NCNN) . We also learn observations noise variance using a similar UNet. The design of this UNet is shown in Figure 8 and it is identical for both networks. It is worth mentioning that this network has only 3 scales compared to original UNet which has 4 scale, since we found empirically that the 4th scale does not improve the estimation. The number of channels per convolution layer was significantly reduced for computational efficiency.

The choice of the activation for the last layer is crucial since it must produce valid range of values for confidences [0,∞[\left[0,\infty\right[. We choose the SoftPlus function (Shown in Figure 8) due to its similarity to the ReLU activation. However, it does not suffer from the gradient discontinuity at zeros.

Ensemble methods

In the main paper, we evaluate different fusion schemes for an ensemble of our network pNCNN. We showed that all fusion schemes utilizing our proposed uncertainty measure outperform the commonly used fusion using the standard mean. Here, we give the definition for the evaluated fusion schemes.

The Mean fusion method refers to the average over the predictions yiky_{i}^{k} at pixel ii:

2 The Weighted Mean

Since the mean fusion does not take into account the uncertainties, we weight the predictions using their confidences cikc_{i}^{k}:

3 Max Voting

Another commonly used voting scheme is to select the most confident prediction ki=arg⁡mmax⁡cimk_{i}=\arg_{m}\max c_{i}^{m}

4 Maximum Likelihood Estimate

We can interpret our predictors as components of a Gaussian Mixture Model. If the prediction corresponds to the mean and the confidence corresponds to the unnormalized mixture weights, we can write the likelihood of a prediction x^\hat{x} given predictions yky^{k} from the networks as:

We can formulate an inference procedure based on the MLE for each pixel ii as:

Optimization Procedure The likelihood function of a Gaussian Mixture Model is in general non-convex. However, for the 1D case, the number of modes is constrained to at most the number of components in the mixture [carreira2003number]. Since it is guaranteed that the global maxima will be found if all local maximas are explored, we optimize the objective starting from each of the predictions. We use the ADAM optimizer with a maximum amount of steps set to 500. And we select the maximum of the local maximas which were found. Note that since we do not explicitly estimate the variances of the components we set v2v^{2} to 0.1 for our experiments.

Additional Results

In this section, we show additional results for all the experiments in the paper. First, we show some qualitative examples on the KITTI-Depth dataset . Then, we show the sparsificiation plots for our proposed uncertainty measure that were used to calculate the AUSE metric. Afterwards, we show some qualitative examples for multi-path interference correction and sparse optical flow rectification. Finally, we show illustrations on the NYU dataset for the case of undisturbed input data.

Figure 9 and 11 show qualitative examples for NCNN , our proposed NCNN-Conf-L1, pNCNN, and pNCNN-Exp from the selected validation set of the KITTI-Depth dataset. NCNN assigns binary confidence to the input, which results in artifacts at regions with disturbed measurements especially edges (indicated with red squares). Our proposed NCNN-Conf-L1 on the other hand, learns a proper input confidence which discards input measurements that causes the prediction error to increase. This causes the final prediction to be artifact-free and sharp along edges. It is worth mentioning that our input confidence estimation learned to discard some of the true measurements (indicated with the white squares) as well in order to produce smoother surfaces. Those discarded measurements are compensated for using other measurements on the end points of the same surface.

It is clear the output confidence from NCNN-Conf-L1 is a densified version of the estimated input confidence. But it does not provide full uncertainty information for all observations in the prediction. Our proposed pNCNN addresses this problem and produces a reliable uncertainty measure for all observations. However, the prediction error at some disturbed measurements increase where the presumed Gaussian error model does not hold (indicated with the red squares if Figure 9). By applying the exponential function to sis_{i} in the data term of the loss in pNCNN-Exp, the network focuses more on minimizing the prediction error for those disturbed measurements and produces a better prediction. Note that the range for the certainty measure changes with pNCNN-Exp due to the exponential scaling.

2 The Quality of the Proposed Uncertainty Measure

To examine the quality of our proposed uncertainty measure, we look at the commonly used sparsification plots . Sparsification plots show how efficiently the uncertainty measure discards the erroneous measurements. The baseline in this case is the prediction error itself, which is denoted as the oracle. Sparsification plots for NCNN-Conf-L1, pNCNN, and pNCNN-Exp are shown in Figure 10. The uncertainty measure from NCNN-Conf-L1 is not correlated with the oracle as the classical normalized convolution framework does not constitute any probabilistic properties. Our proposed probabilistic normalized convolution pNCNN on the other hand, produces an accurate uncertainty measure that is very similar to the error oracle. The modified version pNCNN-Exp also produces a high-quality uncertainty measure, but with a better handling of outliers.

3 Multi-Path Interference Correction

Figure 12 shows two qualitative results for the FLAT dataset. The first row, shows a scene with small areas of missing data. These areas are well handled by the pNCNN and the confidences clearly shows the uncertainty that exist in these areas and on edges. The scene in the second row illustrates the effect of larger areas of missing data. These areas are missing too much data for the network to handle. As such, the output confidences is used to mask these parts of the signal. This illustrates the strength of our formulation in handling both smaller areas were the missing data can be extrapolated and larger areas where high uncertainty is assigned.

4 What happens when the input is undisturbed?

Figure 13 shows some qualitative examples on the NYU dataset for our NCNN-Conf-L1 compared to the standard NCNN . In these examples, the sparse input is undisturbed and NCNN should perform well using the binary input confidences. However, NCNN struggles along edges due to equally trusting the background and the foreground. Our NCNN-Conf-L1 on the other hand, learns proper input confidences that preserve edges similar to non-linear filtering.

5 Sparse Optical Flow Rectification

We include more results for the sparse optical flow rectification to demonstrate the generalization capabilities of our approach to other types of data. Qualitative examples are shown in Figure 14 and 15. Our method successfully removes noisy flow vectors despite the fact that they look completely random. This demonstrates the generalization capabilities of our approach in identifying the inherent noise in the data in a self-supervised manner.