Evaluating Scalable Bayesian Deep Learning Methods for Robust Computer Vision
Fredrik K. Gustafsson, Martin Danelljan, Thomas B. Schön
Introduction
Deep Neural Networks (DNNs) have become the standard paradigm within most computer vision problems due to their astonishing predictive power compared to previous alternatives. Current applications include many safety-critical tasks, such as street-scene semantic segmentation , 3D object detection and depth completion . Since erroneous predictions can have disastrous consequences, such applications require an accurate measure of the predictive uncertainty. The vast majority of these DNN models do however fail to properly capture the uncertainty inherent in their predictions. They are thus not fully capable of the type of uncertainty-aware reasoning that is highly desired e.g. in automotive applications.
The approach of Bayesian deep learning aims to address this issue in a principled manner. Here, predictive uncertainty is commonly decomposed into two distinct types, which both should be captured by the learned DNN . Epistemic uncertainty accounts for uncertainty in the DNN model parameters, while aleatoric uncertainty captures inherent and irreducible data noise. Input-dependent aleatoric uncertainty about the target arises due to e.g. noise and ambiguities inherent in the input . This is present for instance in street-scene semantic segmentation, where image pixels at object boundaries are inherently ambiguous, and in 3D object detection where the location of a distant object is less certain due to noise and limited sensor resolution. In many computer vision applications, this aleatoric uncertainty can be effectively estimated by letting a DNN directly output the parameters of a certain probability distribution, modeling the conditional distribution of the target given the input. For classification tasks, a predictive categorical distribution is commonly realized by a softmax output layer, although recent work has also explored Dirichlet models . For regression, Laplace and Gaussian models have been employed .
Directly predicting the conditional distribution with a DNN does however not capture epistemic uncertainty, as information about the uncertainty in the model parameters is disregarded. This often leads to highly confident predictions that are incorrect, especially for inputs that are not well-represented by the training distribution . For instance, a DNN can fail to generalize to unfamiliar weather conditions or environments in automotive applications, but still generate confident predictions. Reliable estimation of epistemic uncertainty is thus of great importance. However, this task has proven to be highly challenging, largely due to the vast dimensionality of the parameter space, which renders standard Bayesian inference approaches intractable. To tackle this problem, a wide variety of approximations have been explored , but only a small number have been demonstrated to be applicable even to the large-scale DNN models commonly employed in real-world computer vision tasks. Among such scalable methods, MC-dropout and ensembling are clearly the most widely employed, due to their demonstrated effectiveness and simplicity. While scalable techniques for epistemic uncertainty estimation recently have emerged, the research community however lacks a common and comprehensive evaluation framework for such methods. Consequently, both researchers and practitioners are currently unable to properly assess and compare newly proposed methods. In this work, we therefore accept this task and set out to design exactly such an evaluation framework, aiming to benefit and inspire future research in the field.
Previous studies have provided only partial insight into the performance of different scalable methods for epistemic uncertainty estimation. Kendall and Gal evaluated MC-dropout alone on the tasks of semantic segmentation and monocular depth regression, providing mainly qualitative results. Lakshminarayanan et al. introduced ensembling as a non-Bayesian alternative and found it to generally outperform MC-dropout. Their experiments were however based on relatively small-scale models and datasets, limiting the real-world applicability. Ilg et al. compared ensembling and MC-dropout on the task of optical-flow estimation, but only in terms of the AUSE metric which is a relative measure of the uncertainty estimation quality. While finding ensembling to be advantageous, their experiments were also limited to a fixed number () of ensemble members and MC-dropout forward passes, not allowing a completely fair comparison. Ovadia et al. also fixed the number of ensemble members, and moreover only considered classification tasks. We improve upon this previous work and propose an evaluation framework that actually enables a conclusive ranking of the compared methods.
Contributions We propose a comprehensive evaluation framework for scalable epistemic uncertainty estimation methods in deep learning. The proposed framework is specifically designed to test the robustness required in real-world computer vision applications, and employs state-of-the-art DNN models on the tasks of depth completion (regression) and street-scene semantic segmentation (classification). It also employs a novel combination of quantitative evaluation metrics which explicitly measures the reliability and practical usefulness of estimated predictive uncertainties. We apply our proposed framework to provide the first properly extensive and conclusive comparison of the two current state-of-the-art scalable methods: ensembling and MC-dropout. This comparison demonstrates that ensembling consistently outperforms the highly popular MC-dropout method. Our work thus suggests that ensembling should be considered the new go-to approach, and encourages future research to understand and further improve its efficacy. Figure 1 shows example predictive uncertainty estimates generated by ensembling. Our framework can also directly be applied to compare other scalable methods, and we encourage external usage with publicly available code.
In our proposed framework, we predict the conditional distribution in order to estimate input-dependent aleatoric uncertainty. The methods for epistemic uncertainty estimation are then compared by quantitatively evaluating the estimated predictive uncertainty in terms of the relative AUSE metric and the absolute measure of uncertainty calibration. Our evaluation is the first to include both these metrics, and furthermore we apply them to both regression and classification tasks. To provide a deeper and more fair analysis, we also study all metrics as functions of the number of samples , enabling a highly informative comparison of the rate of improvement. Moreover, we simulate challenging real-world conditions found e.g. in automotive applications, where robustness to out-of-domain inputs is required to ensure safety, by training our models exclusively on synthetic data and evaluating the predictive uncertainty on real-world data. By analyzing this important domain shift problem, we significantly increase the practical applicability of our evaluation. We also complement our real-world analysis with experiments on illustrative toy regression and classification problems. Lastly, to demonstrate the evaluation rigor necessary to achieve a conclusive comparison, we repeat each experiment multiple times and report results together with the observed variation.
Predictive Uncertainty Estimation using Bayesian Deep Learning
In regression, the most common approach is to let the DNN directly predict targets, . The parameters are learned by minimizing e.g. the or loss over the training dataset . However, such direct regression does not model aleatoric uncertainty. Instead, recent work has explored predicting the distribution , similar to the classification case above. For instance, can be parameterized by a Gaussian distribution , giving the following model in the 1D case,
Here, the DNN predicts the mean and variance of the target . The variance is naturally interpreted as a measure of input-dependent aleatoric uncertainty. As in classification, the model parameters are learned by minimizing the negative log-likelihood .
Epistemic Uncertainty While the above models can capture aleatoric uncertainty, stemming from the data, they are agnostic to the uncertainty in the model parameters . A principled means to estimate this epistemic uncertainty is to perform Bayesian inference. The aim is to utilize the posterior distribution , which is obtained from the data likelihood and a chosen prior by applying Bayes’ theorem. The uncertainty in the parameters is then marginalized out to obtain the predictive posterior distribution,
Here, the generally intractable integral in (3) is approximated using Monte Carlo samples , ideally drawn from the posterior. In practice however, obtaining samples from the true posterior is virtually impossible, requiring an approximate posterior to be used. We thus obtain the approximate predictive posterior as,
which enables us to estimate both aleatoric and epistemic uncertainty of the prediction. The quality of the approximation (4) depends on the number of samples and the method employed for generating . Prior work on such approximate Bayesian inference methods is discussed in Section 3. For the Categorical model (1), , . For the Gaussian model (2), is a uniformly weighted mixture of Gaussian distributions. We approximate this mixture with a single Gaussian, see Appendix A for details.
Illustrative Example To visualize and provide intuition for the problem of predictive uncertainty estimation with DNNs, we consider the problem of regressing a sinusoid corrupted by input-dependent Gaussian noise,
Training data is only given in the interval $yx^{\star}~{}\in~{}x^{\star}\in|x^{\star}|>3p(\theta)~{}=~{}\mathcal{N}(0,I_{P})M=1\thinspace 000$ samples obtained via Hamiltonian Monte Carlo , is additionally able to predict more reasonable uncertainties in the region with no available training data, see Figure 2(d).
Related Work
Here, we discuss prior work on approximate Bayesian inference. We also note that ensembling, which is often considered a non-Bayesian alternative, in fact can naturally be viewed as an approximate Bayesian inference method.
Approximate Bayesian Inference The method employed for approximating the posterior is a crucial choice, determining the quality of the approximate predictive posterior in (4). There exists two main paradigms for constructing , the first one being Markov chain Monte Carlo (MCMC) methods. Here, samples approximately distributed according to the posterior are obtained by simulating a Markov chain with as its stationary distribution. For DNNs, this approach was pioneered by Neal , who employed Hamiltonian Monte Carlo (HMC) on small feed-forward neural networks. HMC entails performing Metropolis-Hastings updates using Hamiltonian dynamics based on the potential energy . To date, it is considered a “gold standard” method for approximate Bayesian inference, but does not scale to large DNNs or large-scale datasets. Therefore, Stochastic Gradient MCMC (SG-MCMC) methods have been explored, in which stochastic gradients are utilized in place of their full-data counterparts. SG-MCMC variants include Stochastic Gradient Langevin Dynamics (SGLD) , where samples are collected from the parameter trajectory given by the update equation , where and is the stochastic gradient of . Save for the noise term , this update is identical to the conventional SGD update when minimizing the maximum-a-posteriori (MAP) objective . Similarly, Stochastic Gradient HMC (SGHMC) corresponds to SGD with momentum injected with properly scaled noise. Given a limited computational budget, SG-MCMC methods can however struggle to explore the high-dimensional and highly multi-modal posteriors of large DNNs. To mitigate this problem, Zhang et al. proposed to use a cyclical stepsize schedule to help escaping local modes in .
The second paradigm is that of Variational Inference (VI) . Here, a distribution parameterized by variational parameters is explicitly chosen, and the best possible approximation is found by minimizing the Kullback-Leibler (KL) divergence with respect to the true posterior . While principled, VI methods generally require sophisticated implementations, especially for more expressive variational distributions . A particularly simple and scalable method is MC-dropout . It entails using dropout also at test time, which can be interpreted as performing VI with a Bernoulli variational distribution . The approximate predictive posterior in (4) is obtained by performing stochastic forward passes on the same input.
Ensembling Lakshminarayanan et al. created a parametric model of the conditional distribution using a DNN , and learned multiple point estimates by repeatedly minimizing the MLE objective with random initialization. They then averaged over the corresponding parametric models to obtain the following predictive distribution,
The authors considered this a non-Bayesian alternative to predictive uncertainty estimation. However, since the point estimates always can be seen as samples from some distribution , we note that (6) is virtually identical to the approximate predictive posterior in (4). Ensembling can thus naturally be viewed as approximate Bayesian inference, where the level of approximation is determined by how well the implicit sampling distribution approximates the posterior . Ideally, should be distributed exactly according to . Since is highly multi-modal in the parameter space for DNNs , so is . By minimizing multiple times, starting from randomly chosen initial points, we are likely to find different local optima. Ensembling can thus generate a compact set of samples that, even for small values of , captures this important aspect of multi-modality in .
Experiments
We conduct experiments both on illustrative toy regression and classification problems (Section 4.1), and on the real-world computer vision tasks of depth completion (Section 4.2) and street-scene semantic segmentation (Section 4.3). Our evaluation is motivated by real-world conditions found e.g. in automotive applications, where robustness to varying environments and weather conditions is required to ensure safety. Since images captured in these different circumstances could all represent distinctly different regions of the vast input image space, it is infeasible to ensure that all encountered inputs will be well-represented by the training data. Thus, we argue that robustness to out-of-domain inputs is crucial in such applications. To simulate these challenging conditions and test the robustness required for such real-world scenarios, we train all models on synthetic data and evaluate them on real-world data. To improve rigour of our evaluation, we repeat each experiment multiple times and report results together with the observed variation. A more detailed description of all results are found in the Appendix (Appendix B.3, C.2, D.2). All experiments are implemented in PyTorch .
2 Depth Completion
Results A comparison of ensembling and MC-dropout in terms of AUSE, AUCE and RMSE on the KITTI depth completion validation dataset is found in Figure 6. We observe in Figure 6(a) that ensembling consistently outperforms MC-dropout in terms of AUSE. However, the curves decrease as a function of in a similar manner. Sparsification plots and sparsification error curves are found in Appendix C.3. A ranking of the methods can be more readily conducted based on Figure 6(b), where we observe a clearly improving trend as increases for ensembling, whereas MC-dropout gets progressively worse. This result is qualitatively supported by the calibration plots found in Appendix C.3 and Figure 7. Note that corresponds to the baseline of only estimating aleatoric uncertainty.
3 Street-Scene Semantic Segmentation
Evaluation Metrics As for depth completion, we evaluate the methods in terms of the AUSE metric. In this classification setting, we compare the “oracle” ordering of predictions with the one induced by the predictive entropy. We compute AUSE in terms of Brier score and based on all pixels in the evaluation dataset. We also evaluate in terms of calibration by the ECE metric . All predictions are here partitioned into bins based on the maximum assigned confidence. For each bin, the difference between the average predicted confidence and the actual accuracy is then computed, and ECE is obtained as the weighted average of these differences. We use bins of equal size.
Results A comparison of ensembling and MC-dropout in terms of AUSE, ECE and mIoU on the Cityscapes validation dataset is found in Figure 8. We observe that the metrics clearly improve as functions of for both ensembling and MC-dropout, demonstrating the importance of epistemic uncertainty estimation. The rate of improvement is generally greater for ensembling. For ECE, we observe in Figure 8(b) a drastic improvement for ensembling as is increased, followed by a distinct plateau. According to the condensed reliability diagrams in Appendix D.3, this corresponds to a transition from clear model over-confidence to slight over-conservatism. For MC-dropout, the corresponding diagrams suggest a stagnation while the model still is somewhat over-confident. Example reliability diagrams for are shown in Figure 9, in which this over-confidence for MC-dropout can be observed. Note that the relatively low mIoU scores reported in Figure 8(c), obtained by models trained exclusively on Synscapes, are expected and caused by the intentionally challenging domain gap between synthetic and real-world data.
Discussion & Conclusion
We proposed a comprehensive evaluation framework for scalable epistemic uncertainty estimation methods in deep learning. The proposed framework is specifically designed to test the robustness required in real-world computer vision applications. We applied our proposed framework and provided the first properly extensive and conclusive comparison of ensembling and MC-dropout, the results of which demonstrates that ensembling consistently provides more reliable and practically useful uncertainty estimates. We attribute the success of ensembling to its ability, due to the random initialization, to capture the important aspect of multi-modality present in the posterior distribution of DNNs. MC-dropout has a large design-space compared to ensembling, and while careful tuning of MC-dropout potentially could close the performance gap on individual tasks, the simplicity and general applicability of ensembling must be considered key strengths. The main drawback of both methods is the computational cost at test time that grows linearly with , limiting real-time applicability. Here, future work includes exploring the effect of model pruning techniques on predictive uncertainty quality. For ensembling, sharing early stages of the DNN among ensemble members is also an interesting future direction. A weakness of ensembling is the additional training required, which also scales linearly with . The training of different ensemble members can however be performed in parallel, making it less of an issue in practice given appropriate computing infrastructure. In conclusion, our work suggests that ensembling should be considered the new go-to method for scalable epistemic uncertainty estimation.
Acknowledgments This research was supported by the Swedish Foundation for Strategic Research via the project ASSEMBLE and by the Swedish Research Council via the project Learning flexible models for nonlinear dynamics.
References
Supplementary Material
This is the supplementary material for Evaluating Scalable Bayesian Deep Learning Methods for Robust Computer Vision. It consists of Appendix A-D.
Appendix A Approximating a Mixture of Gaussian Distributions
For the Gaussian model (2), in (4) is a uniformly weighted mixture of Gaussian distributions. We approximate this mixture with a single Gaussian parameterized by the mixture mean and variance:
Appendix B Illustrative Toy Problems
In this appendix, further details on the illustrative toy problems experiments (Section 4.1) are provided.
For both regression and classification, HMC with prior and samples is implemented using Pyro . Specifically, we use pyro.infer.mcmc.MCMC with pyro.infer.mcmc.NUTS as kernel, and .
B.2 Implementation Details
For regression, we use the Gaussian model (2) with two separate feed-forward neural networks outputting and . Both neural networks have hidden layers of size .
For classification, we use the Categorical model (1) with a feed-forward neural network with hidden layers of size .
For the MC-dropout comparison, we place a dropout layer after the first hidden layer of each neural network. For regression, we use a drop probability . For classification, we use .
For ensembling, we train all ensemble models for epochs with the Adam optimizer, a batch size of and a fixed learning rate of .
For MC-dropout, we train models for epochs with the Adam optimizer, a batch size of and a fixed learning rate of .
For classification, where (one-hot encoded) and is the Softmax output, it corresponds to the following loss:
For SGLD, we extract samples from the parameter trajectory given by the update equation:
where , is the stochastic gradient of and is the stepsize. We run it for a total number of steps corresponding to epochs with a batch size of . The stepsize is decayed according to:
where is the total number of steps, (the initial stepsize) for regression and for classification. samples are extracted starting at step , ending at step and spread out evenly between.
For SGHMC, we extract samples from the parameter trajectory given by the update equation:
where , is the stochastic gradient of , is the stepsize and . We run it for a total number of steps corresponding to epochs with a batch size of . The stepsize is decayed according to:
where is the total number of steps, (the initial stepsize) for regression and for classification. samples are extracted starting at step , ending at step and spread out evenly between.
For all models, we randomly initialize the parameters using the default initializer in PyTorch.
B.3 Description of Results
The results in Figure 5(a), 5(b) were obtained in the following way:
Ensembling: models were trained using the same training procedure, the mean and standard deviation was computed based on unique sets of models for .
MC-dropout: models were trained using the same training procedure, based on which the mean and standard deviation was computed.
SGLD: models were trained using the same training procedure, based on which the mean and standard deviation was computed.
SGHMC: models were trained using the same training procedure, based on which the mean and standard deviation was computed.
B.4 Additional Results
Figure 10 and Figure 11 show the same comparison as Figure 5(a), 5(b), but using SGD and SGD with momentum for ensembling and MC-dropout, respectively. We observe that ensembling consistently outperforms the compared methods for classification, but that SGLD and SGHMC has better performance for regression in these cases. SGLD and SGHMC are however trained for 256 times longer than each ensemble model, complicating the comparison somewhat. If SGLD and SGHMC instead are trained for just 64 times longer than each ensemble model, we observe in Figure 12 that they are consistently outperformed by ensembling.
For MC-dropout using Adam, we also varied the drop probability and chose the best performing variant. These results are found in Figure 13, in which * marks the chosen variant.
B.5 Qualitative Results
Here, we show visualizations of predictive distributions obtained by the different methods. Figure 14, 18 for ensembling, Figure 15, 19 for MC-dropout, Figure 16, 20 for SGLD, and Figure 17, 21 for SGHMC.
Appendix C Depth Completion
In this appendix, further details on the depth completion experiments (Section 4.2) are provided.
For both ensembling and MC-dropout, we train all models for steps with the Adam optimizer, a batch size of , a fixed learning rate of and weight decay of . We use a smaller batch size and train for fewer steps than Ma et al. to enable an extensive evaluation with repeated experiments. For the same reason, we also train on randomly selected image crops of size . The only other data augmentation used is random flipping along the vertical axis. We follow Ma et al. and randomly initialize all network weights from and all network biases with s. Models are trained on a single NVIDIA TITAN Xp GPU with GB of RAM.
C.2 Description of Results
The results in Figure 6 (Section 4.2) were obtained in the following way:
Ensembling: models were trained using the same training procedure, the mean and standard deviation was computed based on (), () or () sets of randomly drawn models. The same set could not be drawn more than once.
MC-dropout: models were trained using the same training procedure, based on which the mean and standard deviation was computed.
C.3 Additional Results
Here, we show sparsification plots, sparsification error curves and calibration plots. Examples of sparsification plots are found in Figure 22 for ensembling and Figure 23 for MC-dropout. Condensed sparsification error curves are found in Figure 24 for ensembling and Figure 25 for MC-dropout. Condensed calibration plots are found in Figure 26 for ensembling and Figure 27 for MC-dropout.
Appendix D Street-Scene Semantic Segmentation
In this appendix, further details on the street-scene semantic segmentation experiments (Section 4.3) are provided.
For ensembling, we train all ensemble models for steps with SGD + momentum (), a batch size of and weight decay of . The learning rate is decayed according to:
where and (the initial learning rate). We train on randomly selected image crops of size . We choose a smaller crop size than Yuan and Wang to enable an extensive evaluation with repeated experiments. The only other data augmentation used is random flipping along the vertical axis and random scaling in the range . The ResNet101 backbone is initialized with weightshttp://sceneparsing.csail.mit.edu/model/pretrained_resnet/resnet101-imagenet.pth. from a model pretrained on the ImageNet dataset, all other model parameters are randomly initialized using the default initializer in PyTorch. Models are trained on two NVIDIA TITAN Xp GPUs with GB of RAM each. For MC-dropout, models are instead trained for steps.
D.2 Description of Results
The results in Figure 8 (Section 4.3) were obtained in the following way:
Ensembling: models were trained using the same training procedure, the mean and standard deviation was computed based on sets of randomly drawn models for . The same set could not be drawn more than once.
MC-dropout: models were trained using the same training procedure, based on which the mean and standard deviation was computed.
D.3 Additional Results
Here, we show sparsification plots, sparsification error curves and reliability diagrams. Examples of sparsification plots are found in Figure 28 for ensembling and Figure 29 for MC-dropout. Condensed sparsification error curves are found in Figure 30 for ensembling and Figure 31 for MC-dropout. Examples of reliability diagrams with histograms are found in Figure 32 for ensembling and Figure 33 for MC-dropout. Condensed reliability diagrams are found in Figure 34 for ensembling and Figure 35 for MC-dropout.