Evaluating Bayesian Deep Learning Methods for Semantic Segmentation
Jishnu Mukhoti, Yarin Gal
Introduction
Deep learning has had tremendous success in quite a few fields including computer vision , natural language processing , speech recognition , bioinformatics and others. However, most deep learning models produce point-estimates as outputs and hence we do not gain any knowledge about the confidence of the model in its predictions. With the increasing use of AI systems in real-life scenarios like autonomous driving and medical diagnosis , there are many cases in which additional knowledge about the model’s confidence, i.e. capturing whether the model is essentially ‘guessing at random’, becomes not only useful but essential .
The development of new computationally light-weight, scalable methods of performing approximate Bayesian inference in deep neural networks has enabled these models to estimate their uncertainty in addition to making predictions . However, as we do not have ground truth uncertainties, we cannot use the conventional methods of evaluation to compare and benchmark these models. Furthermore, the metrics designed for evaluating model performance are often task dependent. For instance, the intersection-over-union (IOU) metric is heavily used in computer vision problems like object detection and semantic segmentation. Extending on these ideas, in this work, we propose new specialised metrics to evaluate Bayesian models designed for the task of semantic segmentation.
Semantic segmentation is a difficult problem in computer vision which requires pixel-level understanding of an image. A Bayesian model for semantic segmentation will not only produce predictions for each pixel but also generate pixel-wise uncertainty estimates. As mentioned above, evaluating such a Bayesian model is a challenging task because unlike predictions, we do not have a strong definition of what a good uncertainty estimate is. Hence, we have to judge the quality of the model uncertainty based on how accurate the model is for the same input. We require metrics which look at both the model predictions and uncertainties and take into account general desiderata about when a model should be uncertain about its predictions. In particular, we use the following two intuitive desiderata:
Desideratum 1: if a model is confident about its prediction, it should be accurate on the same.
Desideratum 2: if a model is not confident about its prediction, it may or may not be accurate.
This hints at an inverse relation between model accuracy and uncertainty, a property which we exploit when designing metrics for performance evaluation.
There has been a lot of research on deep architectures for semantic segmentation . In this work, we choose DeepLab-v3+ , one of the state-of-the-art neural networks for this purpose and create two probabilistic versions of it using dropout based approximate inference techniques: MC dropout and Concrete dropout . We train these models on the well-known Cityscapes dataset which contains many images of urban street scenes. Finally, we evaluate and compare the trained models using the metrics which we propose in this work. In Figure 1, we present a high level overview of the evaluation system which we implement.
In a nutshell, the main contributions of this paper are:
We propose three novel metrics which can be used to evaluate Bayesian models for semantic segmentation.
We create two probabilistic versions of the DeepLab-v3+ network, which can produce pixel-wise uncertainty estimates in addition to semantic segmentation results.
Finally, we evaluate the MC dropout and Concrete dropout inference techniques using the metrics mentioned above, thereby laying down benchmarks against which other models can be compared.
Related Work
In this section, we discuss some of the recent works on semantic segmentation as well as those on approximate inference in Bayesian Deep Learning.
The work by Long et al. was the first of its kind where a convolutional neural network without any fully connected layers was trained in an end-to-end manner directly mapping images to their corresponding segmentation results. This enabled the network to segment images of varying sizes. However, the fully convolutional networks still suffered due to the presence of pooling layers which ignore the positional information of objects in an attempt to reduce the dimensions of feature maps.
In order to get around this issue, researchers have followed two primary threads of thought: the encoder-decoder architecture and the dilated/atrous convolutions . The encoder-decoder networks first reduce the spatial dimensions of the feature maps with repeated applications of convolution and pooling layers in the encoder module. Next, in the decoder module, the spatial dimensions are gradually recovered using de-convolution and upsampling layers. In order to have sharper segmentation results, skip connections are often introduced between the encoder and decoder modules. Few popular works in this category include U-Net , SegNet and RefineNet .
The second class of architectures use dilated or atrous convolutions to have a larger field of view over the input feature maps without a decrease in spatial dimensions. One of the most popular set of deep neural networks which follow this policy is DeepLab . In this work, we adopt a combination of both policies and use DeepLab-v3+ as the base network for semantic segmentation. DeepLab-v3+ uses atrous convolution layers as well as a simple decoder module to have fine-grained segmentation.
2 Approximate Inference in Deep Neural Nets
The idea behind Bayesian modelling is to find the probability of each set of model parameters given a dataset. In order to do so, an initial distribution known as the prior is assumed over the model parameters. Next, with the input of data, this distribution is updated to capture parameters which are more likely to have generated the dataset. This update is done by applying Bayes’ theorem. Once the entire dataset is processed, the distribution over the model parameters, known as the posterior, captures the updated belief about the optimal set of parameters to represent the data.
However, getting the posterior for large neural networks is computationally intractable. Hence, many methods to approximate the posterior have been proposed. One such class of methods include Markov Chain Monte Carlo (MCMC) techniques . Another popular set of techniques uses variational inference where the posterior is approximated using a variational distribution. In this work, we have used MC dropout and Concrete dropout as methods of approximate inference. Both these methods are based on the variational inference approach and are standard well-performing techniques which are easy to implement in deep neural networks.
3 Existing metrics for evaluating uncertainty
There is a thread of work which focuses on generating calibrated probability estimates from a deep neural network as a measure of model confidence. There also exist popular metrics like the expected calibration error (ECE) and the maximum calibration error (MCE) which can be used to quantitatively measure model calibration. However, these metrics are based on softmax probabilities which cannot capture epistemic or model uncertainty . Furthermore, there are very simple post-processing techniques like temperature scaling which can make a deterministic and a probabilistic model equally calibrated, thereby rendering the ECE and MCE metrics unable to detect whether a model is "guessing at random", a property which our metrics are designed to capture. We empirically validate these observations in Section 5.
Finally, we conclude this section with a discussion of Bayesian SegNet , which is one of several works to have applied approximate Bayesian inference in the context of semantic segmentation. The authors of modify the SegNet architecture using MC dropout to obtain uncertainties in addition to segmentation results. Furthermore, they present a few accuracy-vs-uncertainty plots in their work, which are good sanity checks for a Bayesian neural network. These sanity checks are qualitative though, and do not allow us to compare and choose between BDL models for semantic segmentation. It is precisely this gap that we fill by developing quantitative measures as well.
Bayesian DeepLab
In this section we describe the variant of the DeepLab-v3+ network architecture which we have implemented.
DeepLab-v3+ is one of the state-of-the-art deep neural networks designed for the problem of semantic segmentation. There have been multiple versions of DeepLab, namely DeepLab-v1 and v2 , DeepLab-v3 and DeepLab-v3+ , over the years. However, all these versions possess certain common architectural traits. Firstly, they propose atrous or dilated convolutions as a way to widen the field of view over the input feature maps without increasing the number of parameters or using pooling layers. Secondly, they deal with the problem of objects present at different scales in the image using methods like image pyramid , atrous spatial pyramid pooling (ASPP) , cascaded atrous modules and encoder-decoder architectures . Thirdly, even though there are no pooling layers in the network, due to the presence of multiple convolution layers with strides of 1 or more, the resulting output dimensions are reduced. In order to regain the original dimensions, the output is passed through a fully connected CRF or resized using bilinear interpolation or passed through a decoder with learnable parameters .
Finally, the above-mentioned architectural features can be applied to any base network as long as it is fully convolutional. Some of the popular CNNs which have been used for DeepLab are VGG-16 , ResNet-101 and Xception . In this work, we implement DeepLab-v3+ using Xception as the base network. The Xception architecture enjoys the simplicity of VGG with multiple convolution layers stacked on top of one another. Furthermore, Xception modules use skip connections similar to ResNet and are also based on the Inception hypothesis which postulates the separation of convolution operations performed on spatial (height and width) dimensions and those on the depthwise (or cross-channel) dimensions.
2 Approximate inference in DeepLab-v3+
In Section 2.2 we listed some techniques for approximate inference in deep neural networks. The purpose of this work is to develop and present metrics with which these inference techniques can be evaluated and benchmarked for the task of semantic segmentation. To do this, we have to first create a probabilistic deep neural network for semantic segmentation. As mentioned before, we use DeepLab-v3+ as the network of choice. Furthermore, we use MC dropout as the primary method of approximate inference. It is one of the current standard baselines and has already been applied to multiple problems in computer vision .
between the actual posterior and the variational distribution is minimized. In MC dropout, the variational distribution defined over the network weights is a Bernoulli distribution. In , the authors observed that placing a Bernoulli distribution with parameter over the weights of a hidden layer is equivalent to performing dropout on that layer with a dropout rate of . Furthermore, they also noted that minimising the well-known cross-entropy loss function using standard optimisation algorithms like stochastic gradient descent has the desired effect of minimizing the KL divergence term in equation 1. Hence, in order to perform approximate inference, one first needs to train a network with dropout. However, unlike common practice, these dropout layers are kept active even during the test phase. The idea is to get samples from the posterior distribution and as the dropout layers place a Bernoulli distribution over the network weights, performing a stochastic forward pass through a trained network can be interpreted as generating a Monte Carlo sample from the posterior distribution. Therefore, multiple forward passes on the same input generate multiple such Monte Carlo samples, the mean of which can then be used as the network prediction and the variance can be interpreted as an uncertainty estimate.
Ideally, in a Bayesian neural network, a dropout layer should be inserted after every hidden layer of the network. However, as was observed by Kendall et al. in for the SegNet architecture and as we observe for DeepLab-v3+, insertion of dropout layers after every convolution layer in a large neural network regularises it to an extent that makes training prohibitively slow. Therefore, the dropout layers have to be inserted only in certain regions of the network. This gives rise to multiple probabilistic variants of the Bayesian DeepLab architecture depending on where in the network, the dropout layers are inserted. However, for the sake of simplicity, in this work, we insert dropout layers only in the middle flow of the DeepLab-v3+ network. We do this because of the hypothesis proposed in which states that low level features in shallower layers of the network are mostly consistent across the distribution of models and hence can be represented using deterministic weights whereas the higher level features in deeper layers are better modelled using probabilistic weights.
With the incorporation of MC dropout in DeepLab-v3+, we end up with Bayesian DeepLab, a probabilistic deep neural network designed specifically for semantic segmentation. However, in this work, our goal is to develop evaluation metrics for such networks and lay out a few benchmarks based on these metrics. Thus, in order to compare performance with MC dropout, we implement an alternative inference technique, namely Concrete dropout . We choose Concrete dropout in particular because both the approximate inference techniques are dropout based methods. They are also simple to implement and require minimal changes to the network architecture. Concrete Dropout is a modification on the MC dropout method where the network tunes the dropout rates during training. Similar to the MC dropout variant, we place the concrete dropout layers in the middle flow of the DeepLab-v3+ network.
3 Network Architecture
The backbone framework of Bayesian DeepLab is similar to DeepLab-v3+ where the inputs are first passed through an extended Xception network followed by an ASPP module for multi-scale image processing and finally a decoder module to resize the images to the original input dimensions and to produce sharp segmentation results. The differences between the network architecture which we use in this work and the original DeepLab-v3+ network proposed in are as follows:
Bayesian DeepLab has dropout layers at different points in the network. In particular we insert a dropout layer after every 4 Xception modules in the middle flow of the network. As there are 16 Xception modules in the middle flow, there are a total of 4 dropout layers in the network. We use a dropout rate of 0.5 in each of the dropout layers. However, these rates are hyperparameters which can be fine-tuned further.
We do not use cascaded atrous modules or image pyramids. The only method for multi-scale image processing adopted by Bayesian DeepLab is Atrous Spatial Pyramid Pooling (ASPP). We do this primarily for the sake of simplicity and to reduce training time.
We present the Bayesian DeepLab architecture in Figure 7 in the appendix.
4 Uncertainty Metrics
There are two types of uncertainties which we study in this work. Epistemic uncertainty, also known as model uncertainty represents what the model does not know due to insufficient training data. This kind of uncertainty can be explained away with more training data. Aleatoric uncertainty is caused due to noisy measurements in the data and can be explained away with increased sensor precision (but cannot be explained away with increase in training data). The two uncertainties combined form the predictive uncertainty of the network. In , the author suggests some information theoretic metrics which can be used as measures of uncertainty in classification problems. In our work, we use two such metrics, namely the entropy of the predictive distribution (also known as predictive entropy) and the mutual information between the predictive distribution and the posterior over network weights.
We choose the predictive entropy and mutual information metrics as they capture different kinds of uncertainty. As observed in , mutual information captures epistemic or model uncertainty whereas predictive entropy captures predictive uncertainty which combines both epistemic and aleatoric uncertainties. It is worth noting here that in semantic segmentation, we produce pixel-wise classification results and hence, we also produce pixel-wise uncertainty estimates. Thus, we get uncertainty maps which have the same dimensions as that of the input image.
Performance Evaluation Metrics
As mentioned in Section 1, we measure the performance of Bayesian models using metrics which capture properties which we want the model to satisfy. In particular, we assume that if a model is confident about its prediction, it should be accurate on the same. This also implies that if a model is inaccurate on an output, it should be uncertain about the same output. It is worth noting that the converse of these assumptions may not hold. For instance, a model may have a high epistemic uncertainty on a class which appears infrequently in the training set but can still be accurate on its prediction. Given the above assumptions, we can define the following two conditional probabilities:
In order to implement the above metrics, we first choose a patch/window size and traverse the predicted labels, actual labels and uncertainty maps using windows of dimensions , much like the traversal in a convolution operation. Technically, the patch dimensions can be as small as a single pixel (i.e., a patch) or as large as the entire image. However, labelling the entire image as accurate or uncertain is not very useful. Furthermore, the uncertainties occur in regions within the image comprising multiple pixels in close neighbourhoods. Capturing these uncertain regions is useful for downstream tasks and having single pixel patches does not help in this regard. Thus, we have and we compute the accuracy of each patch from the predicted and actual labels. This accuracy can be computed using any standard technique. In this work we use the pixel accuracy metric defined in . If the patch accuracy is above a certain threshold, we mark the patch as accurate.
Similarly, from the corresponding patch obtained from the uncertainty map, we compute the average patch uncertainty. If this uncertainty value is above a given threshold, we label the patch as uncertain. There can be multiple ways of setting the uncertainty threshold. One simple way would be to find the average uncertainty of all pixels over a validation set and use that value as the threshold. Other ways could include computing the minimum and maximum uncertainty values over validation set pixels and setting the uncertainty threshold as:
Finally, we combine both the good cases of (accurate, certain) and (inaccurate, uncertain) patches into a single metric, the Patch Accuracy vs Patch Uncertainty (PAvPU), defined as follows:
Clearly, a model with a higher value of the above metrics is a better performer. As the values of the above metrics depend on three parameters: the accuracy threshold, the uncertainty threshold, and the patch dimensions, an interesting experiment would be to observe how the metrics vary with these parameters. We show this in Figures 3 and 4. In Figure 2, we provide an illustrative example of computing the above three metrics.
Experiments and Results
In this section we first evaluate the trained models on their segmentation performance using the pixel accuracy, mean accuracy and mean IOU metrics defined in . Next, we compare and benchmark MC dropout and Concrete dropout inference using the metrics proposed in Section 4. We also compare the above two Bayesian models with a deterministic DeepLab-v3+ baseline where we use the entropy of the softmax distribution as a measure of uncertainty. All experiments have been performed on Cityscapes dataset. We provide further details about the training infrastructure in Appendix B.
In Table 1 we report the semantic segmentation results obtained from the Bayesian DeepLab variants using both MC dropout and Concrete dropout inference and compare these results with different versions of DeepLab on the Cityscapes val set. We observe that Concrete dropout performs better than MC dropout with respect to all the three metrics: pixel accuracy, mean accuracy and mean IOU. Furthermore, both our models with mean IOU values of 78.05 and 79.12 outperform all the DeepLab versions except DeepLab-v3+ which has a mean IOU of 79.14. It is worth noting that we compare our models only with those versions of DeepLab which use the same hyperparameters as us. To be precise, we compare with the version of DeepLab-v3 which uses an output stride of 16 and DeepLab-v3+ which uses Xception-65 with ASPP and decoder modules but without image level features.
In Figure 5, we present some qualitative results on Cityscapes val set images. It is interesting to note the difference between the uncertainty maps provided by the predictive entropy and mutual information metrics. In case of mutual information, we observe high uncertainty inside the boundaries of objects which the model is confused about. The bonnet of the Mercedes at the bottom of each image is one such example. However, in predictive entropy maps, we also see high uncertainties on the edges of objects like pedestrians or cars. These edges are regions where the presence of noise in the dataset is highly likely. This observation supports the explanation in that mutual information captures epistemic or model uncertainty and predictive entropy captures aleatoric uncertainty.
2 Evaluation of Bayesian DeepLab
As seen in Table 2, the dropout based models outperform the deterministic model and Concrete dropout consistently outperforms MC dropout on all the metrics. It is worth noting that we cannot compute the metrics for mutual information for a deterministic model as the mutual information is always 0, thereby indicating that the deterministic model cannot capture epistemic uncertainty. In Figures 3 and 4, we plot these metrics for varying thresholds of uncertainty. We can draw the following observations from the plots:
It is interesting to note that the sum of the PAvPU values at uncertainty thresholds 0 and 100% is 1. Furthermore, the PAvPU value at 100% uncertainty threshold is significantly greater than the value at 0. This indicates that the number of accurate patches is much higher than the number of inaccurate patches.
We observe that the Bayesian models perform significantly better than the deterministic baseline. One reason behind this is that the softmax entropy in a deterministic model is only able to capture aleatoric uncertainty and regions of high epistemic uncertainty are ignored, resulting in a relatively high number of inaccurate but certain patches. This does not occur in the Bayesian models as predictive entropy captures both aleatoric and epistemic uncertainties. Finally, the plots indicate the superior performance of Concrete dropout over MC dropout which is consistent with the results in Table 2.
As mentioned in section 2.2, we compare the ECE and MCE values of the three models after calibrating them using temperature scaling . With 15 bins, we obtain an ECE of 0.0182, 0.0161 and 0.0165 at optimal temperatures 4.7, 3.1 and 3.2 respectively for the deterministic, MC dropout and Concrete dropout models. Similarly, we obtain MCE values of 0.1969, 0.1823 and 0.1785 at the same optimal temperatures. Thus, after temperature scaling, the ECE and MCE metrics are non-indicative of which model to choose and are not able to detect when a model is guessing at random (i.e., has high epistemic uncertainty). This, however is clearly captured in the proposed metrics as can be seen from Figures 3 and 4 and Table 2. Finally, it is worth noting that the proposed metrics evaluate uncertainty values and hence, should be used in addition to conventional metrics of measuring accuracy like mean IOU .
Conclusions
In this work we have developed metrics to evaluate Bayesian models for the task of semantic segmentation. We have created two probabilistic variants of the DeepLab-v3+ network and have evaluated them using these metrics, thereby providing benchmarks which can be used for future comparisons. For experiments, we have used the Cityscapes dataset which is particularly suited for applications like autonomous driving. An interesting future work would be to develop metrics which measure performance of Bayesian models based on how effective the segmentation outputs and uncertainty estimates are in making safe and correct autonomous driving decisions. This idea can be further extended to include other downstream applications as well where semantic segmentation is a useful intermediate tool.
References
Appendix A Bayesian DeepLab Network Architecture
In Figure 7, we present the network architecture of Bayesian DeepLab using MC dropout inference. We use Xception as the base network. However, there are certain differences between the original Xception architecture and the one which we use. The differences are as follows:
There are no pooling layers in our network. We use separable convolution filters with a stride of 2 instead of max pooling layers. This helps in dense predictions. Furthermore, following the FCN philosophy, our network is fully convolutional and hence can segment images of arbitrary sizes.
The middle flow in our network has 16 modules instead of 8 as described in the original Xception paper .
Appendix B Training Infrastructure
We train all our networks on the Cityscapes dataset, one of the most popular datasets for urban scene understanding. It has 5000 images collected from street scenes in 50 different cities. There are 2975 images in the training set, 500 images in the validation set and 1525 test images with dimensions . In order to train the networks, we set the following parameters:
a list of atrous rates for the ASPP module which we set to $$ following the DeepLab-v3+ paper,
the output stride and the decoder output stride which we set as 16 and 4 respectively,
the crop size for training images which we set to and
the training batch size which we set to 16.
Furthermore, the Concrete dropout layers require two additional hyperparameters: the weight regulariser which is set to and the dropout regulariser which we set to , where is the number of training images and and are the height and width of the image respectively, following the paper . We train all the networks for 90,000 iterations on 8 NVIDIA Tesla P100-SXM2 GPUs. The networks finish training in approximately 3 days.