Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processes
Sebastian W. Ober, Laurence Aitchison
Introduction
Deep models, formed by stacking together many simple layers, give rise to extremely powerful machine learning algorithms, from deep neural networks (DNNs) to deep Gaussian processes (DGPs) (Damianou & Lawrence, 2013). One approach to reason about uncertainty in these models is to use variational inference (VI) (Jordan et al., 1999). VI in Bayesian neural networks (BNNs) requires the user to specify a family of approximate posteriors over the weights, with the classical approach being independent Gaussian distributions over each individual weight (Hinton & Van Camp, 1993; Graves, 2011; Blundell et al., 2015). Later work has considered more complex approximate posteriors, for instance using a Matrix-Normal distribution as the approximate posterior for a full weight-matrix (Louizos & Welling, 2016; Ritter et al., 2018) and hierarchical variational inference (Louizos & Welling, 2017; Dusenberry et al., 2020). By contrast, DGPs use an approximate posterior defined over functions – the standard approach is to specify the inputs and outputs at a finite number of “inducing” points (Damianou & Lawrence, 2013; Salimbeni & Deisenroth, 2017).
Critically, these classical BNN and DGP approaches define approximate posteriors over functions that are independent across layers. An approximate posterior that factorises across layers is problematic, because what matters for a deep model is the overall input-output transformation for the full model, not the input-output transformation for individual layers. This raises the question of what family of approximate posteriors should be used to capture correlations across layers. One approach for BNNs would be to introduce a flexible “hypernetwork”, used to generate the weights (Krueger et al., 2017; Pawlowski et al., 2017). However, this approach is likely to be suboptimal as it does not sufficiently exploit the rich structure in the underlying neural network. For guidance, we consider the optimal approximate posterior over the top-layer units in a deep network for regression. Remarkably, the optimal approximate posterior for the last-layer weights given the earlier weights can be obtained in closed form without choosing a restrictive family of distributions. In particular, the optimal approximate posterior is given by propagating the training inputs through lower layers to compute the top-layer representation, then using Bayesian linear regression to map from the top-layer representation to the outputs.
Inspired by this result, we use Bayesian linear regression to define a generic family of approximate posteriors for BNNs. In particular, we introduce learned “pseudo-outputs” at every layer, and compute the posterior over the weights by performing linear regression from the inputs (propagated from lower layers) onto the pseudo-outputs. We reduce the burden of working with many training inputs by summarising the posterior using a small number of “inducing” points. We find that these approximate posteriors give excellent performance in the non-tempered, no-data-augmentation regime, with performance on datasets such as CIFAR-10 reaching , comparable to SGMCMC witout tempering but with data augmentation (88%) (Wenzel et al., 2020). Our approach can be extended to DGPs, and we explore connections to the inducing point GP literature, showing that inference in the two classes of models can be unified.
We propose an approximate posterior for BNNs based on Bayesian linear regression that naturally induces correlations between layers (Sec. 2.1).
We provide an efficient implementation of this posterior for convolutional layers (Sec. 2.2).
We introduce new BNN priors that allow for more flexibility with inferred hyperparameters (Sec. 2.3).
We show how our approximate posterior can be naturally extended to DGPs, resulting in a unified approach for inference in BNNs and DGPs (Sec. 2.4).
Methods
Rearranging these terms (App. A), we find that all dependence can be written in terms of the KL divergence between the approximate posterior of interest and the true posterior,
Thus, the optimal approximate posterior is the true last-layer posterior conditioned on the previous layers’ weights,
where the final proportionality comes by applying Bayes theorem and exploiting the model’s conditional independencies. For regression, the likelihood is Gaussian,
where is the value of a single output channel for all training inputs, and is a precision matrix. Thus, the posterior is given in closed form by Bayesian linear regression (Rasmussen & Williams, 2006):
While this result may be neither particularly surprising nor novel, it neatly highlights our motivation for the rest of the paper. In particular, it shows that for regression, we can always obtain the optimal conditional top-layer posterior for regression, which to the best of our knowledge has not been used before in BNN inference. Moreover, doing top-layer Bayesian linear regression based on the propagated features from the previous layers naturally introduces correlations between layers.
Substituting this approximate posterior and the factorised prior into the ELBO (Eq. 3), the full ELBO can be written,
2 Efficient convolutional Bayesian linear regression
3 Priors
We consider four priors in this work, which we refer to using the class names in the bayesfunc library published alongside this paper. We are careful to ensure that all parameters in the model have a prior and approximate posterior, which is necessary to ensure that ELBOs are comparable across models.
First, we consider a Gaussian prior with fixed scale, NealPrior, so named because it is necessary to obtain meaningful results when considering infinite networks (Neal, 1996),
though it bears strong similarities to the “He” initialisation (He et al., 2015). NealPrior is defined so as to ensure that the activations retain a sensible scale as they propagate through the network. We compare this to the standard (StandardPrior), which causes the activations to increase exponentially as they propagate through network layers (see Eq. 2):
Next, we consider ScalePrior, which defines a prior and approximate posterior over the scale,
where is the inverse-Wishart distribution, and the non-negative real number, , and the positive definite matrix, , are learned parameters of the approximate posterior (see Appendix D).
4 Extension to DGPs
It is a remarkable but underappreciated fact that BNNs are special cases of DGPs, with a particular choice of kernel (Louizos & Welling, 2016; Aitchison, 2019). Combining Eqs. (1) and (2),
To obtain more appropriate approximate posteriors, we consider the optimal top-layer posterior for DGPs, which involves GP regression from activations propagated from lower layers onto the output data (Appendix F). Inspired by the form of the optimal posterior we again define an approximate posterior by taking the product of the prior and a “inducing-likelihood”,
In summary, we propose an approximate posterior over inducing outputs that takes the form
We provide a full description of our method as applied to DGPs in App. E.
5 Asymptotic complexity
In the fully-connected BNN case, we have three terms, . The first term, , where corresponds to the width of the network, arises from taking the inverse of the covariance matrix in Eq. (11), but is also the complexity e.g. for propagating the inducing points from layer to the next (Eq. 10). The second term, , comes from computing that covariance in Eq. (11), by taking the product of input features with themselves. The third term comes from multiplying the training inputs/minibatch by the sampled inputs (Eq. 1).
Results
We describe our experiments and results to assess the performance of global inducing points (‘gi’) against local inducing points (‘li’) and the fully factorised (‘fac’) approximation family. We additionally consider models where we use one method up to the last layer and another for the last layer, which may have computational advantages; we denote such models ‘method1 method2’.
We demonstrate the use of local and global inducing point methods in a toy 1-D regression problem, comparing it with fully factorised VI and Hamiltonian Monte Carlo (HMC; (Neal et al., 2011)). Following Hernández-Lobato & Adams (2015), we generate 40 input-output pairs with the inputs sampled i.i.d. from and the outputs generated by , where . We then normalised the inputs and outputs. Note that we have introduced a ‘gap’ in the inputs, following recent work (Foong et al., 2019b; Yao et al., 2019; Foong et al., 2019a) that identifies the ability to express ‘in-between’ uncertainty as an important quality of approximate inference algorithms. We evaluated the inference algorithms using fully-connected BNNs with 2 hidden layers of 50 ReLU hidden units, using the NealPrior. For the inducing point methods, we used 100 inducing points per layer.
The predictive distributions for the toy experiment can be seen in Fig. 1. We observe that of the variational methods, the global inducing method produces predictive distributions closest to HMC, with good uncertainty in the gap. Meanwhile, factorised and local inducing fit the training data, but do not produce reasonable error bars, demonstrating an important limitation of methods lacking correlation structure between layers.
We provide additional experiments looking at the effect of the number of inducing points in Appendix I, and experiments looking at compositional uncertainty (Ustyuzhaninov et al., 2020) in both BNNs and DGPs for 1D regression in Appendix J.
2 Depth-dependence in deep linear networks
The lack of correlations between layers might be expected to become more problematic in deeper networks. To isolate the effect of depth on different approximate posteriors, we considered deep linear networks trained on data generated from a toy linear model: 5 input features were mapped to 1 output feature, where the 1000 training and 100 test inputs are drawn IID from a standard Gaussian, and the true outputs are drawn using a weight-vector drawn IID from a Gaussian with variance , and with noise variance of . We could evaluate the model evidence under the true data generating process which forms an upper bound (in expectation) on the model evidence and ELBO for all models.
We found that the ELBO for methods that factorise across layers — factorised and local inducing — drops rapidly as networks get deeper and wider (Fig. 2). This is undesirable behaviour, as we know that wide, deep networks are necessary for good performance on difficult machine learning tasks. In contrast, we found that methods with global inducing points at the last layer decay much more slowly with depth, and perform better as networks get wider. Remarkably, global-inducing points gave good performance even with lower-layer weights drawn at random from the prior, which is not possible for any method that factorises across layers. We believe that fac gi performed poorly at width due to optimisation issues as rand gi performs better yet is a special case of fac gi.
3 Regression benchmark: UCI
We benchmark our methods on the UCI datasets in Hernández-Lobato & Adams (2015), popular benchmark regression datasets for BNNs and DGPs. Following the standard approach (Gal & Ghahramani, 2015), each dataset uses 20 train-test ‘splits’ (except for protein with 5 splits) and the inputs and outputs are normalised to have zero mean and unit standard deviation. We focus on the five smallest datasets, as we expect Bayesian methods to be most relevant in small-data settings (see App. K andM for all datasets). We consider two-layer fully-connected ReLU networks, using fully factorised and global inducing approximating families, as well as two- and five-layer DGPs with doubly-stochastic variational inference (DSVI) (Salimbeni & Deisenroth, 2017) and global inducing. For the BNNs, we consider the standard prior and ScalePrior.
We display ELBOs and average test log likelihoods for the un-normalised data in Fig. 3, where the dots and error bars represent the means and standard errors over the test splits, respectively. We observe that global inducing obtains better ELBOs than factorised and DSVI in almost every case, indicating that it does indeed approximate the true posterior better (since the ELBO is the marginal likelihood minus the KL to the posterior). While this is the case for the ELBOs, this does not always translate to a better test log likelihood due to model misspecification, as we see that occasionally DSVI outperforms global inducing by a very small margin. The very poor results for factorised on ScalePrior indicate that it has difficulty learning useful prior hyperparameters for prediction (see also Blundell et al. (2015)), which is due to the looseness of its bound to the marginal likelihood. We provide experimental details, as well as additional results with additional architectures, priors, datasets, and RMSEs, in Appendices K and M, for BNNs and DGPs, respectively.
4 Convolutional benchmark: CIFAR-10
For CIFAR-10, we considered a ResNet-inspired model consisting of conv2d-relu-block-avgpool2-block-avgpool2-block-avgpool-linear, where the ResNet blocks consisted of a shortcut connection in parallel with conv2d-relu-conv2d-relu, using 32 channels in all layers. In all our experiments, we used no data augmentation and 500 inducing points. Our training scheme (see App. O) ensured that our results did not reflect a ‘cold posterior’ (Wenzel et al., 2020). Our results are shown in Table 1. We achieved remarkable performance of predictive accuracy, with global inducing points used for all layers, and with a spatial inverse Wishart prior on the weights. These results compare very favourably with comparable Bayesian approaches, i.e. those without data augmentation or posterior sharpening: past work with deep GPs obtained (Shi et al., 2019), and work using infinite-width neural networks to define a GP obtained accuracy (Li et al., 2019). Remarkably, with only 500 inducing points we are approaching the accuracy of sampling-based methods (Wenzel et al., 2020), which are in principle able to more closely approximate the true posterior. Furthermore, we see that global inducing performs the best in terms of ELBO (per datapoint) by a wide margin, demonstrating that it gets far closer to the true posterior than the other methods. We provide additional results on uncertainty calibration and out-of-distribution detection in Appendix L.
Finally, while we have focused this work on achieving good results with a fully principled, Bayesian approach, we briefly consider training a full ResNet-18 (He et al., 2016) using more popular techniques such as data augmentation and tempering. While data augmentation and tempering are typically viewed as clouding the Bayesian perspective, there is work attempting to formalise both within the context of modified probabilistic generative models (Aitchison, 2020; Nabarro et al., 2021). Using a cold posterior with a factor of 20 reduction on the KL term, with horizontal flipping and random cropping, we obtained a test accuracy of and a test log likelihood of nats with a standard fully factorised Gaussian posterior using SpatialIWPrior. Using fac gi, we obtain a significant improvement of test accuracy and a test log likelihood of nats. We attempted to train a full global inducing model; however, we encountered difficulties in scaling the method to the large widths of ResNet-18 layers. We believe this presents an avenue for fruitful future work in scaling global inducing point posteriors.
Related work
Note that some work on BNNs reports better performance on datasets such as CIFAR-10. However, to the best of our knowledge, no variational Bayesian method outperforms ours without modifying the BNN model or some form of posterior tempering (Wenzel et al., 2020), where the KL term in the ELBO is down-weighted relative to the likelihood (Zhang et al., 2017; Bae et al., 2018; Osawa et al., 2019; Ashukha et al., 2020), which often increases the test accuracy. However, tempering clouds the Bayesian perspective, as the KL to the posterior is no longer minimised and the resulting objective is no longer a lower bound on the marginal likelihood. By contrast, we use the untempered ELBO, thereby retaining the full Bayesian perspective. Dusenberry et al. (2020) report better performance on CIFAR-10 without tempering, but only perform variational inference over a rank-1 perturbation to the weights, and maximise over all the other parameters, which may risk overfitting (Ober et al., 2021). Our approach retains a full-rank parameterisation of the weight matrices.
Karaletsos & Bui (2020) propose a very different hybrid GP-BNN which uses an RBF-GP prior over weights, whereas we use standard DGP models (which do not have any weights) and BNN models with IID weight priors. Our BNN and DGP map from global-inducing inputs defined only at the input layer onto activations, while their GPs map from inducing inputs in a “neuron-embedding” space onto weights. Unlike Karaletsos & Bui (2020), our approach scales directly to large ResNets, and allows us to capture the optimal top-layer posterior. Note they use “global” vs. “local” for different priors, while we use it for different approximate posteriors. While their models allow for the possibility of correlations between layers, these arise from modifying the prior structure, which are then reflected in the approximate posterior. In contrast, we use standard Gaussian priors over weights that are uncorrelated across layers, and our across-layer dependencies arise only in the approximate posterior.
Alternative approaches to introducing more flexibility and dependence in approximate posteriors are to use hierarchical variational inference (HVI), i.e. to introduce auxiliary random variables which can be integrated out to give a richer variational posterior (Ranganath et al., 2016; Louizos & Welling, 2017; Sobolev & Vetrov, 2019). This has been used in a number of prior works to introduce correlations between parameters (Dusenberry et al., 2020; Louizos & Welling, 2017; Ghosh & Doshi-Velez, 2017). As we do not introduce auxiliary variables, our approach is not HVI. However, our approximate posterior does include dependencies across latent variables and is thus an instance of structured stochastic variational inference (Hoffman & Blei, 2015).
The general idea of marginalising over the Gaussian process latent function in order to compute the marginal likelihood is common in the Gaussian process literature (e.g. Heinonen et al., 2016), and is related to, but very different from our approach of using the top-layer optimal posterior to inspire an approximate posterior for the weights of a Bayesian neural network, or the Gaussian process function.
Conclusions
We derived optimal top-layer variational approximate posteriors for BNNs and deep GPs, and used them to develop generic, scalable approximate posteriors. These posteriors make use of global inducing points, which are learned only at the bottom layer and are propagated through the network. This leads to extremely flexible posteriors, which even allow the lower-layer weights to be drawn from the prior. We showed that these global inducing variational posteriors lead to improved performance with better ELBOs, and state-of-the-art performance for variational BNNs on CIFAR-10.
We would like to thank the anonymous reviewers for their valuable feedback which greatly improved the manuscript. We thank David R. Burt, Andrew Y. K. Foong, Siddharth Swaroop, Adrià Garriga-Alonso, Vidhi Lalchand, and Carl E. Rasmussen for valuable discussions and feedback. SWO acknowledges the support of the Gates Cambridge Trust in funding his doctoral studies. LA would like to thank the University of Bristol’s Advanced Computing Research Centre (ACRC) for computational resources.
References
Appendix A Optimal approximate posterior derivation
Starting with Eq (3), we start by combining the prior and likelihood for ,
then split out from and ,
Then notice that the inner expectation is a KL-divergence,
Appendix B Reparameterised variational inference
In variational inference, the ELBO objective takes the form,
As the distribution over which the expectation is taken is now independent of , we can form unbiased estimates of the gradient of by drawing one or a few samples of .
Appendix C Efficient convolutional linear regression
Working in one dimension for simplicity, the standard form for a convolution in deep learning is
where is the input image/feature-map, is the output feature-map, is the convolutional weights, indexes images, and index channels, indexes the location within the image, and indexes the location within the convolutional patch. Later, we will swap the identity of the “patch location” and the “image location” and to facilitate this, we define them both to be centred on zero,
where is the size of an image and is the size of a patch, such that, for example for a size 3 kernel, .
To use the expressions for the fully-connected case, we can form a new input, by cutting out each image patch,
Note that we have used commas to group pairs of indices ( and ) that may be combined into a single index (e.g. using a reshape operation). Indeed, combining and into a single index and combining and into a single index, this expression can be viewed as standard matrix multiplication,
This means that we can directly apply the approximate posterior we derived for the fully-connected case in Eq. (11) to the convolutional case. To allow for this, we take
While we can explicitly compute by extracting image patches, this imposes a very large memory cost (a factor of for a kernel, with stride , because there are roughly as many patches as pixels in the image, and a patch requires 9 times the storage of a pixel. To implement convolutional linear regression with a more manageable memory cost, we instead compute the matrices required for linear regression directly as convolutions of the input feature-maps, with themselves, and as convolutions of the with the output feature maps, , which we describe here.
For linear regression (Eq. 11), we first need to compute,
This can be directly viewed as the convolution of and , where we treat as the “convolutional weights”, as the location within the now very large (size ) “convolutional patch”, and as the location in the resulting output. Once we realise that the computation is a spatial convolution, it is possible to fit it into standard convolution functions provided by deep-learning frameworks (albeit with some rearrangement of the tensors).
To treat this as a convolution, we first need exact translational invariance, which can be achieved by using circular boundary conditions. Note that circular boundary conditions are not typically used in neural networks for images, and we therefore only use circular boundary conditions to define the approximate posterior over weights. The variational framework does not restrict us to also using circular boundary conditions within our feedforward network, and as such, we use standard zero-padding. With exact translational invariance, we can write this expression directly as a convolution,
i.e. for a size kernel, , where we treat as the “convolutional weights”, as the location within the “convolutional patch”, and as the location in the resulting output “feature-map”.
Finally, note that this form offers considerable benefits in terms of memory consumption. In particular, the output matrices are usually quite small — the number of channels is typically or , and the number of locations within a patch is typically , giving a very manageable total size that is typically smaller than .
Appendix D Wishart distributions with real-valued degrees of freedom
The classical description of the Wishart distribution,
However, for the purposes of defining learnable approximate posteriors, we need to be able to sample and evaluate the probability density when is positive real.
To do this, consider the alternative, much more efficient means of sampling from a Wishart distribution, using the Bartlett decomposition (Bartlett, 1933). The Bartlett decomposition gives the probability density for the Cholesky of a Wishart sample. In particular,
Here, is usually considered to be sampled from a distribution, but we have generalised this slightly using the equivalent Gamma distribution to allow for real-valued . Following Chafaï (2015), We need to change variables to rather than ,
Thus, the probability density for under the Bartlett sampling operation is
To convert this to a distribution on , we need the volume element for the transformation from to ,
which can be obtained directly by computing the log-determinant of the Jacobian for the transformation from to , or by taking the ratio of Eq. (54) and the usual Wishart probability density (with integral ). Thus,
where the second-to-last line was obtained by noting that covers the same range of integers as under the product. Finally, using the definition of the multivariate Gamma function,
We thereby re-obtain the probability density for the standard Wishart distribution,
Appendix E Full description of our method for deep Gaussian process
Appendix F Motivating the approximate posterior for deep GPs
Our original motivation for the approximate posterior was for the BNN case, which we then extended to deep GPs. Here, we show how the same approximate posterior can be motivated from a deep GP perspective. As with the BNN case, we first derive the form of the optimal approximate posterior for the last layer, in the regression case. Without inducing points, the ELBO becomes
Some straightforward rearrangements lead to a similar form to before,
where is the precision,
This can be understood as kernelized Bayesian linear regression conditioned on the features from the previous layers. Finally, as is usual in GPs (Rasmussen & Williams, 2006), the predictive distribution for test points can be obtained by conditioning using this approximate posterior.
Appendix G Comparing our deep GP approximate posterior to previous work
Appendix H Parameter scaling for ADAM
The standard optimiser for variational BNNs is ADAM (Kingma & Ba, 2014), which we also use. Considering similar RMSprop updates for simplicity (Tieleman & Hinton, 2012),
where the expectation over is approximated using a moving-average of past gradients. Thus, absolute parameter changes are going to be of order . This is fine if all the parameters have roughly the same order of magnitude, but becomes a serious problem if some of the parameters are very large and others are very small. For instance, if a parameter is around and , then a single ADAM step can easily double the parameter estimate, or change it from positive to negative. In contrast, if a parameter is around , then ADAM, with can make proportionally much smaller changes to this parameter, (around ). Thus, we need to ensure that all of our parameters have the same scale, especially as we mix methods, such as combining factorised and global inducing points. We thus design all our new approximate posteriors (i.e. the inducing inputs and outputs) such that the parameters have a scale of around . The key issue is that the mean weights in factorised methods tend to be quite small — they have scale around . To resolve this issue, we store scaled weights, and we divide these stored, scaled mean parameters by the fan-in as part of the forward pass,
This scaling forces us to use larger learning rates than are typically used.
Appendix I Exploring the effect of the number of inducing points
In this section, we briefly consider the effect of changing the number of inducing points, , used in global inducing. We reconsider the toy problem from Sec. 1, and plot predictive posteriors obtained with global inducing as the number of inducing points increases from 2 to 40 (noting that in Fig. 1 we used 100 inducing points).
We plot the results of our experiment in Fig. A2. While two inducing points are clearly not sufficient, we observe that there is remarkably very little difference between the predictive posteriors for 10 or more inducing points. This observation is reflected in the ELBOs per datapoint (listed above each plot), which show that adding more points beyond 10 gains very little in terms of closeness to the true posterior.
However, we note that this is a very simple dataset: it consists of only two clusters of close points with a very clear trend. Therefore, we would expect that for more complex datasets more inducing points would be necessary. We leave a full investigation of how many inducing points are required to obtain a suitable approximate posterior, such as that found in Burt et al. (2020) for sparse GP regression, to future work.
Appendix J Understanding compositional uncertainty
In this section, we take inspiration from the experiments of Ustyuzhaninov et al. (2020) which investigate the compositional uncertainty obtained by different approximate posteriors for DGPs. They noted that methods which factorise over layers have a tendency to cause the posterior distribution for each layer to collapse to a (nearly) deterministic function, resulting in worse uncertainty quantification within layers and worse ELBOs. In contrast, they found that allowing the approximate posterior to have correlations between layers allows those layers to capture more uncertainty, resulting in better ELBOs and therefore a closer approximation to the true posterior. They argue that this then allows the model to better discover compositional structure in the data.
We first consider a toy problem consisting of 100 datapoints generated by sampling from a two-layer DGP of width one, with squared-exponential kernels in each layer. We then fit two two-layer DGPs to this data - one using local inducing, the other using global inducing. The results of this experiment can be seen in Fig. A3, which show the final fit, along with the learned posteriors over intermediate functions and . These results mirror those observed by Ustyuzhaninov et al. (2020) on a similar experiment: local inducing, which factorises over layers, collapses to a nearly deterministic posterior over the intermediate functions, whereas global inducing provides a much broader distribution over functions for the two layers. Therefore, global inducing leads to a wider range of plausible functions that could explain the data via composition, which can be important in understanding the data. We observe that this behaviour directly leads to better uncertainty quantification for the out-of-distribution region, as well as better ELBOs.
To illustrate a similar phenomenon in BNNs, we reconsider the toy problem of Sec. 1. As it is not meaningful to consider neural networks with only one hidden unit per layer, instead of plotting intermediate functions we instead look at the mean variance of the functions at random input points, following roughly the experiment Dutordoir et al. (2019) in Table 1. For each layer, we consider the quantity
where the expectation is over random input points, which we sample from a standard normal. We expect that for methods which introduce correlations across layers, this quantity will be higher, as there will be a wider range of intermediate functions that could plausibly explain the data. We confirm this in Table 2, which indicates that global inducing leads to many more compositions of functions being considered as plausible explanations of the data. This is additionally reflected in the ELBO, which is far better for global inducing than the other, factorised methods. However, we note that the variances are far closer than we might otherwise expect for the second layer. We hypothesise that this is due to the pruning effects described in Trippe & Turner (2018), where a layer has many weights that are close to the prior that are then pruned out by the following layer by having the outgoing weights collapse to zero. In fact, we note that the variances in the last layer are small for the factorised methods, which supports this hypothesis. By contrast, global inducing leads to high variances across all layers.
We believe that understanding the role of compositional uncertainty in variational inference for deep Bayesian models can lead to important conclusions about both the models being used and the compositional structure underlying the data being modelled, and is therefore an important direction for future work to consider.
Appendix K UCI results with Bayesian neural networks
For this Appendix, we consider all of the UCI datasets from (Hernández-Lobato & Adams, 2015), along with four approximation families: factorised (i.e. mean-field), local inducing, global inducing, and facgi, which may offer some computational advantages to global inducing. We also considered three priors: the standard prior, NealPrior, and ScalePrior. The test LLs and ELBOs for BNNs applied to UCI datasets are given in Fig. A4. Note that the ELBOs for the global inducing methods (both global inducing and facgi) are almost always better than those for baseline methods, often by a very large margin. However, as noted earlier, this does not necessarily correspond to better test log likelihoods due to model misspecification: there is not a straightforward relationship between the ELBO and the predictive performance, and so it is possible to obtain better test log likelihoods with worse inference. We present all the results, including for the test RMSEs, in tabulated form in Appendix P.
The architecture we considered for all BNN UCI experiments were fully-connected ReLU networks with 2 hidden layers of 50 hidden units each. We performed a grid search to select the learning rate and minibatch size. For the fully factorised approximation, we selected the learning rate from and the minibatch size from , optimising for 25000 gradient steps; for the other methods we selected the learning rate from and fixed the minibatch size to 10000 (as in Salimbeni & Deisenroth, 2017), optimising for 10000 gradient steps. For all methods we selected the hyperparameters that gave the best ELBO. We trained the models using 10 samples from the approximate posterior, while using 100 for evaluation. For the inducing point methods, we used the selected batch size for the number of inducing points per layer. For all methods, we initialised the log noise variance at -3, but use the scaling trick in Appendix H to accelerate convergence, scaling by a factor of 10. Note that for the fully factorised method we used the local reparameterisation trick (Kingma et al., 2015); however, for fac gi we cannot do so because the inducing point methods require that covariances be propagated through the network correctly. For the inducing point methods, we additionally use output channel-specific precisions, , which effectively allows the network to prune unnecessary neurons if that benefits the ELBO. However, we only parameterise the diagonal of these precision matrices to save on computational and memory cost.
Appendix L Uncertainty calibration & out-of-distribution detection for CIFAR-10
To assess how well our methods capture uncertainty, we consider calibration, as well as the predictive entropy for out-of-distribution data. Calibration is assessed by comparing the model’s probabilistic assessment of its confidence with its accuracy — the proportion of the time that it is actually correct. For instance, gathering model predictions with some confidence (e.g. softmax probabilities in the range 0.9 to 0.95), and looking at the accuracy of these predictions, we would expect the model to be correct with probability 0.925; a higher or lower value would represent miscalibration.
We begin by plotting calibration curves (for the small ‘ResNet’ model) in Fig. A5, obtained by binning the predictions in 20 equal bins and assessing the mean accuracy of the binned predictions. For well-calibrated models, we expect the line to lie on the diagonal. A line above the diagonal indicates the model is underconfident (the model is performing better than it expects), whereas a line below the diagonal indicates it is overconfident (it is performing worse than it expects). While it is difficult to draw strong conclusions from these plots, it appears generally that factorised is poorly calibrated for both priors, that SpatialIWPrior generally improves calibration over ScalePrior, and that local inducing with SpatialIWPrior performs very well.
To come to more quantitative conclusions, we use expected calibration error (ECE; (Naeini et al., 2015; Guo et al., 2017)), which measures the expected absolute difference between the model’s confidence and accuracy. Confirming the results from the plots (Fig. A5), we find that using the more sophisticated SpatialIWPrior gave considerable improvements in calibration. While, as expected, we find that our most accurate prior, SpatialIWPrior, in combination with global inducing points did very well (ECE of 0.021), the model with the best ECE is actually local inducing with SpatialIWPrior, albeit by a very small margin. We leave investigation of exactly why this is to future work. Finally, note our final ECE value of 0.021 is a considerable improvement over those for uncalibrated models in Guo et al. (2017) (Table 1), which are in the region of 0.03-0.045 (although considerably better calibration can be achieved by post-hoc scaling of the model’s confidence).
In addition to calibration, we consider out-of-distribution performance. Given out-of-distribution data, we would hope that the network would give high output entropies (i.e. low confidence in the predictions), whereas for in-distribution data, we would hope for low entropy predictions (i.e. high confidence in the predictions). To evaluate this, we consider the mean predictive entropies for both the CIFAR-10 test set and the SVHN (Netzer et al., 2011) test set, when the network has been trained on CIFAR-10. We compare the ratio (CIFAR-10 entropy/ SVHN entropy) of these mean predictive entropies for each model in Table 4; a lower ratio indicates that the model is doing a better job of differentiating the datasets. We see that for both priors we considered, global inducing performs the best.
Appendix M UCI results with deep Gaussian processes
In this appendix, we again consider all of the UCI datasets from Hernández-Lobato & Adams (2015) for DGPs with depths ranging from two to five layers. We compare DSVI (Salimbeni & Deisenroth, 2017), local inducing, and global inducing. While local inducing uses the same inducing-point architecture as Salimbeni & Deisenroth (2017), the actual implementation and parameterisation is very different. As such, we do expect to see differences between local inducing and Salimbeni & Deisenroth (2017).
We show our results in Fig. A6. Here, the results are not as clear-cut as in the BNN case. For the smaller datasets (i.e. boston, concrete, energy, wine, but with the notable exception of yacht), global inducing generally outperforms both local inducing and DSVI, as noted in the main text, especially when considering the ELBOs. We do however observe that for power, protein, yacht, and one model for kin8nm, the local approaches sometimes outperform global inducing, even for the ELBOs. We believe this is due to the fact that relatively few inducing points were used (100), in combination with the fact that global inducing has far fewer variational parameters than the local approaches. This may make optimisation harder in the global inducing case, especially for larger datasets where the model uncertainty does not matter as much, as the posterior concentration will be stronger. Importantly, however, our results on CIFAR-10 indicate that these issues do not arise in very large-scale, high-dimensional datasets, which are of most interest for future work. Surprisingly, local inducing generally significantly outperforms DSVI, even though they are different parameterisations of the same approximate posterior. We leave consideration of why this is for future work.
We provide tabulated results, including for RMSEs, in Appendix P.
Here, we matched the experimental setup in Salimbeni & Deisenroth (2017) as closely as possible. In particular, we used 100 inducing points, and full-covariance observation noise. However, our parameterisation is still somewhat different from theirs, in part because our approximate posterior is defined in terms of noisy function-values, while their approximate posterior was defined in terms of the function-values themselves.
As the original results in Salimbeni & Deisenroth (2017) used different UCI splits, and did not provide the ELBO, we reran their code https://github.com/ICL-SML/Doubly-Stochastic-DGP (changing the number of epochs and noise variance to reflect the values in the paper), which gave very similar log likelihoods to those in the paper.
Appendix N MNIST 500
For MNIST, we considered a LeNet-inspired model consisting of two conv2d-relu-maxpool blocks, followed by conv2d-relu-linear, where the convolutions all have kernels with 64 channels. We trained all models using a learning rate of .
When training on very small datasets, such as the first 500 training examples in MNIST, we can see a variety of pathologies emerge with standard methods. To help build intuition for these pathologies, we introduce a sanity check for the ELBO. In particular, we could imagine a model that sets the distribution over all lower-layer parameters equal to the prior, and sets the top-layer parameters so as to ensure that the predictions are uniform. With classes, this results in an average test log likelihood of , and an ELBO (per datapoint) of approximately . We found that many combinations of the approximate posterior/prior converged to ELBOs near this baseline. Indeed, the only approximate posterior to escape this baseline for ScalePrior and SpatialIWPrior is global inducing points. This is because ScalePrior and SpatialIWPrior both offer the flexibility to shrink the prior variance, and hence shrink the weights towards zero, giving uniform predictions, and potentially zero KL divergence. In contrast, NealPrior and StandardPrior do not offer this flexibility: you always have to pay something in KL divergence in order to give uniform predictions. We believe that this is the reason that factorised performs better than expected with NealPrior, despite having an ELBO that is close to the baseline. Furthermore, it is unclear why local inducing gives very test log likelihood and performance, despite having an ELBO that is similar to factorised. For StandardPrior, all the ELBOs are far lower than the baseline, and far lower than for any other priors. Despite this, factorised and fac gi in combination with StandardPrior appear to transiently perform better in terms of predictive accuracy than any other method. These results should sound a note of caution whenever we try to use factorised approximate posteriors with fixed prior covariances (e.g. Blundell et al., 2015; Farquhar et al., 2020). We leave a full investigation of these effects for future work.
Appendix O Additional experimental details
All the methods were implemented in PyTorch. We ran the toy experiments on CPU, with the UCI experiments being run on a mixture of CPU and GPU. The remaining experiments – linear, CIFAR-10, and MNIST 500 – were run on various GPUs. For CIFAR-10, the most intensive of our experiments, we trained the models on one NVIDIA Tesla P100-PCIE-16GB. We optimised using ADAM (Kingma & Ba, 2014) throughout.
Inducing point methods
For global inducing, we initialise the inducing inputs, , and pseudo-outputs for the last layer, , using the first batch of data, except for the toy experiment, where we initialise using samples from (since we used more inducing points than datapoints). For the remaining layers, we initialise the pseudo-outputs by sampling from . We initialise the log precision to , except for the last layer, where we initialise it to 0. We additionally use a scaling factor of 3 as described in Appendix H. For local inducing, the initialisation is largely the same, except we initialise the pseudo-outputs for every layer by sampling from . We additionally sample the inducing inputs for every layer from .
Toy experiment
For each variational method, we optimise the ELBO over 5000 epochs, using full batches for the gradient descent. We use a learning rate of . We fix the noise variance at its true value, to help assess the differences between each method more clearly. We use 10 samples from the variational posterior for training, using 100 for testing. For HMC, we use 10000 samples to burn in, and 10000 samples for evaluation, which we subsequently thin by a factor of 10. We initialise the samples from a standard normal distribution, and use 20 leapfrog steps for each sample. We hand-tune the leapfrog step sizes to be 0.0007 and 0.003 for the burn-in and sampling phases, respectively.
Deep linear network
We use 10 inducing points for the inducing point methods. We use 1 sample from the approximate posterior for training and 10 for testing, training for 40 periods of 1000 gradient steps, using full batches for each step, with a learning rate of .
UCI experiments
The splits that we used for the UCI datasets can be found at https://github.com/yaringal/DropoutUncertaintyExps.
CIFAR-10
MNIST 500
The MNIST dataset (http://yann.lecun.com/exdb/mnist/) is a dataset of grayscale handwritten digits, each pixels, with 10 classes. It comprises 60,000 training images and 10,000 test images. For the MNIST 500 experiments, we trained using the first 500 images from the training dataset and discarded the rest. We normalised the images using the full training dataset’s statistics.
Deep GPs
As mentioned, we largely follow the approach of Salimbeni & Deisenroth (2017) for hyperparameters. For global inducing, we initialise the inducing inputs to the first batch of training inputs, and we initialise the pseudo-outputs for the last layer to the respective training outputs. For the remaining layers, we initialise the pseudo-outputs by sampling from a standard normal distribution. We initialise the precision matrix to be diagonal with log precision zero for the output layer, and log precision -4 for the remaining layers. For local inducing, we initialise inducing inputs and pseudo-outputs by sampling from a standard normal for every layer, and initialise the precision matrices to be diagonal with log precision zero.