Model Selection in Bayesian Neural Networks via Horseshoe Priors

Soumya Ghosh, Finale Doshi-Velez

Introduction

Bayesian Neural Networks (BNNs) are increasingly the de-facto approach for modeling stochastic functions. By treating the weights in a neural network as random variables, and performing posterior inference on these weights, BNNs can avoid overfitting in the regime of small data, provide well-calibrated posterior uncertainty estimates, and model a large class of stochastic functions with heteroskedastic and multi-modal noise. These properties have resulted in BNNs being adopted in applications ranging from active learning and reinforcement learning .

While there have been many recent advances in training BNNs , model-selection in BNNs has received relatively less attention. Unfortunately, the consequences for a poor choice of architecture are severe: too few nodes, and the BNN will not be flexible enough to model the function of interest; too many nodes, and the BNN predictions will have large variance because the posterior uncertainty in the weights will remain large. In other approaches to modeling stochastic functions such as Gaussian Processes (GPs), such concerns can be addressed via optimizing continuous kernel parameters; In BNNs, the number of nodes in a layer is a discrete quantity. Practitioners typically perform model-selection via onerous searches over different layer sizes.

In this work, we demonstrate that we can perform computationally-efficient and statistically-effective model selection in Bayesian neural networks by placing Horseshoe (HS) priors over the variance of weights incident to each node in the network. The HS prior has heavy tails and supports both zero values and large values. Fixing the mean of the incident weights to be zero, nodes with small variance parameters are effectively turned off—all incident weights will be close to zero—while nodes with large variance parameters can be interpreted as active. In this way, we can perform model selection over the number of nodes required in a Bayesian neural network.

While they mimic a spike-and-slab approach that would assign a discrete on-off variable to each node, the continuous relaxation provided by the Horseshoe prior keeps the model differentiable; with appropriate parameterization, we can take advantage of recent advances in variational inference (e.g. ) for training. We demonstrate that our approach avoids under-fitting even with the required number of nodes in the network is grossly over-estimated; we learn compact network structures without sacrificing—and sometimes improving—predictive performance.

Bayesian Neural Networks

A Bayesian neural network captures uncertainty in the weight parameters W\mathcal{W} by endowing them with distributions W∼p(W)\mathcal{W}\sim p(\mathcal{W}). Given a dataset of NN observation-response pairs D={xn,yn}n=1N\mathcal{D}=\{x_{n},y_{n}\}_{n=1}^{N}, we are interested in estimating the posterior distribution,

and leveraging the learned posterior for predicting responses to unseen data x∗x_{*}, p(y∗∣x∗)=∫p(y∗∣f(W,x∗))p(W∣D)dWp(y_{*}\mid x_{*})=\int p(y_{*}\mid f(\mathcal{W},x_{*}))p(\mathcal{W}\mid\mathcal{D})d\mathcal{W}. The prior p(W)p(\mathcal{W}) allows one to encode problem-specific beliefs as well as general properties about weights. In the past, authors have used fully factorized Gaussians on each weight wk′klw_{k^{\prime}kl}, structured Gaussians on each layer WlW_{l} as well as a two component scale mixture of Gaussians on each weight wk′klw_{k^{\prime}kl}. The scale mixture prior has been shown to encourage weight sparsity. In this paper, we show that by using carefully constructed infinite scale mixtures of Gaussians, we can induce heavy-tailed priors over network weights. Unlike previous work, we force all weights incident into a unit to share a common prior allowing us to induce sparsity at the unit level and prune away units that do not help explain the data well.

Automatic Model Selection through Horseshoe Priors

The distribution over weights in Equation 2 is called horseshoe prior . It exhibits Cauchy-like flat, heavy tails while maintaining an infinitely tall spike at zero. Consequently, it has the desirable property of allowing sufficiently large node weight vectors wklw_{kl} to escape un-shrunk—by having a large scale parameter—while providing severe shrinkage to smaller weights. This is in contrast to Lasso style regularizers and their Bayesian counterparts that provide uniform shrinkage to all weights. By forcing the weights incident on a unit to share scale parameters, the prior in Equation 2 induces sparsity at the unit level, turning off units that are unnecessary for explaining the data well. Intuitively, the shared layer wide scale υl\upsilon_{l} pulls all units in layer ll to zero, while the heavy tailed unit specific τkl\tau_{kl} scales allow some of the units to escape the shrinkage.

While a direct parameterization of the Half-Cauchy distribution in Equation 2 is possible, it leads to challenges during variational learning. Standard exponential family variational approximations struggle to capture the thick Cauchy tails, while a Cauchy approximating family leads to high variance gradients. Instead, we use a more convenient auxiliary variable parameterization ,

The joint distribution of the Horseshoe Bayesian neural network is then given by,

where p(yn∣f(W,xn))p(y_{n}\mid f(\mathcal{W},x_{n})) is an appropriate likelihood function, and, r(a,λ∣b)=Inv-Gamma(a∣1/2,1/λ)Inv-Gamma(λ∣1/2,1/b2)r(a,\lambda\mid b)=\text{Inv-Gamma}(a\mid 1/2,1/\lambda)\text{Inv-Gamma}(\lambda\mid 1/2,1/b^{2}), with θ={W,T,κ,ρκ}\mathcal{\theta}=\{\mathcal{W},\mathcal{T},\kappa,\rho_{\kappa}\}, T={{τkl}k=1,l=1K,L,{υl}l=1L,{λkl}k=1,l=1K,L,{ϑl}l=1L}\mathcal{T}=\{\{\tau_{kl}\}_{k=1,l=1}^{K,L},\{\upsilon_{l}\}_{l=1}^{L},\{\lambda_{kl}\}_{k=1,l=1}^{K,L},\{\vartheta_{l}\}_{l=1}^{L}\}.

Parameterizing for More Robust Inference: Non-Centered Parameterization

The horseshoe prior of Equation 2 exhibits strong correlations between the weights wklw_{kl} and the scales τklυl\tau_{kl}\upsilon_{l}. Indeed, its favorable sparsity inducing properties stem from this coupling. However, an unfortunate consequence is a strongly coupled posterior that exhibits pathological funnel shaped geometries and is difficult to reliably sample or approximate. Fully factorized approximations are particularly problematic and can lead to non-sparse solutions erasing the benefits of using the horseshoe prior.

Recent work suggests that the problem can be alleviated by adopting non-centered parameterizations. Consider a reformulation of Equation 2,

where the distribution on the scales are left unchanged. Such a parameterization is referred to as non-centered since the scales and weights are sampled from independent prior distributions and are marginally uncorrelated. The coupling between the two is now introduced by the likelihood, when conditioning on observed data. Non-centered parameterizations are known to lead to simpler posterior geometries . Empirically, we find that adopting a non-centered parameterization significantly improves the quality of our posterior approximation and helps us better find sparse solutions. Figure 1 summarizes the conditional dependencies assumed by the centered and the non-centered Horseshoe Bayesian neural networks model.

Learning Bayesian Neural Networks with Horseshoe priors

We restrict the variational distribution for the non-centered weight βij,l\beta_{ij,l} between units ii in layer l−1l-1 and jj in layer ll, q(βij,l∣ϕβijl)q(\beta_{ij,l}\mid\phi_{\beta_{ijl}}) to the Gaussian family N(βij,l∣μij,l,σij,l2)\mathcal{N}(\beta_{ij,l}\mid\mu_{ij,l},\sigma^{2}_{ij,l}). We will use β\beta to denote the set of all non-centered weights in the network. The non-negative scale parameters τkl\tau_{kl} and υl\upsilon_{l} and the variance of the output layer weights are constrained to the log-Normal family, q(ln τkl∣ϕτkl)=N(μτkl,στkl2)q(\text{ln }\tau_{kl}\mid\phi_{\tau_{kl}})=\mathcal{N}(\mu_{\tau_{kl}},\sigma^{2}_{\tau_{kl}}), q(ln υl∣ϕυl)=N(μυl,συl2)q(\text{ln }\upsilon_{l}\mid\phi_{\upsilon_{l}})=\mathcal{N}(\mu_{\upsilon_{l}},\sigma^{2}_{\upsilon_{l}}), and q(ln κ∣ϕκ)=N(μκ,σκ2)q(\text{ln }\kappa\mid\phi_{\kappa})=\mathcal{N}(\mu_{\kappa},\sigma^{2}_{\kappa}). We do not impose a distributional constraint on the variational approximations of the auxiliary variables ϑl\vartheta_{l}, λkl\lambda_{kl}, or ρκ\rho_{\kappa}, but we will see that conditioned on the remaining variables the optimal variational family for these latent variables follow inverse Gamma distributions. Evidence Lower Bound The resulting evidence lower bound (ELBO),

is challenging to evaluate. The non-linearities introduced by the neural network and the potential lack of conjugacy between the neural network parameterized likelihoods and the Horseshoe priors render the first expectation in Equation 7 intractable. Consequently, the traditional prescription of optimizing the ELBO by cycling through a series of fixed point updates is no longer available.

Recent progress in black box variational inference provides a recipe for subverting this difficulty. These techniques provide noisy unbiased estimates of the gradient ∇ϕL^(ϕ)\nabla_{\phi}\hat{\mathcal{L}}(\phi), by approximating the offending expectations with unbiased Monte-Carlo estimates and relying on either score function estimators or reparameterization gradients to differentiate through the sampling process. With the unbiased gradients in hand, stochastic gradient ascent can be used to optimize the ELBO. In practice, reparameterization gradients exhibit significantly lower variances than their score function counterparts and are typically favored for differentiable models. The reparameterization gradients rely on the existence of a parameterization that separates the source of randomness from the parameters with respect to which the gradients are sought. For our Gaussian variational approximations, the well known non-centered parameterization, ζ∼N(μ,σ2)⇔ϵ∼N(0,1),ζ=μ+σϵ\zeta\sim\mathcal{N}(\mu,\sigma^{2})\Leftrightarrow\epsilon\sim\mathcal{N}(0,1),\zeta=\mu+\sigma\epsilon, allows us to compute Monte-Carlo gradients,

for any differentiable function gg and ϵ(s)∼N(0,1)\epsilon^{(s)}\sim\mathcal{N}(0,1). Further, as shown in , the variance of the gradient estimator can be provably lowered by noting that the weights in a layer only affect L(ϕ)\mathcal{L}(\phi) through the layer’s pre-activations and directly sampling from the relatively lower-dimensional variational posterior over pre-activations.

Recall that the pre-activation of node kk in layer ll, uklu_{kl} in our non-centered model is ukl=τklυlβklT[a,1]Tu_{kl}=\tau_{kl}\upsilon_{l}\beta_{kl}^{T}[a,1]^{T}. The variational posterior for the pre-activations is given by,

Algorithm

We now have all the tools necessary for optimizing Equation 7. By recursively sampling from the variational posterior of Equation 9 for each layer of the network, we are able to forward propagate information through the network. Owing to the reparameterizations (Equation 8), we are also able to differentiate through the sampling process and use reverse mode automatic differentiation tools to compute the relevant gradients. With the gradients in hand, we optimize L(ϕ)\mathcal{L}(\phi) with respect to the variational weights ϕβ\phi_{\beta}, per-unit scales ϕτkl\phi_{\tau_{kl}}, per-layer scales ϕυl\phi_{\upsilon_{l}}, and the variational scale for the output layer weights, ϕκ\phi_{\kappa} using Adam . Conditioned on these, the optimal variational posteriors of the auxiliary variables ϑl\vartheta_{l}, λkl\lambda_{kl}, and ρκ\rho_{\kappa} follow Inverse Gamma distributions. Fixed point updates that maximize L(ϕ)\mathcal{L}(\phi) with respect to ϕϑl,ϕλkl,ϕρκ\phi_{\vartheta_{l}},\phi_{\lambda_{kl}},\phi_{\rho_{\kappa}}, holding the other variational parameters fixed are available. The overall algorithm, involves cycling between gradient and fixed point updates to maximize the ELBO in a coordinate ascent fashion.

Related Work

Early work on Bayesian neural networks can be traced back to . These early approaches relied on Laplace approximation or Markov Chain Monte Carlo (MCMC) for inference. They do not scale well to modern architectures or the large datasets required to learn them. Recent advances in stochastic variational methods , black-box variational and alpha-divergence minimization , and probabilistic backpropagation have reinvigorated interest in BNNs by allowing inference to scale to larger architectures and larger datasets.

Work on learning structure in BNNs remains relatively nascent. In the authors use a cascaded Indian buffet process to learn the structure of sigmoidal belief networks. While interesting, their approach appears susceptible to poor local optima and their proposed Markov Chain Monte Carlo based inference does not scale well. More recently, introduce a mixture-of-Gaussians prior on the weights, with one mixture tightly concentrated around zero, thus approximating a spike and slab prior over weights. Their goal of turning off edges is very different than our approach, which performs model selection over the appropriate number of nodes. Further, our proposed Horseshoe prior can be seen as an extension of their work, where we employ an infinite scale mixture-of-Gaussians. Beyond providing stronger sparsity, this is attractive because it obviates the need to directly specify the mixture component variances or the mixing proportion as is required by the prior proposed in . Only the prior scales of the variances needs to be specified and in our experiments, we found results to be relatively robust to the values of these scale hyper-parameters. Recent work indicates that further gains may be possible by a more careful tuning of the scale parameters. Others have noticed connections between Dropout and approximate variational inference. In particular, show that the interpretation of Gaussian dropout as performing variational inference in a network with log uniform priors over weights leads to sparsity in weights. This is an interesting but orthogonal approach, wherein sparsity stems from variational optimization instead of the prior.

There also exists work on learning structure in non-Bayesian neural networks. Early work pruned networks by analyzing second-order derivatives of the objectives. More recently, describe applications of structured sparsity not only for optimizing filters and layers but also computation time. Closest to our work in spirit, , and who use group sparsity to prune groups of weights—e.g. weights incident to a node. However, these approaches don’t model the uncertainty in weights and provide uniform shrinkage to all parameters. Our horseshoe prior approach similarly provides group shrinkage while still allowing large weights for groups that are active.

Experiments

In this section, we present experiments that evaluate various aspects of the proposed Bayesian neural network with horseshoe priors (HS-BNN). We begin with experiments on synthetic data that showcase the model’s ability to guard against under fitting and recover the underlying model. We then proceed to benchmark performance on standard regression and classification tasks. For the regression problems we use Gaussian likelihoods with an unknown precision γ\gamma, p(yn∣f(W,xn),γ)=N(yn∣f(W,xn),γ−1)p(y_{n}\mid f(\mathcal{W}_{,}x_{n}),\gamma)=\mathcal{N}(y_{n}\mid f(\mathcal{W}_{,}x_{n}),\gamma^{-1}). We place a vague prior on the precision,γ∼Gamma(6,6)\gamma\sim\text{Gamma}(6,6) and approximate the posterior over γ\gamma using a Gamma distribution. The corresponding variational parameters are learned via a gradient update during learning. We use a Categorical likelihood for the classification problems. In a preliminary study, we found larger mini-batch sizes improved performance, and in all experiments we use a batch size of 512512. The hyper parameters b0b_{0} and bgb_{g} are both set to one.

We begin with a one-dimensional non linear regression problem shown in Figure 2. To explore the effect of additional modeling capacity on performance, we sample twenty points uniformly at random in the interval [−4,+4][-4,+4] from the function yn=xn3+ϵ,ϵ∼N(0,9)y_{n}=x_{n}^{3}+\epsilon,\epsilon\sim\mathcal{N}(0,9) and train single layer Bayesian neural networks with 5050, 100100 and 10001000 units each. We compare HS-BNN against a BNN with Gaussian priors on weights, wij,l∼N(0,κ),κ∼C+(0,5)w_{ij,l}\sim\mathcal{N}(0,\kappa),\kappa\sim C^{+}(0,5), training both for a 10001000 iterations. The performance of the BNN with Gaussian priors quickly deteriorates with increasing capacity as a result of under fitting the limited amount of training data. In contrast, HS-BNN by pruning away additional capacity is more robust to model misspecification showing only a marginal drop in predictive performance with increasing number of units.

Non-centered parameterization

Next, we explore the benefits of the non-centered parameterization. We consider a simple two dimensional classification problem generated by sampling data uniformly at random from [−1,+1]×[−1,+1][-1,+1]\times[-1,+1] and using a 2-2-1 network, whose parameters are known a-priori to generate the class labels. We train three Bayesian neural networks with a 1515 unit layer on this data, with Gaussian priors, with horseshoe priors but employing a centered parameterization, and with the non-centered horseshoe prior. Each model is trained till convergence. We find that all three models are able to easily fit the data and provide high predictive accuracy. However, the structure learned by the three models are very different. In Figure 2 we visualize the distribution of weights incident onto a unit. Unsurprisingly, the BNN with Gaussian priors does not exhibit sparsity. In contrast, models employing the horseshoe prior are able to prune units away by setting all incident weights to tiny values. It is interesting to note that even for this highly stylized example the centered parameterization struggles to recover the true structure of the underlying network. The non-centered parameterization however does significantly better and prunes away all but two units. Further experiments provided in the supplement demonstrate the same effect for wider 100 unit networks. The non-centered parameterized model is again able to recover the two active units.

2 Classification and Regression experiments

We benchmark classification performance on the MNIST dataset. Additional experiments on a gesture recognition task are available in the supplement. We compare HS-BNN against the variational matrix Gaussian (VMG) , a BNN with a two-component scale mixture (SM-BNN) prior on weights proposed in and a BNN with Gaussian prior (BNN) on weights. VMG uses a structured variational approximation, while the other approaches all use fully factorized approximations and differ only in the type of prior used. These approaches constitute the state-of-the-art in variational learning for Bayesian neural networks.

We preprocessed the images in the MNIST digits dataset by dividing the pixel values by 126. We explored networks with varying widths and depths all employing rectified linear units. For HS-BNN we used Adam with a learning rate of 0.0050.005 and 500500 epochs. We did not use a validation set to monitor validation performance or tune hyper-parameters. We used the parameter settings recommended in the original papers for the competing methods. Figure 3 summarizes our findings. We showcase results for three architectures with two hidden layers each containing 400400, 800800 and 12001200 rectified linear hidden units. Across architectures, we find our performance to be significantly better than BNN, comparable to SM-BNN, and worse than VMG. The poor performance with respect to VMG likely stems from the structured matrix variate variational approximation employed by VMG.

Discussion and Conclusion

In Section 6, we demonstrated that a properly parameterized horseshoe prior on the scales of the weights incident to each node is a computationally efficient tool for model selection in Bayesian neural networks. Decomposing the horseshoe prior into inverse gamma distributions and using a non-centered representation ensured a degree of robustness to poor local optima. While we have seen that the horseshoe prior is an effective tool for model selection, one might wonder about more common alternatives. We lay out a few obvious choices and contrast their deficiencies. One starting point is to observe that a node can be pruned if all its incident weights are zero (in this case, it can only pass on the same bias term to the rest of the network). Such sparsity can be encouraged by a simple exponential prior on the weight scale, but without heavy tails all scales are forced artificially low and prediction suffers and has been noted in the context of learning sparse neural networks . In contrast, simply using a heavy-tail prior on the scale parameter, such as a half-Cauchy, will not apply any pressure to set small scales to zero, and we will not have sparsity. Both the shrinkage to zero and the heavy tails of the horseshoe prior are necessary to get the model selection that we require. And importantly, using a continuous prior with the appropriate statistical properties is simple to incorporate with existing inference, unlike an explicit spike and slab model. Another alternative is to observe that a node can be pruned if the product z⋅wz\cdot w is nearly constant for all inputs zz—having small weights is sufficient to achieve this property; weights ww that are orthogonal to the variation in zz is another. Thus, instead of putting a prior over the scale of ww, one could put a prior over the scale of the variation in z⋅wz\cdot w. While we believe this is more general, we found that such a formulation has many more local optima and thus harder to optimize.

References

Appendix A Fixed point updates

The ELBO corresponding to the non-centered HS model is,

The auxiliary variables ρκ\rho_{\kappa}, ϑl\vartheta_{l} and ϑl\vartheta_{l} all follow inverse Gamma distributions. Here we derive for λkl\lambda_{kl}, the others follow analogously. Consider,

Appendix B Additional Experiments

Here we provide an additional experiment with the data setup in Section 6.2. We use the same linearly separable data, but train larger networks with 100 units each. Figure 4 shows the inferred weights under the different models. Observe that the non-centered HS-BNN is again able to prune away extra capacity and recover two active nodes.

B.2 Further Exploration of Model Selection Properties

B.3 Gesture Recognition

We also experimented with a gesture recognition dataset that consists of 24 unique aircraft handling signals performed by 20 different subjects, each for 20 repetitions. The task consists of recognizing these gestures from kinematic, tracking and video data. However, we only use kinematic and tracking data. A couple of example gestures are visualized in Figure 7. The dataset contains 96009600 gesture examples.

A 12-dimensional vector of body features (angular joint velocities for the right and left elbows and wrists), as well as an 8 dimensional vector of hand features (probability values for hand shapes for the left and right hands) collected by Song et al. are provided as features for all frames of all videos in the dataset. We additionally used the 20 dimensional per-frame tracking features made available in . We constructed features to represent each gesture by first extracting frames by sampling uniformly in time and then concatenating the per-frame features of the selected frames to produce 600-dimensional feature vectors.

This is a much smaller dataset than MNIST and recent work has demonstrated that a BNN with Gaussian priors performs well on this task. Figure 7 compares the performance of HS-BNN with competing methods. We train a two layer HS-BNN with each layer containing 400 units. The error rates reported are a result of averaging over 5 random 75/25 splits of the dataset. Similar to MNIST, HS-BNN significantly outperforms BNN and is competitive with VMG and SM-BNN. We also see strong sparsity, just as in MNIST.