Sylvester Normalizing Flows for Variational Inference

Rianne van den Berg, Leonard Hasenclever, Jakub M. Tomczak, Max Welling

INTRODUCTION

Stochastic variational inference (Hoffman et al., 2013) allows for posterior inference in increasingly large and complex problems using stochastic gradient ascent. In continuous latent variable models, variational inference can be made particularly efficient through the amortized inference, in which inference networks amortize the cost of calculating the variational posterior for a data point (Gershman and Goodman, 2014). A particularly successful class of models is the variational autoencoder (VAE) in which both the generative model and the inference network are given by neural networks, and sampling from the variational posterior is efficient through the non-centered parameterization (Kingma and Welling, 2014), also known as the reparameterization trick (Kingma and Welling, 2013; Rezende et al., 2014).

Despite its success, variational inference has drawbacks compared to other inference methods such as MCMC. Variational inference searches for the best posterior approximation within a parametric family of distributions. Hence, the true posterior distribution can only be recovered exactly if it happens to be in the chosen family. In particular, with widely used simple variational families such as diagonal covariance Gaussian distributions, the variational approximation is likely to be insufficient. More complex variational families enable better posterior approximations, resulting in improved model performance. Therefore, designing tractable and more expressive variational families is an important problem in variational inference (Nalisnick et al., 2016; Salimans et al., 2015; Tran et al., 2015).

Rezende and Mohamed (2015) introduced a general framework for constructing more flexible variational distributions, called normalizing flows. Normalizing flows transform a base density through a number of invertible parametric transformations with tractable Jacobians into more complicated distributions. They proposed two classes of normalizing flows: planar flows and radial flows. While effective for small problems, these can be hard to train and often many transformations are required to get good performance. For planar flows, Kingma et al. (2016) argue that this is due to the fact that the transformation used acts as a bottleneck, warping one direction at a time. Having a large number of flows makes the inference network very deep and harder to train, empirically resulting in suboptimal performance. Kingma et al. (2016) proposed inverse auto-regressive flows (IAF), achieving state of the art results on dynamically binarized MNIST at the time of publication. While very successful, each transformation in IAF only depends on the datapoint x\mathbf{x} through a context vector, with flow parameters that are independent of the datapoint.

Paper contribution In this paper, we use Sylvester’s determinant identity to introduce Sylvester normalizing flows (SNFs). This family of flows is a generalization of planar flows, removing the bottleneck. We compare a number of different variants of SNFs and show that they compare favorably against planar flows and IAFs. We show that one specific variant of SNF is related to IAF, with the main difference being the amortization strategy of the flow parameters. Besides the usual requirement of having flexible transformations, this demonstrates the importance of having data-dependent flow parameters. Note that this concept generalizes to applying normalizing flows to any conditional distribution, in the sense that the transformation parameters should be functions of the conditioning variable.

VARIATIONAL INFERENCE

Consider a probabilistic model with observations x\mathbf{x} and continuous latent variables z\mathbf{z} and model parameters θ\theta. In generative modeling we are often interested in performing maximum (marginal) likelihood learning of the parameters θ\mathbf{\theta} of the latent-variable model pθ(x,z)p_{\mathbf{\theta}}(\mathbf{x},\mathbf{z}). This requires marginalization over the unobserved latent variables z\mathbf{z}. Unfortunately, this integration is generally intractable. Variational inference (Jordan et al., 1999) instead introduces a variational approximation q(z∣x)q(\mathbf{z}|\mathbf{x}) to the posterior, to construct a lower bound on the log marginal likelihood:

This bound is known as the evidence lower bound (ELBO) and F\mathcal{F} is referred to as the variational free energy. In equation (2), the first term represents the reconstruction error, and the second term is the Kullback-Leibler (KL) divergence from the approximate posterior to the prior distribution, which acts as a regularizer. In this paper we consider variational autoencoders (VAEs), where both pθ(x∣z)p_{\theta}(\mathbf{x}|\mathbf{z}) and q(z∣x)q(\mathbf{z}|\mathbf{x}) are distributions whose parameters are given by neural networks. That is, we perform amortized inference such that q(z∣x)=qϕ(z∣x)q(\mathbf{z}|\mathbf{x})=q_{\phi}(\mathbf{z}|\mathbf{x}) with ϕ\phi the parameters of the encoder neural network. The parameters θ\theta and ϕ\phi of the generative model and inference model, respectively, are trained jointly through stochastic minimization of F(θ,ϕ)\mathcal{F}(\theta,\phi), which can be made efficient through the reparameterization trick (Kingma and Welling, 2013; Rezende et al., 2014).

From equation (1) we see that the better the variational approximation to the posterior the tighter the ELBO. The simplest, but probably most widely used choice of variation distribution qϕ(z∣x)q_{\phi}(\mathbf{z}|\mathbf{x}) is diagonal-covariance Gaussians of the form N(z∣μϕ(x), σϕ2(x))\mathcal{N}(\mathbf{z}|\boldsymbol{\mu}_{\phi}(\mathbf{x}),~{}\boldsymbol{\sigma}_{\phi}^{2}(\mathbf{x}))

However, with such simple variational distributions the ELBO will be fairly loose, resulting in biased maximum likelihood estimates of the model parameters θ\mathbf{\theta} (see Fig. 1) and harming generative performance. Thus, for variational inference to work well, more flexible approximate posterior distributions are needed.

Rezende and Mohamed (2015) propose a way to construct more flexible posteriors by transforming a simple base distribution with a series of invertible transformations (known as normalizing flows) with easily computable Jacobians. The resulting transformed density after one such transformation ff is as follows (Tabak and Turner, 2013; Tabak and Vanden-Eijnden, 2010):

This strategy is used in variational inference as follows: first, a stochastic variable is drawn from a simple base posterior distribution such as a diagonal Gaussian N(z0∣μ(x),σ2(x))\mathcal{N}(\mathbf{z}_{0}|\boldsymbol{\mu}(\mathbf{x}),\boldsymbol{\sigma}^{2}(\mathbf{x})). The sample is then transformed with a number of flows. After applying KK flows, the final latent stochastic variables are given by zK=fK∘…f2∘f1(z0)\mathbf{z}_{K}=f_{K}\circ\dots f_{2}\circ f_{1}(\mathbf{z}_{0}). The corresponding log-density is then given by:

where λk\lambda_{k} are the parameters of the kk-th transformation. Note that in order to achieve a flexible amortization strategy, the flow parameters λk\lambda_{k} can be made dependent on the input data: λk=λk(x)\lambda_{k}=\lambda_{k}(\mathbf{x}) (Rezende and Mohamed, 2015). Given a variational posterior qϕ(z∣x)=qK(z∣x)q_{\phi}(\mathbf{z}|\mathbf{x})=q_{K}(\mathbf{z}|\mathbf{x}) parametrized by a normalizing flow of length K, the variational objective can be rewritten as:

Normalizing flows are frequently applied to amortized variational inference. Instead of learning the parameters of the posterior distribution for each data point, (such as μ\boldsymbol{\mu} and σ\boldsymbol{\sigma} for a Gaussian posterior), the input-dependence of the posterior distribution parameters is modeled through an encoder/inference network. When performing amortized inference for normalizing flows, the flow parameters determine the final distribution, and should thus also be considered functions of the datapoint x\mathbf{x}. This can be achieved through the use of hypernetworks (Ha et al., 2016).

Rezende and Mohamed (2015) introduced a normalizing flow, called planar flow, for which the Jacobian determinant could be computed efficiently. A single transformation of the planar flow is given by:

By the Matrix determinant lemma the Jacobian of this transformation is given by:

where h′h^{\prime} denotes the derivative of hh and which can be computed in O(D)O(D) time.

In practice, many planar flow transformations are required to transform a simple base distribution into a flexible distribution, especially for high dimensional latent spaces. Kingma et al. (2016) argue that this is related to the term uh(wTz+b)\mathbf{u}h(\mathbf{w}^{T}\mathbf{z}+b) in Eq. (8), which effectively acts as a single-neuron MLP. In the next section we will derive a generalization of planar flows, which does not have a single-neuron bottleneck, while still maintaining the property of an efficiently computable Jacobian determinant.

SYLVESTER NORMALIZING FLOWS

Consider the following more general transformation similar to a single layer MLP with MM hidden units and a residual connection:

where IM\mathbf{I}_{M} and ID\mathbf{I}_{D} are MM and DD-dimensional identity matrices, respectively.

When M<DM<D, the computation of the determinant of a D×DD\times D matrix is thus reduced to the computation of the determinant of an M×MM\times M matrix.

Using Sylvester’s determinant identity, the Jacobian determinant of the transformation in Eq. (10) is given by:

Since Sylvester’s determinant identity plays a crucial role in the proposed family of normalizing flows, we will refer to them as Sylvester normalizing flows.

In general, the transformation in (10) will not be invertible. Therefore, we propose the following special case of the above transformation:

Let us now consider the case when R\mathbf{R} is an upper triangular matrix. By the argument for the diagonal case above, it suffices to consider the effect of the transformation in W\mathcal{W}. Multiplying (13) by QT\mathbf{Q}^{T} from the left gives:

Since fMf_{M} is invertible we can write vM=fM−1(vM′)v_{M}=f^{-1}_{M}(v_{M}^{\prime}). Now suppose we have expressed {vj,∀j>k}\{v_{j},\forall j>k\} in terms of {vj′,∀j>k}\{v^{\prime}_{j},\forall j>k\}. Then

Thus we have expressed {vj,∀j≥k}\{v_{j},\forall j\geq k\} in terms of {vj′,∀j≥k}\{v^{\prime}_{j},\forall j\geq k\}. By induction, we can express {vj,∀j}\{v_{j},\forall j\} in terms of {vj′,∀j}\{v^{\prime}_{j},\forall j\} and hence the transformation is invertible.

Hence the transformation in (22) is invertible. ∎

2 PRESERVING ORTHOGONALITY OF 𝐐𝐐\mathbf{Q}

Orthogonality is a convenient property, mathematically, but hard to achieve in practice. In this paper we consider three different flows based on the theorem above and various ways to preserve the orthogonality of Q\mathbf{Q}. The first two use explicit differentiable constructions of orthogonal matrices, while the third variant assumes a specific fixed permutation matrix as the orthogonal matrix.

First, we consider a Sylvester flow using matrices with MM orthogonal columns (O-SNF). In this flow we can choose M<DM<D, and thus introduce a flexible bottleneck. Similar to (Hasenclever et al., 2017), we ensure orthogonality of Q\mathbf{Q} by applying the following differentiable iterative procedure proposed by (Björck and Bowie, 1971; Kovarik, 1970):

Since this orthogonalization procedure is differentiable, it allows for the calculation of gradients with respect to Q(0)\mathbf{Q}^{(0)} by backpropagation, allowing for any standard optimization scheme such as stochastic gradient descent to be used for updating the flow parameters.

It is worth noting that performing a single Householder transformation is very cheap to compute, as it only requires DD parameters. Chaining together several Householder transformations results in more general orthogonal matrices, and it can be shown (Bischof and Sun, 1997; Sun and Bischof, 1995) that any M×MM\times M orthogonal matrix can be written as the product of M−1M-1 Householder transformations. In our Householder Sylvester flow, the number of Householder transformations HH is a hyperparameter that trades off the number of parameters and the generality of the orthogonal transformation. Note that the use of Householder transformations forces us to use M=DM=D, since Householder transformation result in square matrices.

3 DATA-DEPENDENT FLOW PARAMETERS

RELATED WORK

A number of invertible transformations with tractable Jacobians have been proposed in recent years. Rezende and Mohamed (2015) first discussed such transformations in the context of stochastic variation inference, coining the term normalizing flows.

Rezende and Mohamed (2015) proposed two different parametric families of transformations with tractable Jacobians: planar and radial flows. While effective for small problems, these transformations are hard to scale to large latent spaces and often require a large number of transformations. The transformation corresponding to planar flows is given in Eq. (8).

More recently, a successful class of flows called Inverse Autoregressive Flows was introduced in (Kingma et al., 2016). As the name suggests, one IAF transformation can be seen as the inverse of an autoregressive transformation. Consider the following autoregressive transformation:

with ϵ∼N(0,I)\boldsymbol{\epsilon}\sim\mathcal{N}(\mathbf{0},\mathbf{I}). This transformation models the distribution over the variable z\mathbf{z} with an autoregressive factorization p(z)=p(z0)∏i=1Dp(zi∣zi−1,…,z0)p(\mathbf{z})=p(z_{0})\prod_{i=1}^{D}p(z_{i}|z_{i-1},\ldots,z_{0}). Since the parameters of transformation for ziz_{i} are dependent on z1:i−1\mathbf{z}_{1:i-1}, this procedure requires DD sequential steps to sample a single vector z\mathbf{z}. This is undesirable for variational inference, where sampling occurs for every forward pass.

However, the inverse transformation (which exists if σˉi>0\bar{\sigma}_{i}>0 ∀i\forall i) is easy to sample from:

For this inverse transformation, ϵi\epsilon_{i} is no longer dependent on the transformation of ϵj\epsilon_{j} for j≠ij\neq i. Hence, this transformation can be computed in parallel: ϵ=(z−μˉ(z))/σˉ(z)\boldsymbol{\epsilon}=(\mathbf{z}-\boldsymbol{\bar{\mu}}(\mathbf{z}))/\boldsymbol{\bar{\sigma}}(\mathbf{z}). Rewriting σi(z1:i−1)=1/σˉi(z1:i−1)\sigma_{i}(z_{1:i-1)}=1/\bar{\sigma}_{i}(z_{1:i-1)} and μi(z1:i−1)=−μˉ(z1:i−1)/σˉi(z1:i−1)\mu_{i}(z_{1:i-1)}=-\bar{\mu}(z_{1:i-1})/\bar{\sigma}_{i}(z_{1:i-1)}, yields the IAF transformation:

Starting from z0∼N(0,I)\mathbf{z}^{0}\sim\mathcal{N}(\mathbf{0},\mathbf{I}), multiple IAF transformations can be stacked on top of each other to produce flexible probability distributions.

If μt\boldsymbol{\mu}^{t} and σt\boldsymbol{\sigma}^{t} depend on zt−1\mathbf{z}^{t-1} linearly, IAF can model full covariance Gaussian distributions. In order to move away from Gaussian distributions to more flexible distributions, it is important that μt\boldsymbol{\mu}^{t} and σt\boldsymbol{\sigma}^{t} are nonlinear functions of zt−1\mathbf{z}^{t-1}.

In practice, wide MADEs (Germain et al., 2015) or deep PixelCNN layers (van den Oord et al., 2016) are needed to increase the flexibility of IAF transformations. This results in transformations with a large number of parameters. As shown in Figure 2 (right), amortization is achieved through a context h(x)\mathbf{h}(\mathbf{x}) that is fed into the autoregressive networks as an additional input at every IAF step.

Householder Sylvesters flows can also be seen as a non-linear extension of Householder flows (Tomczak and Welling, 2016). Householder flows are volume-preserving flows, which transform the variational posterior with a diagonal covariance matrix to a full-covariance posterior. Householder flows are a special case of H-SNF if h(z)=zh(\mathbf{z})=\mathbf{z}, R\mathbf{R} is the identity matrix, and the residual connection in Eq. (13) is left out.

2 NORMALIZING FLOWS FOR DENSITY ESTIMATION

A number of invertible transformations have been proposed in the context of density estimation. Note that density estimation requires the inverse of the flow to be tractable. Having a provably invertible transformation is not the same as being able to compute the inverse.

For density estimation with normalizing flows, we are interested maximizing the log-likelihood of the data:

Thus, the goal is to transform a complicated data distribution back to a simple distribution. In general, both directions of an invertible transformations need not be tractable. Hence, methods developed for density estimation are generally not directly applicable to variational inference.

Non-linear independent component estimation (NICE, Dinh et al. (2014)) and the related Real NVP (Dinh et al., 2016), and Masked Autoregressive Flow (MAF, Papamakarios et al. (2017)) are recent examples of normalizing flows for density estimation.

In NICE, each transformation splits the variables into two disjoint subsets zA,zB\mathbf{z}_{A},\mathbf{z}_{B}. One of the subsets is transformed as zA′=zA+f(zB)\mathbf{z}_{A}^{\prime}=\mathbf{z}_{A}+f(\mathbf{z}_{B}), while zB\mathbf{z}_{B} is left unchanged. In the next transformation a different subset of variables is transformed. This results in a transformation which is trivially invertible and has a tractable Jacobian. Real NVP uses the same fundamental idea. Appealingly, because of the tractable inverse, NICE and real NVP can generate data and estimate density with one forward pass. However due to fact that only a subset of variables is updated in each transformation many transformations are needed in practice. Rezende and Mohamed (2015) compared NICE to planar flows in the context of variational inference and found that planar flows empirically perform better.

Finally, Papamakarios et al. (2017) showed that fitting an MAF can be seen as fitting an implicit IAF from the data distribution to the base distribution. However, generating data from an MAF density model requires DD passes, making it unappealing for variational inference.

NUMBER OF PARAMETERS

Here, we briefly compare the number of parameters needed by planar flows, IAF and the three Sylvester normalizing flows. We denote the size of the stochastic variables zz with DD, and the number of output units of the inference network with EE.

For the implementation of IAF as described in Section 6, the inference network needs to produce a context of size CC, where CC denotes the width of the MADE layers. The total number of flow related learnable parameters then comes down to EC+K×(C2+3CD)EC+K\times(C^{2}+3CD).

In the case of Orthogonal Sylvester flows with a bottleneck of size MM, we require KE×(MD+2M2+M)KE\times(MD+2M^{2}+M) parameters. For Householder Sylvester flows with HH Householder reflections per flow transformation, KE×(HD+2D2+D)KE\times(HD+2D^{2}+D) parameters are needed. Finally, for triangular Sylvester flows KE×(2D2+D)KE\times(2D^{2}+D) parameters require optimization.

Planar flows require the smallest number of parameters but generally result in worse results. IAFs on the other hand require a number of parameters that is quadratic in the width of the MADE layers. For good results this has to be quite large. In contrast, for SNFs the number of parameters is quadratic in the dimension of the latent space and while large, this can still be amortized.

EXPERIMENTS

We perform empirical studies of the performance of Sylvester flows on four datasets: statically binarized MNIST, Freyfaces, Omniglot and Caltech 101 Silhouettes. The baseline model is a plain VAE with a fully factorized Gaussian distribution. We furthermore compare against planar flows and Inverse Autoregressive Flows of different sizes.

We use annealing to optimize the lower bound, where the prefactor of the KL divergence is linearly increased from 0 to 1 during 100 epochs as suggested by Bowman et al. (2015) and Sønderby et al. (2016). A learning rate of 0.00050.0005 was used in all experiments. In order to obtain estimates for the negative log likelihood we used importance sampling (as proposed in (Rezende et al., 2014)). Unless otherwise stated, 5000 importance samples were used.

In order to assess the performance of the different flows properly, we use the same base encoder and decoder architecture for all models. We use gated convolutions and transposed convolutions as base layers for the encoder and decoder architecture respectively. The inference network consists of several gated convolution layers that produce a hidden unit vector. After being flattened, these hidden units act as an input to two fully connected layers that predict the mean and variance of z0\mathbf{z}^{0}.

Figure 3 shows the dependence of the negative evidence lower bound (or free energy) on the number of flows and the type of flow for static MNIST. The exact numbers corresponding to the figure are shown in Section B in the appendix.

For all models the performance improves as a functions of the number of flows. For 4 flows the difference between the baseline VAE and planar flows is very small. However, planar flows clearly benefit from more flow transformations.

For IAF three different widths of the MADE layers were used: C=320C=320, 640 and 1280. Surprisingly, for 4 flows the widest IAF with 1280 hidden units is outperformed by an IAF with 640 hidden units in the MADE layers. We expect this to be due to the fact that this model has more parameters and can therefore be harder to train, as indicated by the larger standard deviation for this model.

All three Sylvester flows outperform IAF and planar flows. For Orthogonal Sylvester flows, we show results for M=16M=16 and M=32M=32 orthogonal vectors per orthogonal matrix, thus corresponding to bottlenecks of size 16 and 32 respectively for a latent space of size D=64D=64. Clearly, a larger bottleneck improves performance. For Householder Sylvester flows we experimented with H=4H=4 and H=8H=8 Householder reflections per orthogonal matrix. Since the results were nearly indistinguishable between these two variants, we have left out the curve for H=4H=4 to avoid clutter. O-SNF with M=32M=32, H-SNF and T-SNF seem to perform on par.

In Table 1, the negative evidence lower bound and the estimated negative log-likelihood are shown for the baseline VAE, together with all flow models for 16 flows. The reported result for IAF is for a MADE width of 1280. The O-SNF model has a bottleneck of M=32M=32, and H-SNF contains 8 Householder reflections per orthogonal matrix. Again, all Sylvester flows outperform planar flows and IAF, both in terms of the free energy and the negative log-likelihood.

As discussed in Section 4, T-SNF is closely related to mean-only IAF, but with the MADE parameters produced by a hypernetwork that depends on the input data x\mathbf{x}. The fact that T-SNF outperforms IAF indicates that having data-dependent flow parameters directly leads to a more flexible transformation compared to taking a very wide MADE with a data-dependent context as an additional input.

2 FREYFACES, OMNIGLOT AND CALTECH 101 SILHOUETTES

We further assess the performance of the different models on Freyfaces, Omniglot and Caltech 101 Silhouettes. The results are shown in Table 2. The model settings are the sameFor Caltech 101 Silhouettes we used 2000 importance samples for the estimation of the negative log-likelihood. as those used for Table 1.

Freyfaces is a very small dataset of around 2000 faces. All normalizing flows increase the performance, with planar flows yielding the best result, closely followed by Triangular and Householder Sylvester flows. We expect planar flows to perform the best in this case since it is the least sensitive to overfitting.

For Omniglot and Caltech 101 Silhouettes the results are clearer, with the Sylvester normalizing flows family resulting in the best performance. Both H-SNF and T-SNF perform better than O-SNF. This could be attributed to the fact that O-SNF has a bottleneck of M=32M=32 for a latent space size of D=64D=64. The IAF scores for Caltech 101 are surprisingly bad. We expect this could be the case due to the large number of parameters that need to be trained for IAF(1280). Therefore we also evaluated the result for MADEs of width 320 for 16 flows. The resulting free energy and estimated negative log-likelihood are 111.23±0.45111.23\pm 0.45 and 99.74±0.2899.74\pm 0.28 respectively, only slightly improving on the results of 1280 wide IAFs.

CONCLUSION

We present a new family of normalizing flows: Sylvester normalizing flows. These flows generalize planar flows, while maintaining an efficiently computable Jacobian determinant through the use of Sylvester’s determinant identity. We ensure invertibility of the flows through the use of orthogonal and triangular parameter matrices. Three variants of Sylvester flows are investigated. First, orthogonal Sylvester flows use an iterative procedure to maintain orthogonality of parameter matrices. Second, Householder Sylvester flows use Householder reflections to construct orthogonal matrices. Third, triangular Sylvester flows alternate between fixed permutation and identity matrices for the orthogonal matrices. We show that the triangular Sylvester flows are closely related to mean-only IAF, with data-dependent MADE parameters. While performing comparably with planar flows and IAF for the Freyfaces dataset, our proposed family of flows improve significantly upon planar flows and IAF on the three other datasets.

We would like to thank Christos Louizos for helping with the implementation of inverse autoregressive flows, and Diederik Kingma for fruitful discussions. LH is funded by the UK EPSRC OxWaSP CDT through grant EP/L016710/1. JMT is funded by the European Commission within the MSC-IF (Grant No. 702666). RvdB is funded by SAP SE.

References

Appendix A Architecture

In the experiments we used convolutional layers for both the encoder and the decoder. Moreover, we used the gated activation function for convolutional layers:

where hl−1\mathbf{h}_{l-1} and hl\mathbf{h}_{l} are inputs and outputs of the ll-th layer, respectively, Wl,Vl\mathbf{W}_{l},\mathbf{V}_{l} are weights of the ll-th layer, bl,cl\mathbf{b}_{l},\mathbf{c}_{l} denote biases, ∗* is the convolution operator, and σ(⋅)\sigma(\cdot) is the sigmoid activation function.

Notice the last layer acts as a fully-connected layer. Eventually, fully-connected linear layers were used to parameterized diagonal Gaussian distribution and amortized parameters of a flow.

In the experimetns we used the following four image datasets: static MNISThttp://yann.lecun.com/exdb/mnist/, OMNIGLOThttps://github.com/yburda/iwae/blob/master/datasets/OMNIGLOT/chardata.mat., Caltech 101 Silhouetteshttps://people.cs.umass.edu/~marlin/data/caltech101_silhouettes_28_split1.mat., and Frey Faceshttp://www.cs.nyu.edu/~roweis/data/frey_rawface.mat. Frey Faces contains images of size 28×2028\times 20 and all other datasets contain 28×2828\times 28 images.

MNIST consists of hand-written digits split into 60,000 training datapoints and 10,000 test sample points. In order to perform model selection we put aside 10,000 images from the training set.

OMNIGLOT is a dataset containing 1,623 hand-written characters from 50 various alphabets. Each character is represented by about 20 images that makes the problem very challenging. The dataset is split into 24,345 training datapoints and 8,070 test images. We randomly pick 1,345 training examples for validation. During training we applied dynamic binarization of data similarly to dynamic MNIST.

Caltech 101 Silhouettes contains images representing silhouettes of 101 object classes. Each image is a filled, black polygon of an object on a white background. There are 4,100 training images, 2,264 validation datapoints and 2,307 test examples. The dataset is characterized by a small training sample size and many classes that makes the learning problem ambitious.

Frey Faces is a dataset of faces of a one person with different emotional expressions. The dataset consists of nearly 2,000 gray-scaled images. We randomly split them into 1,565 training images, 200200 validation images and 200200 test images. We repeated the experiment 33 times.

Appendix B MNIST experiments

The exact numbers for the evidence lower bound as shown in Fig. 3 are listed in Table 3.