Information Dropout: Learning Optimal Representations Through Noisy Computation

Alessandro Achille, Stefano Soatto

Introduction

We call “representation” any function of the data that is useful for a task. An optimal representation is most useful (sufficient), parsimonious (minimal), and minimally affected by nuisance factors (invariant). Do deep neural networks approximate such sufficient invariants?

The cross-entropy loss most commonly used in deep learning does indeed enforce the creation of sufficient representations, but the other defining properties of optimal representations do not seem to be explicitly enforced by the commonly used training procedures. However, we show that this can be done by adding a regularizer, which is related to the injection of multiplicative noise in the activations, with the surprising result that noisy computation facilitates the approximation of optimal representations. In this paper we establish connections between the theory of optimal representations for classification tasks, variational inference, dropout and disentangling in deep neural networks. Our contributions can be summarized in the following steps:

We define optimal representations using established principles of statistical decision and information theory: sufficiency, minimality, invariance (cf. ) (Section 3).

We relate the defining properties of optimal representations for classification to the loss function most commonly used in deep learning, but with an added regularizer (Section 4, eq. 3).

We show that, counter-intuitively, injecting multiplicative noise to the computation improves the properties of a representation and results in better approximation of an optimal one (Section 6).

We relate such a multiplicative noise to the regularizer, and show that in the special case of Bernoulli noise, regularization reduces to dropout , thus establishing a connection to information theoretic principles. We also provide a more efficient alternative, called Information Dropout, that makes better use of limited capacity, adapts to the data, and is related to Variational Dropout (Section 6).

We show that, when the task is reconstruction, the procedure above yields a generalization of the Variational Autoencoder, which is instead derived from a Bayesian inference perspective . This establishes a connection between information theoretic and Bayesian representations, where the former explains the use of a multiplier used in practice but unexplained by Bayesian theory (Section 7).

We show that “disentanglement of the hidden causes,” an often-cited but seldom formalized desideratum for deep networks, can be achieved by assuming a factorized prior for the components of the optimal representation. Specifically, we prove that computing the regularizer term under the simplifying assumption of an independent prior has the effect of minimizing the total correlation of the components, a phenomenon previously observed empirically by (Section 5).

We validate the theory with several experiments including: improved insensitivity/invariance to nuisance factors using Information Dropout using (a) Cluttered MNIST and (b) MNIST+CIFAR, a newly introduced dataset to test sensitivity to occlusion phenomena critical in Vision applications; (c) we show improved efficiency of Information Dropout compared to regular dropout for limited capacity networks, (d) we show that Information Dropout favors disentangled representations; (e) we show that Information Dropout adapts to the data and allows different amounts of information to flow between different layers in a deep network (Section 8).

In the next section we introduce the basic formalism to make the above statements more precise, which we do in subsequent sections.

Preliminaries

In the general supervised setting, we want to learn the conditional distribution p(y∣x)p(\mathbf{y}|\mathbf{x}) of some random variable y\mathbf{y}, which we refer to as the task, given (samples of the) input data x\mathbf{x}. In typical applications, x\mathbf{x} is often high dimensional (for example an image or a video), while y\mathbf{y} is low dimensional, such as a label or a coarsely-quantized location. In such cases, a large part of the variability in x\mathbf{x} is actually due to nuisance factors that affect the data, but are otherwise irrelevant for the task . Since by definition these nuisance factors are not predictive of the task, they should be disregarded during the inference process. However, it often happens that modern machine learning algorithms, in part due to their high flexibility, will fit spurious correlations, present in the training data, between the nuisances and the task, thus leading to poor generalization performance.

In view of this, argue that the success of deep learning is in part due to the capability of neural networks to build incrementally better representations that expose the relevant variability, while at the same time discarding nuisances. This interpretation is intriguing, as it establishes a connection between machine learning, probabilistic inference, and information theory. However, common training practice does not seem to stem from this insight, and indeed deep networks may maintain even in the top layers dependencies on easily ignorable nuisances (see for example Figure 2).

To bring the practice in line with the theory, and to better understand these connections, we introduce a modified cost function, that can be seen as an approximation of the Information Bottleneck Lagrangian of , which encourages the creation of representations of the data which are increasingly disentangled and insensitive to the action of nuisances, and we show that this loss can be minimized using a new layer, which we call Information Dropout, that allows the network to selectively introduce multiplicative noise in the layer activations, and thus to control the flow of information. As we show in various experiments, this method improves the generalization performance by building better representations and preventing overfitting, and it considerably improves over binary dropout on smaller models, since, unlike dropout, Information Dropout also adapts the noise to the structure of the network and to the individual sample at test time.

Apart from the practical interest of Information Dropout, one of our main results is that Information Dropout can be seen as a generalization to several existing dropout methods, providing a unified framework to analyze them, together with some additional insights on empirical results. As we discuss in Section 3, the introduction of noise to prevent overfitting has already been studied from several points of view. For example the original formulation of dropout of , which introduces binary multiplicative noise, was motivated as a way of efficiently training an ensemble of exponentially many networks, that would be averaged at testing time. introduce Variational Dropout, a dropout method which closely resemble ours, and is instead derived from a Bayesian analysis of neural networks. Information Dropout gives an alternative information-theoretic interpretation of those methods.

As we show in Section 7, other than being very closely related to Variational Dropout, Information Dropout directly yields a variational autoencoder as a special case when the task is the reconstruction of the input. This result is in part expected, since our loss function seeks an optimal representation of the input for the task of reconstruction, and the representation given by the latent variables of a variational autoencoder fits the criteria. However, it still rises the question of exactly what and how deep are the links between information theory, representation learning, variational inference and nuisance invariance. This work can be seen as a small step in answering this question.

Related work

The main contribution of our work is to establish how two seemingly different areas of research, namely dropout methods to prevent overfitting, and the study of optimal representations, can be linked through the Information Bottleneck principle.

Dropout was introduced by Srivastava et al. . The original motivation was that by randomly dropping the activations during training, we can effectively train an ensemble of exponentially many networks, that are then averaged during testing, therefore reducing overfitting. Wang et al. suggested that dropout could be seen as performing a Monte-Carlo approximation of an implicit loss function, and that instead of multiplying the activations by binary noise, like in the original dropout, multiplicative Gaussian noise with mean 1 can be used as a way of better approximating the implicit loss function. This led to a comparable performance but faster training than binary dropout.

Kingma et al. take a similar view of dropout as introducing multiplicative (Gaussian) noise, but instead study the problem from a Bayesian point of view. In this setting, given a training dataset D={(xi,yi)}i=1,…,N\mathcal{D}=\left\{(\mathbf{x}_{i},\mathbf{y}_{i})\right\}_{i=1,\ldots,N} and a prior distribution p(w)p(\mathbf{w}), we want to compute the posterior distribution p(w∣D)p(\mathbf{w}|\mathcal{D}) of the weights w\mathbf{w} of the network. As is customary in variational inference, the true posterior can be approximated by minimizing the negative variational lower bound L(θ)\mathcal{L}(\theta) of the marginal log-likelihood of the data,

This minimization is difficult to perform, since it requires to repeatedly sample new weights for each sample of the dataset. As an alternative, suggest that the uncertainty about the weights that is expressed by the posterior distribution pθ(w∣D)p_{\theta}(\mathbf{w}|\mathcal{D}) can equivalently be encoded as a multiplicative noise in the activations of the layers (the so called local reparametrization trick). As we will see in the following sections, this loss function closely resemble the one of Information Dropout, which however is derived from a purely information theoretic argument based on the Information Bottleneck principle. One difference is that we allow the parameters of the noise to change on a per-sample basis (which, as we show in the experiments, can be useful to deal with nuisances), and that we allow a scaling constant β\beta in front of the KL-divergence term, which can be changed freely. Interestingly, even if the Bayesian derivation does not allow a rescaling of the KL-divergence, notice that choosing a different scale for the KL-divergence term can indeed lead to improvements in practice. A related method, but derived from an information theoretic perspective was also suggested previously by .

The interpretation of deep neural network as a way of creating successively better representations of the data has already been suggested and explored by many. Most recently, Tishby et al. put forth an interpretation of deep neural networks as creating sufficient representations of the data that are increasingly minimal. In parallel simultaneous work, approximate the information bottleneck similarly to us, but focus on empirical analysis of robustness to adversarial perturbations rather than tackling disentanglement, invariance and minimality analytically. Some have focused on creating representations that are maximally invariant to nuisances, especially when they have the structure of a (possibly infinite-dimensional) group acting on the data, like , or, when the nuisance is a locally compact group acting on each layer, by successive approximations implemented by hierarchical convolutional architectures, like and . In these cases, which cover common nuisances such as translations and rotations of an image (affine group), or small diffeomorphic deformations due to a slight change of point of view (group of diffeomorphisms), the representation is equivalent to the data modulo the action of the group. However, when the nuisances are not a group, as is the case for occlusions, it is not possible to achieve such equivalence, that is, there is a loss. To address this problem, defined optimal representations not in terms of maximality, but in terms of sufficiency, and characterized representations that are both sufficient and invariant. They argue that the management of nuisance factors common in visual data, such as changes of viewpoint, local deformations, and changes of illumination, is directly tied to the specific structure of deep convolutional networks, where local marginalization of simple nuisances at each layer results in marginalization of complex nuisances in the network as a whole.

Our work fits in this last line of thinking, where the goal is not equivalence to the data up to the action of (group) nuisances, but instead sufficiency for the task. Our main contribution in this sense is to show that injecting noise into the layers, and therefore using a non-deterministic function of the data, can actually simplify the theoretical analysis and lead to disentangling and improved insensitivity to nuisances. This is an alternate explanation to that put forth by the references above.

Optimal representations and the Information Bottleneck loss

Given some input data x\mathbf{x}, we want to compute some (possibly nondeterministic) function of x\mathbf{x}, called a representation, that has some desirable properties in view of the task y\mathbf{y}, for instance by being more convenient to work with, exposing relevant statistics, or being easier to store. Ideally, we want this representation to be as good as the original data for the task, and not squander resources modeling parts of the data that are irrelevant to the task. Formally, this means that we want to find a random variable z\mathbf{z} satisfying the following conditions:

z\mathbf{z} is a representation of x\mathbf{x}; that is, its distribution depends only on x\mathbf{x}, as expressed by the following Markov chain:

y<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mi>x</mi></mrow><annotationencoding="application/x−tex">x</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.4306em;"></span><spanclass="mordmathnormal">x</span></span></span></span></span>zy<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>x</mi></mrow><annotation encoding="application/x-tex">x</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">x</span></span></span></span></span>z (ii) z\mathbf{z} is sufficient for the task y\mathbf{y}, that is I(x;y)=I(z;y)I(\mathbf{x};\mathbf{y})=I(\mathbf{z};\mathbf{y}), expressed by the Markov chain:

y<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mi>z</mi></mrow><annotationencoding="application/x−tex">z</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.4306em;"></span><spanclass="mordmathnormal"style="margin−right:0.044em;">z</span></span></span></span></span>xy<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>z</mi></mrow><annotation encoding="application/x-tex">z</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.044em;">z</span></span></span></span></span>x (iii) among all random variables satisfying these requirements, the mutual information I(x;z)I(\mathbf{x};\mathbf{z}) is minimal. This means that z\mathbf{z} discards all variability in the data that is not relevant to the task.

Using the identity I(x;y)−I(z;y)=H(y∣z)−H(y∣x)I(\mathbf{x};\mathbf{y})-I(\mathbf{z};\mathbf{y})=H(\mathbf{y}|\mathbf{z})-H(\mathbf{y}|\mathbf{x}), where HH denotes the entropy and II the mutual information, it is easy to see that the above conditions are equivalent to finding a distribution p(z∣x)p(\mathbf{z}|\mathbf{x}) which solves the optimization problem

The minimization above is difficult in general. For this reason, Tishby et al. have introduced a generalization known as the Information Bottleneck Principle and the associated Lagrangian to be minimized

where β\beta is a positive constant that manages the trade-off between sufficiency (the performance on the task, as measured by the first term) and minimality (the complexity of the representation, measured by the second term). It is easy to see that, in the limit β→0+\beta\to 0^{+}, this is equivalent to the original problem, where z\mathbf{z} is a minimal sufficient statistic. When all random variables are discrete and z=T(x)\mathbf{z}=T(\mathbf{x}) is a deterministic function of x\mathbf{x}, the algorithm proposed by can be used to minimize the IB Lagrangian efficiently. However, no algorithm is known to minimize the IB Lagrangian for non-Gaussian, high-dimensional continuous random variables.

One of our key results is that, when we restrict to the family of distributions obtained by injecting noise to one layer of a neural network, we can efficiently approximate and minimize the IB Lagrangian.Since we restrict the family of distributions, there is no guarantee that the resulting representation will be optimal. We can, however, iterate the process to obtain incrementally improved approximations. As we will show, this process can be effectively implemented through a generalization of the dropout layer that we call Information Dropout.

To set the stage, we rewrite the IB Lagrangian as a per-sample loss function. Let p(x,y)p(\mathbf{x},\mathbf{y}) denote the true distribution of the data, from which the training set {(xi,yi)}i=1,…,N\left\{(\mathbf{x}_{i},\mathbf{y}_{i})\right\}_{i=1,\ldots,N} is sampled, and let pθ(z∣x)p_{\theta}(\mathbf{z}|\mathbf{x}) and pθ(y∣z)p_{\theta}(\mathbf{y}|\mathbf{z}) denote the unknown distributions that we wish to estimate, parametrized by θ\theta. Then, we can write the two terms in the IB Lagrangian as

where KL{\rm KL} denotes the Kullback-Leibler divergence. We can therefore approximate the IB Lagrangian empirically as

Notice that the first term simply is the average cross-entropy, which is the most commonly used loss function in deep learning. The second term can then be seen as a regularization term. In fact, many classical regularizers, like the L2L_{2} penalty, can be expressed in the form of eq. 3 (see also ). In this work, we interpret the KL term as a reuglarizer that penalizes the transfer of information from x\mathbf{x} to z\mathbf{z}. In the next section, we discuss ways to control such information transfer through the injection of noise.

Aside from being easier to work with, stochastic representations can attain a lower value of the IB Lagrangian than any deterministic representation. For example, consider the task of reconstructing single random bit yy given a noisy observation xx. The only deterministic representations are equivalent to the either the noisy observation itself or to the trivial constant map. It is not difficult to check that for opportune values of β\beta and of the noise, neither realize the optimal tradeoff reached by a suitable stochastic representation.

The quantity I(x;y∣z)=H(y∣z)−H(y∣x)≥0I(x;y|z)=H(y|z)-H(y|x)\geq 0 can be seen as a measure of the distance between p(x,y,z)p(x,y,z) and the closest distribution q(x,y,z)q(x,y,z) such that x→z→y\mathbf{x}\to\mathbf{z}\to\mathbf{y} is a Markov chain. Therefore, by minimizing eq. 2 we find representations that are increasingly “more sufficient”, meaning that they are closer to an actual Markov chain.

Disentanglement

In addition to sufficiency and minimality, “disentanglement of hidden factors” is often cited as a desirable property of a representation, but seldom formalized. We can quantify disentanglement by measuring the total correlation, or multivariate mutual information, defined as

Notice that the components of z\mathbf{z} are mutually independent if and only if TC⁡(z)\operatorname{TC}(z) is zero. Adding this as a penalty in the IB Lagrangian, with a factor γ\gamma yields

In general, minimizing this augmented loss is intractable, since to compute both the KL term and the total correlation, we need to know the marginal distribution pθ(z)p_{\theta}(\mathbf{z}), which is not easily computable. However, the following proposition, that we prove in Appendix B, shows that if we choose γ=β\gamma=\beta, then the problem simplifies, and can be easily solved by adding an auxiliary variable.

is equivalent to the following minimization in two variables

In other words, minimizing the standard IB Lagrangian assuming that the activations are independent, i.e. having q(z)=∏iqi(zi)q(\mathbf{z})=\prod_{i}q_{i}(\mathbf{z}_{i}), is equivalent to enforcing disentanglement of the hidden factors. It is interesting to note that this independence assumption is already adopted often by practitioners on grounds of simplicity, since the actual marginal p(z)=∫xp(x,z)dxp(\mathbf{z})=\int_{\mathbf{x}}p(\mathbf{x},\mathbf{z})d\mathbf{x} is often incomputable. That using a factorized model results in “disentanglement” was also observed empirically by which, however, introduced an ad-hoc metric based on classifiers of low VC-dimension, rather than the more natural Total Correlation adopted here.

In view of the previous proposition, from now on we will assume that the activations are independent and ignore the total correlation term.

Information Dropout

Guided by the analysis in the previous sections, and to emphasize the role of stochasticity, we consider representations z\mathbf{z} obtained by computing a deterministic map f(x)f(\mathbf{x}) of the data (for instance a sequence of convolutional and/or fully-connected layers of a neural network), and then multiplying the result component-wise by a random sample ϵ\epsilon drawn from a parametric noise distribution pαp_{\alpha} with unit mean and variance that depends on the input x\mathbf{x}:

where “⊙\odot” denotes the element-wise product. Notice that, if pα(x)(ε)p_{\alpha(\mathbf{x})}(\varepsilon) is a Bernoulli distribution rescaled to have mean 11, this reduces exactly to the classic binary dropout layer. As we discussed in Section 3, there are also variants of dropout that use different distributions.

x<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mi>f</mi><mostretchy="false">(</mo><mi>x</mi><mostretchy="false">)</mo></mrow><annotationencoding="application/x−tex">f(x)</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:1em;vertical−align:−0.25em;"></span><spanclass="mordmathnormal"style="margin−right:0.1076em;">f</span><spanclass="mopen">(</span><spanclass="mordmathnormal">x</span><spanclass="mclose">)</span></span></span></span></span>z=ε⊙f(x)<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mi>ε</mi></mrow><annotationencoding="application/x−tex">ε</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.4306em;"></span><spanclass="mordmathnormal">ε</span></span></span></span></span>yx<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>f</mi><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">f(x)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="mopen">(</span><span class="mord mathnormal">x</span><span class="mclose">)</span></span></span></span></span>\mathbf{z}=\varepsilon\odot f(\mathbf{x})<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>ε</mi></mrow><annotation encoding="application/x-tex">\varepsilon</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">ε</span></span></span></span></span>y A natural choice for the distribution pα(x)(ε)p_{\alpha(\mathbf{x})}(\varepsilon), which also simplifies the theoretical analysis, is the log-normal distribution pα(x)(ε)=log⁡N(0,αθ2(x))p_{\alpha(\mathbf{x})}(\varepsilon)=\log\mathcal{N}(0,\alpha_{\theta}^{2}(\mathbf{x})). Once we fix this noise distribution, given the above expression for z\mathbf{z}, we can easily compute the distribution pθ(z∣x)p_{\theta}(\mathbf{z}|\mathbf{x}) that appears in eq. 3. However, to be able to compute the KL-divergence term, we still need to fix a prior distribution qθ(z)q_{\theta}(\mathbf{z}). The choice of this prior largely depends on the expected distribution of the activations f(x)f(\mathbf{x}). Recall that, by Section 5, we can assume that all activations are independent, thus simplifying the computation. Now, we concentrate on two of the most common activation functions, the rectified linear unit (ReLU), which is easy to compute and works well in practice, and the Softplus function, which can be seen as a strictly positive and differentiable approximation of ReLU.

A network implemented using only ReLU and a final Softmax layer has the remarkable property of being scale-invariant, meaning that multiplying all weights, biases, and activations by a constant does not change the final result. Therefore, from a theoretical point of view, it would be desirable to use a scale-invariant prior. The only such prior is the improper log-uniform, q(log⁡(z))=cq(\log(z))=c, or equivalently q(z)=c/zq(z)=c/z, which was also suggested by , but as a prior for the weights of the network, rather than the activations. Since the ReLU activations are frequently zero, we also assume q(z=0)=q0q(z=0)=q_{0} for some constant 0≤q0≤10\leq q_{0}\leq 1. Therefore, the final prior has the form q(z)=q0δ0(z)+c/zq(z)=q_{0}\delta_{0}(z)+c/z, where δ0\delta_{0} is the Dirac delta in zero. In Figure 1(a), we compare this prior distribution with the actual empirical distribution p(z)p(z) of a network with ReLU activations.

In a network implemented using Softplus activations, a log-normal is a good fit of the distribution of the activations. This is to be expected, especially when using batch-normalization, since the pre-activations will approximately follow a normal distribution with zero mean, and the Softplus approximately resembles a scaled exponential near zero. Therefore, in this case we suggest using a log-normal distribution as our prior q(z)q(z). In Figure 1(b), we compare this prior with the empirical distribution p(z)p(z) of a network with Softplus activations.

Using these priors, we can finally compute the KL divergence term in eq. 3 for both ReLU activations and Softplus activations. We prove the following two propositions in Appendix A.

Let z=ε⋅f(x)z=\varepsilon\cdot f(x), where ε∼pα(ε)\varepsilon\sim p_{\alpha}(\varepsilon), and assume p(z)=qδ0(z)+c/zp(z)=q\delta_{0}(z)+c/z. Then, assuming f(x)≠0f(x)\neq 0, we have

In particular, if pα(ε)p_{\alpha}(\varepsilon) is chosen to be the log-normal distribution pα(ε)=log⁡N(0,αθ2(x))p_{\alpha}(\varepsilon)=\log\mathcal{N}(0,\alpha^{2}_{\theta}(x)), we have

Let z=ε⋅f(x)z=\varepsilon\cdot f(x), where ε∼pα(ε)=log⁡N(0,αθ2(x))\varepsilon\sim p_{\alpha}(\varepsilon)=\log\mathcal{N}(0,\alpha_{\theta}^{2}(x)), and assume pθ(z)=log⁡N(μ,σ2)p_{\theta}(z)=\log\mathcal{N}(\mu,\sigma^{2}). Then, we have

Substituting the expression for the KL divergence in eq. 4 inside eq. 3, and ignoring for simplicity the special case f(x)=0f(x)=0, we obtain the following loss function for ReLU activations

and a similar expression for Softplus. Notice that the first expectation can be approximated by sampling (in the experiments we use one single sample, as customary for dropout), and is just the average cross-entropy term that is typical in deep learning. The second term, which is new, penalizes the network for choosing a low variance for the noise, i.e. for letting more information pass through to the next layer. This loss can be optimized easily using stochastic gradient descent and the reparametrization trick of to back-propagate the gradient through the sampling operation.

Variational autoencoders and Information Dropout

In this section, we outline the connection between variational autoencoders and Information Dropout. A variational autoencoder (VAE) aims to reconstruct, given a training dataset D={xi}\mathcal{D}=\left\{\mathbf{x}_{i}\right\}, a latent random variable z\mathbf{z} such that the observed data x\mathbf{x} can be thought as being generated by the, usually simpler, variable z\mathbf{z} through some unknown generative process pθ(x∣z)p_{\theta}(\mathbf{x}|\mathbf{z}). In practice, this is done by minimizing the negative variational lower-bound to the marginal log-likelihood of the data

which can be optimized easily using the SGVB method of . Interestingly, when the task is reconstruction, that is when y=x\mathbf{y}=\mathbf{x}, the IB loss function in eq. 3 reduces to

Therefore, by letting β=1\beta=1 in the previous expression, we obtain exactly the loss function of a variational autoencoder, that is, the representation z\mathbf{z} computed by the Information Dropout layer coincides with the latent variable z\mathbf{z} computed by the VAE. This is in part to be expected, since the objective of Information Dropout is to create a representation of the data that is minimal sufficient for the task of reconstruction, and the latent variables of a VAE can be thought as such a representation. The term β\beta in this case can be seen as managing the trade off between the fidelity of the reconstruction of the input from the representation (measured by the cross-entropy), against the compression factor (complexity) of the representation (measured by the KL-divergence). As anticipated, Bayesian theory would prescribe β=1\beta=1, whereas it has been observed empirically that other choices can yield better performance. In the IB framework, the choice of β\beta is for the designer or model selection algorithm to choose.

Taking inspiration by experimental evidence in neuroscience, a contemporary work by Higgins et al. also suggests the use of the loss function in eq. (7) to train a VAE. They prove experimentally that for higher values of β\beta the resulting representation z\mathbf{z} is increasingly disentangled. This result is indeed compatible with our observation in Section 5, and in Section 8 we prove related experimental results in a more general situation.

Experiments

The goal of our experiments is to validate the theory, by showing that indeed increasing noise level yields reduced dependency on nuisance factors, a more disentangled representation, and that by adapting the noise level to the data we can better exploit architectures of limited capacity.

To this end, we first compare Information Dropout with the Dropout baseline on several standard benchmark datasets using different networks architecture, and highlight a few key properties. All the models were implemented using TensorFlow . As also notice, letting the variance of the noise grow excessively leads to poor generalization. To avoid this problem, we constraint α(x)<0.7\alpha(x)<0.7, so that the maximum variance of the log-normal error distribution will be approximately 1, the same as binary dropout when using a drop probability of 0.5. In all experiments we divide the KL-divergence term by the number of training samples, so that for β=1\beta=1 the scaling of the KL-divergence term in similar to the one used by Variational Dropout (see Section 3).

Cluttered MNIST. To visually asses the ability of Information Dropout to create a representation that is increasingly insensitive to nuisance factors, we train the All-CNN-96 network (Table II) for classification on a Cluttered MNIST dataset , consisting of 96×9696\times 96 images containing a single MNIST digit together with 21 distractors. The dataset is divided in 50,000 training images and 10,000 testing images. As shown in Figure 2, for small values of β\beta, the network lets through both the objects of interest (digits) and distractors, to upper layers. By increasing the value of β\beta, we force the network to disregard the least discriminative components of the data, thereby building a better representation for the task. This behavior depends on the ability of Information Dropout to learn the structure of the nuisances in the dataset which, unlike other methods, is facilitated by the ability to select noise level on a per-sample basis.

Occluded CIFAR. Occlusions are a fundamental phenomenon in vision, for which it is difficult to hand-design invariant representations. To assess that the approximate minimal sufficient representation produced by Information Dropout has this invariance property, we created a new dataset by occluding images from CIFAR-10 with digits from MNIST (Figure 4). We train the All-CNN-32 network (Table II) to classify the CIFAR image. The information relative to the occluding MNIST digit is then a nuisance for the task, and therefore should be excluded from the final representation. To test this, we train a secondary network to classify the nuisance MNIST digit using only the the representation learned for the main task. When training with small values of β\beta, the network has very little pressure to limit the effect of nuisances in the representation, so we expect the nuisance classifier to perform better. On the other hand, increasing the value of β\beta we expect its performance to degrade, since the representation will become increasingly minimal, and therefore invariant to nuisances. The results in Figure 4 confirm this intuition.

MNIST and CIFAR-10. Similar to , to see the effect of Information Dropout on different network sizes and architectures, we train on MNIST a network with 3 fully connected hidden layers with a variable number of hidden units, and we train on CIFAR-10 the All-CNN-32 convolutional network described in Table II, using a variable percentage of all the filters. The fully connected network was trained for 80 epochs, using stochastic gradient descent with momentum with initial learning rate 0.07 and dropping the learning rate by 0.1 at 30 and 70 epochs. The CNN was trained for 200 epochs with initial learning rate 0.05 and dropping the learning rate by 0.1 at 80, 120 and 160 epochs. We show the results in Figure 3. Information Dropout is comparable or outperforms binary dropout, especially on smaller networks. A possible explanation is that dropout severely reduces the already limited capacity of the network, while Information Dropout can adapt the amount of noise to the data and to the size of the network so that the relevant information can still flow to the successive layers. Figure 6 shows how the amount of transmitted information also adapts to the size and hierarchical level of the layer.

Disentangling. As we saw Section 6, in the case of Softplus activations, the logarithm of the activations approximately follow a normal distribution. We can then approximate the total correlation using the associated covariance matrix Σ\Sigma. Precisely, we have

where Σ0=diag⁡Σ\Sigma_{0}=\operatorname{diag}\Sigma is the variance of the marginal distribution. In Figure 5 we plot for different values of β\beta the testing error and the total correlation of the representation learned by All-CNN-32 on CIFAR-10 when using 25% of the filters. As predicted, when β\beta increases the total correlation diminishes, that is, the representation becomes disentangled, and the testing error improves, since we prevent overfitting. When β\beta is to large, information flow is insufficient, and the testing error rapidly increases.

VAE. To validate Section 7, we replicate the basic variational autoencoder of , implementing it both with Gaussian latent variables, as in the original, and with an Information Dropout layer. We trained both implementations for 300 epochs dropping the learning rate by 0.1 at 30 and 120 epochs. We report the results in the following table. The Information Dropout implementation has similar performance to the original, confirming that a variational autoencoder can be considered a special case of Information Dropout.

Discussion

We relate the Information Bottleneck principle and its associated Lagrangian to seemingly unrelated practices and concepts in deep learning, including dropout, disentanglement, variational autoencoding. For classification tasks, we show how an optimal representation can be achieved by injecting multiplicative noise in the activation functions, and therefore into the gradient computation during learning.

A special case of noise (Bernoulli) results in dropout, which is standard practice originally motivated by ensemble averaging rather than information-theoretic considerations. Better (adaptive) noise models result better exploitation of limited capacity, leading to a method we call Information Dropout. We also establish connections with variational inference and variational autoencoding, and show that “disentangling of the hidden causes” can be measured by total correlation and achieved simply by enforcing independence of the components in the representation prior.

So, what may be done by necessity in some computational systems (noisy computation), turns out to be beneficial towards achieving invariance and minimality. Analogously, what has been done for convenience (assuming a factorized prior) turns out to be the beneficial towards achieving “disentanglement.”

Another interpretation of Information Dropout is as a way of biasing the network towards reconstructing representations of the data that are compatible with a Markov chain generative model, making it more suited to data coming from hierarchical models, and in this sense is complementary to architectural constraint, such as convolutions, that instead bias the model toward geometric tasks.

It should be noticed that injecting multiplicative noise to the activations can be thought of as a particular choice of a class of minimizers of the loss function, but can also be interpreted as a regularization terms added to the cost function, or as a particular procedure utilized to carry out the optimization. So the same operation can be interpreted as either of the three key ingredients in the optimization: the function to be minimized, the family over which to minimize, and the procedure with which to minimize. This highlight the intimate interplay between the choice of models and algorithms in deep learning.

References

Appendix A Computations

Let z=ε⋅f(x)z=\varepsilon\cdot f(x), where ε∼pα(ε)\varepsilon\sim p_{\alpha}(\varepsilon), and assume p(z)=qδ0(z)+c/zp(z)=q\delta_{0}(z)+c/z. Then, assuming f(x)≠0f(x)\neq 0, we have

In particular, if pα(ε)p_{\alpha}(\varepsilon) is chosen to be the log-normal distribution pα(ε)=log⁡N(0,αθ2(x))p_{\alpha}(\varepsilon)=\log\mathcal{N}(0,\alpha^{2}_{\theta}(x)), we have

If f(x)≠0f(x)\neq 0, then we also have z≠0z\neq 0. Since the KL-divergence is invariant under parameter transformations we can write

For the second part, notice that by definition pα(x)=N(0,αθ2(x))p_{\alpha(x)}=\mathcal{N}(0,\alpha_{\theta}^{2}(x)) and

Finally, if f(x)=0f(x)=0, then also z=0z=0, so p(z∣x)=δ0(z)p(z|x)=\delta_{0}(z). It is then easy to see that

Let z=ε⋅f(x)z=\varepsilon\cdot f(x), where ε∼pα(ε)=log⁡N(0,αθ2(x))\varepsilon\sim p_{\alpha}(\varepsilon)=\log\mathcal{N}(0,\alpha_{\theta}^{2}(x)), and assume pθ(z)=log⁡N(μ,σ2)p_{\theta}(z)=\log\mathcal{N}(\mu,\sigma^{2}). Then, we have

Since the KL divergence is invariant for reparametrizations, the divergence between two log-normal distributions is equal to the divergence between the corresponding normal distributions. Therefore, using the known formula for the KL divergence of normals, we get the desired result. ∎

Appendix B Disentanglement

In this appendix, we show that the minimization problem

which is difficult in general since we do not have access to the joint distribution p(z)p(\mathbf{z}), is equivalent to the following simpler optimization problem in two variables

In the following proposition, for simplicity, we concentrate on discrete random variables.

Let z=(z1,…,zn)\mathbf{z}=(z_{1},\ldots,z_{n}) be a discrete random variable, let p(z∣x)p(\mathbf{z}|\mathbf{x}) be a generic probability distribution, and let q(z)=∏i=1nqi(zi)q(\mathbf{z})=\prod_{i=1^{n}}q_{i}(z_{i}) be a factorized prior distribution. Then, for any function F(p)F(p), a minimization problem in the form

where Ip(z;x)I_{p}(\mathbf{z};\mathbf{x}) is the mutual information and TC⁡p(z)\operatorname{TC}_{p}(\mathbf{z}) is the total correlation of z\mathbf{z}, assuming z∼p(z)\mathbf{z}\sim p(\mathbf{z}).

To prove the proposition, we just need to minimize with respect to qq and substitute back the solution. Adding a Lagrange multiplier for the constrain ∑ziqi(zi)=1\sum_{z_{i}}q_{i}(z_{i})=1, the problem can be rewritten as

Taking the derivative with respect to to pi(zˉi)p_{i}(\bar{z}_{i}) we have

Setting it to zero, we obtain p(zi)=q(zi)p(z_{i})=q(z_{i}), that is, the optimal factorized prior is the product of the marginals. Substituting it back in the second term (the only containing pp), we obtain

Appendix C Additional plots