Learning Hierarchical Features from Generative Models

Shengjia Zhao, Jiaming Song, Stefano Ermon

Introduction

A key property of deep feed-forward networks is that they tend to learn learn increasingly abstract and invariant representations at higher levels in the hierarchy (Bengio, 2009; Zeiler & Fergus, 2014) In the context of image data, low levels may learn features corresponding to edges or basic shapes, while higher levels learn more abstract features, such as object detectors (Zeiler & Fergus, 2014).

Generative models with a hierarchical structure, where there are multiple layers of latent variables, have been less successful compared to their supervised counterparts (Sønderby et al., 2016). In fact, the most successful generative models often use only a single layer of latent variables (Radford et al., 2015; van den Oord et al., 2016), and those that use multiple layers only show modest performance increases in quantitative metrics such as log-likelihood (Sønderby et al., 2016; Bachman, 2016). Because of the difficulties in evaluating generative models (Theis et al., 2015), and the fact that adding network layers increases the number of parameters, it is not always clear whether the improvements truly come from the choice of a hierarchical architecture. Furthermore, the capability of learning a hierarchy of increasingly complex and abstract features has only been demonstrated to a limited extent, with feature hierarchies that are not nearly as rich as the ones learned by feed-forward networks (Gulrajani et al., 2016).

Part of the problem is inherent and unavoidable for any generative model. The heart of the matter is that while highly invariant and local features are often sufficient for classification, generative modeling requires preservation of details (as illustrated in Figure 1). In fact, most latent features in a generative model of images cannot even demonstrate scale and translation invariance. The size and location of a sub-part often has to be dependent on the other sub-parts. For example, an eye should only be generated with the same size as the other eye, at symmetric locations with respect to the center of the face, with appropriate distance between them. The inductive biases that are directly encoded into the architecture of convolutional networks is not sufficient in the context of generative models.

On the other hand, other problems are associated with specific models or design choices, and may be avoided with deeper understanding and careful design. The goal of this paper is to provide a deeper understanding of the design and performance of common hierarchical latent variable models. We focus on variational models, though most of the conclusions can be generalized to adversarially trained models that support inference (Dumoulin et al., 2016; Donahue et al., 2016). In particular, we study two classes of models with a hierarchical structure:

1) Stacked hierarchy: The first type we study is characterized by recursively stacking generative models on top of each other. Most existing models (Sønderby et al., 2016; Gulrajani et al., 2016; Bachman, 2016; Kingma et al., 2016), belong to this class. We show that these models have two limitations. First, we show that if these models can be trained to optimality, then the bottom layer alone contains enough information to reconstruct the data distribution, and the layers above the first one can be ignored. This result holds under fairly general conditions, and does not depend on the specific family of distributions used to define the hierarchy (e.g., Gaussian). Second, we argue that many of the building blocks commonly used to construct hierarchical generative models are unlikely to help us learn disentangled features.

2) Architectural hierarchy: Motivated by these limitations, we turn our attention to single layer latent variable models. We propose an alternative way to learn disentangled hierarchical features by crafting a network architecture that prefers to place high-level features on certain parts of the latent code, and low-level features in others. We show that this approach, called Variational Ladder Autoencoder, allows us to learn very rich feature hierarchies on natural image datasets such as MNIST, SVHN (Netzer et al., 2011) and CelebA (Liu et al., 2015); in contrast, generative models with a stacked hierarchical structure fail to learn such features.

Problem Setting

We consider a family of latent variable models specified by a joint probability distribution pθ(x,z)p_{\theta}({{\bf x}},{{\bf z}}) over a set of observed variables x{{\bf x}} and latent variables z{{\bf z}}. The family of models is assumed to be parametrized by θ\theta. Let pθ(x)p_{\theta}({{\bf x}}) denote the marginal distribution of x{{\bf x}}. We wish to maximize the marginal log-likelihood p(x)p({{\bf x}}) over a dataset X={x(1),…,x(N)}{{\bf X}}=\{{{\bf x}}^{(1)},\ldots,{{\bf x}}^{(N)}\} drawn from some unknown underlying distribution pdata(x)p_{data}({{\bf x}}). Formally we would like to maximize

which is non-convex and often intractable for complex generative models, as it involves marginalization over the latent variables z{{\bf z}}.

We are especially interested in unsupervised feature learning applications, where by maximizing (1) we hope to discover a meaningful representation for the data x{{\bf x}} in terms of latent features given by pθ(z∣x)p_{\theta}({{\bf z}}|{{\bf x}}).

A popular solution (Kingma & Welling, 2013; Jimenez Rezende et al., 2014) for optimizing the intractable marginal likelihood (1) is to optimize the evidence lower bound (ELBO) by introducing an inference model qϕ(z∣x)q_{\phi}({{\bf z}}\lvert{{\bf x}}) parametrized by ϕ\phi We omit the dependency on θ\theta and ϕ\phi for the remainder of the paper.:

2 Hierarchical Variational Autoencoders

A hierarchical VAE (HVAE) can be thought of as a series of VAEs stacked on top of each other. It has the following hierarchy of latent variables z={z1,…,zL}{{\bf z}}=\{{{\bf z}}_{1},\ldots,{{\bf z}}_{L}\}, in addition to the observed variables x{{\bf x}}. We use the notation convention that z1{{\bf z}}_{1} represents the lowest layer closest to x{{\bf x}} and zL{{\bf z}}_{L} the top layer. Using chain rule, the joint distribution p(x,z1,…,zL)p({{\bf x}},{{\bf z}}_{1},\ldots,{{\bf z}}_{L}) can be factored as follows

Several models assume a Markov independence structure on the hidden variables, leading to the following simpler factorization (Jimenez Rezende et al., 2014; Gulrajani et al., 2016; Kaae Sønderby et al., 2016)

We refer to this common but more restrictive formulation as Markov HVAE.

For the inference distribution q(z∣x)q({{\bf z}}|{{\bf x}}) we do not assume any factorized structure to account for complex inference techniques used in recent work (Kaae Sønderby et al., 2016; Bachman, 2016). We also denote q(x,z)=pdata(x)q(z∣x)q({{\bf x}},{{\bf z}})=p_{data}({{\bf x}})q({{\bf z}}|{{\bf x}}).

Both p(x∣z)p({{\bf x}}|{{\bf z}}) and q(z∣x)q({{\bf z}}|{{\bf x}}) are jointly optimized, as before in Equation (2), to maximize the ELBO objective

where we define z0≡x{{\bf z}}_{0}\equiv{{\bf x}}, zL+1≡0{{\bf z}}_{L+1}\equiv{\bf 0}, and HH the entropy of a distribution, and expectation over pdata(x)p_{data}(x) is estimated by the samples in the dataset. This can be interpreted as stacking VAEs on top of each other.

Limitations of Hierarchical VAEs

One of the main reasons deep hierarchical networks are widely used as function approximators is their representational power. It is well known that certain functions can be represented much more compactly with deep networks, requiring exponentially less parameters compared to shallow networks (Bengio et al., 2009). However, we show that under ideal optimization of LELBO\mathcal{L}_{ELBO}, HVAE models do not lead to improved representational power. This is because for a well trained HVAE, a Gibbs chain on the bottom layer, which is a single layer model, can be used to recover pdata(x)p_{data}({{\bf x}}) exactly.

We first show this formally for Markov HVAE with the following proposition

LELBO\mathcal{L}_{ELBO} in Eq.(5) is globally maximized as a function of q(z∣x)q({{\bf z}}|{{\bf x}}) and p(x∣z)p({{\bf x}}|{{\bf z}}) when LELBO=−H(pdata(x))\mathcal{L}_{ELBO}=-H(p_{data}(x)). If LELBO\mathcal{L}_{ELBO} is globally maximized for a Markov HVAE, the following Gibbs sampling chain converges to pdata(x)p_{data}({{\bf x}}) if it is ergodic

By non-negativity of KL-divergence, and the fact that KL divergence is zero if an only if the two distributions are identical, it can be seen that this is uniquely optimized when p(x)=∫zp(x,z)dz=pdata(x)p({{\bf x}})=\int_{{\bf z}}p({{\bf x}},{{\bf z}})d{{\bf z}}=p_{data}({{\bf x}}) and ∀x,q(z∣x)=p(z∣x)\forall{{\bf x}},q({{\bf z}}|{{\bf x}})=p({{\bf z}}|{{\bf x}}) and the optimum is

This also implies that ∀x\forall{{\bf x}}

Because the following Gibbs chain converges to pdata(x)p_{data}({{\bf x}}) when it is ergodic

We can replace q(x∣z1(t))q({{\bf x}}|{{\bf z}}_{1}^{(t)}) with p(x∣z1(t))p({{\bf x}}|{{\bf z}}_{1}^{(t)}) using (7) and the chain still converges to pdata(x)p_{data}({{\bf x}}). ∎

Therefore under the assumptions of Proposition 1 we can sample from pdata(x)p_{data}({{\bf x}}) without using the latent code (z2,⋯ ,zL)({{\bf z}}_{2},\cdots,{{\bf z}}_{L}) at all. Hence, optimization of the LELBO\mathcal{L}_{ELBO} objective and efficient representation are conflicting, in the sense that optimality implies some level of redundancy in the representation.

2 Feature learning

Another significant advantage of hierarchical models for supervised learning is that they learn rich and disentangled hierarchies of features. This has been demonstrated for example using various visualization techniques (Zeiler & Fergus, 2014). However, we show in this section that typical HVAEs do not enjoy this property.

Suppose we train LELBO\mathcal{L}_{ELBO} in Equation (5) to optimality, we would have

Variational Ladder Autoencoders

Given the limitations of hierarchical architectures described in the previous section, we focus on an alternative approach to learn a hierarchy of disentangled features.

Our approach is to define a simple distribution with no hierarchical structure over the latent variables p(z)=p(z1,⋯ ,zL)p({{\bf z}})=p({{\bf z}}_{1},\cdots,{{\bf z}}_{L}). For example, the joint distribution p(z)p({{\bf z}}) can be a white Gaussian. Instead we encourage the latent code z1,⋯ ,zL{{\bf z}}_{1},\cdots,{{\bf z}}_{L} to learn features with different levels of abstraction by carefully choosing the mappings p(x∣z)p({{\bf x}}|{{\bf z}}) and q(z∣x)q({{\bf z}}|{{\bf x}}) between input x{{\bf x}} and latent code z{{\bf z}}. Our approach is based on the following intuition:

Assumption: If zi{{\bf z}}_{i} is more abstract than zj{{\bf z}}_{j}, then the inference mapping q(zi∣x)q({{\bf z}}_{i}|{{\bf x}}) and generative mapping when other layers are fixed p(x∣zi,z¬i=z¬i0)p({{\bf x}}|{{\bf z}}_{i},{{\bf z}}_{\neg i}=z^{0}_{\neg_{i}}) requires a more expressive network to capture.

This informal assumption suggests that we should use neural networks of different level of expressiveness to generate the corresponding features; the more abstract features require more expressive networks, and vice versa. We loosely quantify expressiveness with depth of the network. Based on these assumptions we are able to design an architecture that disentangles hierarchical features for many natural image datasets.

We decompose the latent code into subparts z={z1,z2,…}{{\bf z}}=\{{{\bf z}}_{1},{{\bf z}}_{2},\ldots\}, where z1{{\bf z}}_{1} relates to x{{\bf x}} with a shallow network, and increase network depth up to zL{{\bf z}}_{L}, which relates to x{{\bf x}} with a deep network. In particular, we share parameters with a ladder-like architecture (Valpola, 2015; Pezeshki et al., 2015). Because of this similarity we denote this architecture as Variational Ladder Autoencoder (VLAE). Formally, our model, shown in Figure 4 is defined as follows

1) Generative Network: p(z)=p(z1,⋯ ,zL)p({{\bf z}})=p({{\bf z}}_{1},\cdots,{{\bf z}}_{L}) is a simple prior on all latent variables. We choose it as a standard Gaussian N(0,I)\mathcal{N}(0,I). The conditional distribution p(x∣z1,z2,…,zL)p({{\bf x}}|{{\bf z}}_{1},{{\bf z}}_{2},\ldots,{{\bf z}}_{L}) is defined implicitly as:

2) Inference Network: For the inference network, we choose q(z∣x)q({{\bf z}}\lvert{{\bf x}}) as

3) Learning: For learning we use the ELBO criteria as in Equ.(2):

where p(z)=N(0,I)p({{\bf z}})=\mathcal{N}(0,{{\bf I}}) denotes the prior for z{{\bf z}}. This is tractable if r{{\bf r}} has tractable log likelihood, i.e. when r{{\bf r}} is a Gaussian.

This is essentially the inference and learning framework for a one-layer VAE; the hierarchy is only implicitly defined by the network architecture, therefore we call this model flat hierarchy. Motivated by our earlier theoretical results, we do not use additional layers of latent variables.

2 Comparison with Ladder Variational Autoencoders

Our architecture resembles the ladder variational autoencoder (LVAE) (Sønderby et al., 2016). However the two models are very different. The purpose of our architecture is to connect subparts of the latent code with networks of different expressive power (depth); the model is encouraged to place high-level, complex features at the top, and low-level, simple features at the bottom, in order to reach lower reconstruction error with latent codes of the same capacity. Empirically, this allows the network to learn disentangled factors of variation, corresponding to different subparts of the latent code. Meanwhile, because it is essentially a single-layer flat model, our VLAE does not exhibit the problems we have identified with traditional hierarchical VAE described in Section 3.

Ladder Variational Autoencoders (LVAE) on the other hand, utilize the ladder architecture from the inference/encoding side; its generative model is a standard HVAE. While the ladder inference network performs better than the one used in the original HVAE, ladder variational autoencoders still suffer from the problems we discussed in Section 3. The difference is between our model (VLAE) and LVAE is illustrated in Figure 4

Experiments

We train VLAE over several datasets and visualize the semantic meaning of the latent code. Code is available at https://github.com/ShengjiaZhao/Variational-Ladder-Autoencoder According to our assumptions, complex, high-level information will be learned by latent codes at higher layers, whereas simple, low-level features will be represented by lower layers.

In Figure 5, we visualize generation results from MNIST, where the model is a 3-layer VLAE with 2 dimensional latent code (z{{\bf z}}) at each layer. The visualizations are generated by systematically exploring the 2D latent code for one layer, while randomly sampling other layers. From the visualization, we see that the three layers encode stroke width, digit width and tilt and digit identity respectively. Remarkably, the semantic meaning of a particular latent code is stable with respect to the sampled latent codes from other layers. For example, in the second layer, the left side represents narrow digits whereas the right side represents wide digits. Sampling latent codes at other layers will control the digit identity, but have no influence over the width. This is interesting given that width is actually correlated with the digit identity; for example, digit 1 is typically thin while digit 0 is mostly wide. Therefore, the model will generate more zeros than ones if the latent code at the second layer corresponds to a wide digit, as displayed in the visualization.

Next we evaluate VLAE on the Street View House Number (SVHN, Netzer et al. (2011)) dataset, where it is significantly more challenging to learn interpretable representations since it is relatively noisy, containing certain digits which do not appear in the center. However, as is shown in Figure 6, our model is able to learn highly disentangled features through a 4-layer ladder, which includes color, digit shape, digit context, and general structure. These features are highly disentangled: since the latent code at the bottom layer controls color, modifying the code from other three layers while keeping the bottom layer fixed will generate a set of image which have the same tone in general. Moreover, the latent code learned at the top layer is the most complex one, which captures rich variations lower layers cannot accurately represent.

Finally, we display compelling results from another challenging dataset, CelebA (Liu et al., 2015), which includes 200,000 celebrity images. These images are highly varied in terms of environment and facial expressions. We visualize the generation results in Figure 7. As in the SVHN model, the latent code at the bottom layer learns the ambient color of the environment while keeping the personal details intact. Controlling other latent codes will change the other details of the individual, such as skin color, hair color, identity, pose (azimuth); more complicated features are placed at higher levels of the hierarchy.

Discussions

Training hierarchical deep generative models is a very challenging task, and there are two main successful families of methods. One family defines the destruction and reconstruction of data using a pre-defined process. Among them, LapGANs (Denton et al., 2015) define the process as repeatedly downsampling, and Diffusion Nets (Sohl-Dickstein et al., 2015) defines a forward Markov chain that coverts a complex data distribution to a simple, tractable one. Without having to perform inference, this makes training much easier, but it does not provide latent variables for other downstream tasks (unsupervised learning).

Another line of work focuses on learning a hierarchy of latent variables by stacking single layer models on top of each other. Many models also use more flexible inference techniques to improve performance (Sønderby et al., 2016; Dinh et al., 2014; Salimans et al., 2015; Rezende & Mohamed, 2015; Li et al., 2016; Kingma et al., 2016). However we show that there are limitations to stacked VAEs.

Our work distinguishes itself from prior work by explicitly discussing the purpose of learning such models: the advantage of learning a hierarchy is not in better representation efficiency, or better samples, but rather in the introduction of structure in the features, such as hierarchy or disentanglement. This motivates our method, VLAE, which justifies our intuition that a reasonable network structure can be, by itself, highly effective at learning structured (disentangled) representations. Contrary to previous efforts on hierarchical models, we do not stack VAEs on top of each other, instead we use a “flat” approach. This can be applied in combination with the stacking approach.

The results displayed in the experiments resemble those obtained with InfoGAN (Chen et al., 2016); both frameworks learn disentangled representations from the data in an unsupervised manner. The InfoGAN objective, however, explicitly maximizes the mutual information between the latent variables and the observation; whereas in VLAE, this is achieved through the reconstruction error objective which encourages the use of latent codes. Furthermore we are able to explicitly disentangle features with different level of abstractness.

Conclusions

In this paper, we discussed the potential practical value of learning a hierarchical generative model over a non-hierarchical one. We show that little can be gained in terms of representation efficiency or sample quality. We further show that traditional HVAE models have trouble learning structured features. Based on these insights, we consider an alternative to learning structured features by leveraging the expressive power of a neural network. Empirical results show that we can learn highly disentangled features.

One limitation of VLAE is the inability to learn structures other than hierarchical disentanglement. Future work should consider more principled ways of designing architectures that allow for learning features with more complex structures.

Acknowledgement

This research was supported by Intel Corporation, NSF (#1649208) and Future of Life Institute (#2016-158687).

References

Appendix A Additional Results

LELBO\mathcal{L}_{ELBO} for HVAE in Eq.(5) is optimized when LELBO=−H(pdata(x))\mathcal{L}_{ELBO}=-H(p_{data}(x)). If LELBO\mathcal{L}_{ELBO} is optimized the following Gibbs sampling chain converges to pdata(x)p_{data}({{\bf x}}) if it is ergodic

As in the proof of Proposition 1 when LELBO\mathcal{L}_{ELBO} is optimized, q(z∣x)=p(z∣x)q({{\bf z}}|{{\bf x}})=p({{\bf z}}|{{\bf x}}). Because the following Gibbs chain converges to pdata(x)p_{data}({{\bf x}})

We can replace q(x∣z(t))q({{\bf x}}|{{\bf z}}^{(t)}) with p(x∣z(t))p({{\bf x}}|{{\bf z}}^{(t)}) and the chain still converges to pdata(x)p_{data}({{\bf x}}). ∎

Appendix B Experimental Details

where W1,W2W_{1},W_{2} are trainable linear transformation matrices, and sigmsigm is sigmoid activation function. fl{{\bf f}}_{l} is a two layer dense network. For l=0l=0, we let

where σ\sigma is a hyper-parameter that can be specified apriori or trained. f0{{\bf f}}_{0} is a two layer convolutional network with 1/21/2 stride for spatial up-sampling. For inference we use the same architecture as the generator.

Learning: During training we use the Adam (Kingma & Ba, 2014) optimizer with learning rate 10−410^{-4}. We also anneal the scale the KL-regularization from to 11 to encourage use of latent feature during early stages of training.

B.2 VLAE

For VLAE, we use varying layers of convolution depending on size of input image. However, for the ladder connections we do not use convolution. Because of our argument in introduction and Figure 1, generative models do not benefit from convolutional latent features. Therefore we always flatten convolutional layers and apply linear transformation to reduce dimension for each ladder connection. For implementation details please refer to our code.