Spherical Latent Spaces for Stable Variational Autoencoders

Jiacheng Xu, Greg Durrett

Introduction

Recent work has established the effectiveness of deep generative models for a range of tasks in NLP, including text generation Hu et al. 2017; Yu et al. 2017, machine translation Zhang et al. 2016, and style transfer Shen et al. 2017; Zhao et al. 2017a. Variational autoencoders, which have been explored in past work for text modeling Miao et al. 2016; Bowman et al. 2016, posit a continuous latent variable which is used to capture latent structure in the data. Typical VAE implementations assume the prior of this latent space is a multivariate Gaussian; during training, a Kullback-Leibler (KL) divergence term in loss function encourages the variational posterior to approximate the prior. One major limitation of this approach observed by past work is that the KL term may encourage the posterior distribution of the latent variable to “collapse” to the prior, effectively rendering the latent structure unused Bowman et al. 2016; Chen et al. 2016.

In this paper, we propose to use the von Mises-Fisher (vMF) distribution rather than Gaussian for our latent variable. vMF places a distribution over the unit hypersphere governed by a mean parameter μ\mu and a concentration parameter κ\kappa. Our prior is a uniform distribution over the unit hypersphere (κ=0\kappa=0) and our family of posterior distributions treats κ\kappa as a fixed model hyperparameter. Since the KL divergence only depends on κ\kappa, we can structurally prevent the KL collapse and make our model’s optimization problem easier. We show that this approach is actually more robust than trying to flexibly learn κ\kappa, and a wide range of settings for fixed κ\kappa lead to good performance. Our model systematically achieves better log likelihoods than analogous Gaussian models while having higher KL divergence values, showing that it more successfully makes use of the latent variables at the end of training.

Past work has suggested several other techniques for dealing with the KL collapse in the Gaussian case. Annealing the weight of KL term Bowman et al. 2016 still leaves us with brittleness in the optimization process, as we show in Section 2. Other prior work Yang et al. 2017; Semeniuta et al. 2017 focuses on using CNNs rather than RNNs as the decoder in order to weaken the model and encourage the use of the latent code, but the gains are limited and changing the decoder in this way requires ad hoc model engineering and careful tuning of various decoder capacity parameters. Our method is orthogonal to the choice of the decoder and can be combined with any of these approaches. Using vMF distributions in VAEs also leaves us the flexibility to modify the prior in other ways, such as using a product distribution with a uniform Guu et al. 2018 or piecewise constant term Serban et al. 2017a.

We evaluate our approach in two generative modeling paradigms. For both RNN language modeling and bag-of-words document modeling, we find that vMF is more robust than a Gaussian prior, and our model learns to rely more on the latent variable while achieving better held-out data likelihoods. To better understand the contrast between these models, we design and conduct a series of experiments to understand the properties of the Gaussian and vMF latent code spaces, which make different structural assumptions. Unsurprisingly, these latent code distributions capture much of the same information as in a bag of words, but we show that vMF can more readily go beyond this, capturing ordering information more effectively than a Gaussian code.

Variational Autoencoders for Text

Bowman et al. 2016 propose a variational autoencoder model for generative text modeling inspired by Kingma and Welling 2013. Instead of modeling p(x)p(x) directly as in vanilla language models, VAEs introduce a continuous latent variable zz and take the form p(z)p(x∣z)p(z)p(x|z). To train a VAE, we optimize the marginal likelihood p(x)=∫pθ(z)p(x∣z)dzp(x)=\int p_{\theta}(z)p(x|z)dz. The marginal log likelihood can be written as:

qϕ(z∣x)q_{\phi}(z|x), a variational approximation to the posterior p(z∣x)p(z|x), can be variously interpreted as a recognition model or encoder, parameterized by a neural network to encode the sentence xx into a dense code zz. L(θ,ϕ;x)\mathcal{L}(\theta,\phi;x) is often called the evidence lower bound (ELBO). The first term of ELBO is the KL divergence of the approximate posterior from prior and the second term is an expected reconstruction error.

Since KL divergence is always non-negative, we can use L(θ,ϕ;x)\mathcal{L}(\theta,\phi;x) as a lower bound of marginal likelihood log⁡pθ(x)\log p_{\theta}(x). We optimize L(θ,ϕ;x)\mathcal{L}(\theta,\phi;x), jointly learning the recognition model parameters ϕ\phi and generative model parameters θ\theta.

A Neural Variational RNN (NVRNN) for language modeling is described in Bowman et al. 2016 and depicted in Figure 1. The goal of the NVRNN model is to extract a high level representation of a sentence into zz and reconstruct the sentence with a neural language model.

We denote a sequence of words as x={x1,x2,⋯ ,xn}x=\{x_{1},x_{2},\cdots,x_{n}\}. Unlike in vanilla language modeling, an NVRNN conditions on the latent variable zz at each step of the generation pθ(x∣z)=pθ(x1∣z)∏i=1np(xi∣x1,…,xi−1,z).p_{\theta}(x|z)=p_{\theta}(x_{1}|z)\prod_{i=1}^{n}p(x_{i}|x_{1},\ldots,x_{i-1},z). This probability distribution is modeled using a recurrent model like an LSTM Hochreiter and Schmidhuber 1997 as illustrated in Figure 1. There is nothing unique about this choice; other recurrent sequence models like a CNN or a Transformer Vaswani et al. 2017 could be used.

2 Posterior Collapse

When training a VAE, we update θ\theta and ϕ\phi simultaneously. Optimizing Eq. 1 gives two gradient terms: an update from the reconstruction loss (likelihood of the correct labels) and an update from the KL divergence. While the reconstruction loss term encourages the zz to convey useful information to this model, the KL term consistently tries to regularize q(z∣x)q(z|x) towards the prior on every gradient update. This may trap the model in a bad local optimum where qϕ(z∣x)=pθ(z)q_{\phi}(z|x)=p_{\theta}(z) for all xx: in this case, zz is simply a noise source, which is useless to the model, so the model has learned to ignore it and will not make large enough gradient updates to break q(z∣x)q(z|x) out of this optimum.

Bowman et al. 2016 termed this issue KL collapse and proposed an annealing schedule to handle it, where the weight of the KL term is increased over the course of training. Reweighting the KL term is also used in methods like β\beta-VAE Higgins et al. 2017 and InfoVAE Zhao et al. 2017b. In this way, the model initially learns to use the latent code but is then regularized towards the prior as training progresses. However, this trick is not sufficient to avert KL collapse in all scenarios, particularly when strong decoders are used and zz has a minor impact on pθ(x∣z)p_{\theta}(x|z).

Table 1 shows experiments in a similar setup to that of Bowman et al. 2016. We train an NVRNN model on the Penn Treebank with four different hyperparameter settings. We either use a 3-layer LSTM encoder or a 1-layer LSTM and use or do not use a sigmoid annealing schedule (increase the KL weight from 0 to 1 over the first 20 epochs). We observe the best performance using the 1-layer model with annealing. One might conclude from this table that the annealing trick has worked since both models achieve better performance when annealing is used. But in fact, a vMF-based model can do better than either (NLL of 117), and moreover, we have no way of knowing that a better annealing scheme might not achieve even higher performance after training. Furthermore, the higher-capacity 3-layer model can theoretically do anything the 1-layer model can, so its lower performance indicates that our training is derailed either by overfitting or getting stuck in a local optimum where the latent variable is unused. In our experiments, we found significant variance in collapse frequency due to other hyperparameters including whether the encoder is a unidirectional or bidirectional LSTM.

Getting the best performance out of a VAE is, therefore, a challenging problem that requires careful tuning of the objective function and optimization procedure Bowman et al. 2016; Zhao et al. 2017b; Higgins et al. 2017. Beyond the well-documented problem of KL collapse, an optimizer may simply get stuck in a local optimum during training and as a result, fail to find a model that most effectively exploits the latent variable.

The solution we advocate for in this paper is to change the distribution for the latent space and simplify the optimization problem. In the next section, we describe the von Mises-Fisher distribution and its use in VAE, where it forces the model to put the latent representations on the surface of the unit hypersphere rather than squeezing everything to the origin. Critically, this distribution lets us fix the value of the KL term by fixing the distribution’s concentration parameter κ\kappa; this averts the KL collapse and leads to good model performance across two generative modeling paradigms.

von Mises-Fisher VAE

where IvI_{v} stands for the modified Bessel function of the first kind at order vv.

Figure 1 shows samples from vMF distributions with various μ\mu vectors (arrows), d=3d=3, and κ=100\kappa=100. This is a high κ\kappa value, leading to samples that are tightly clustered around μ\mu, which is the mean and mode of the distribution. When κ=0\kappa=0, the distribution degenerates to a uniform distribution over the hypersphere independent of μ\mu.

Past work has used vMF as an emission distribution in unsupervised clustering models Banerjee et al. 2005, VAE for other domains Davidson et al. 2018; Hasnat et al. 2017, and a generative editing model for text Guu et al. 2018. We focus specifically on the empirical properties of vMF for text modeling and conduct a systematic examination of how this prior affects VAE models compared to using a Gaussian.

We will use vMF as both our prior and variational posterior in our VAE models. Otherwise, the setup for our VAE remains the same as in the Gaussian case established in Section 2. Our prior is the uniform distribution vMF(⋅,κ=0\cdot,\kappa=0). Since true posterior pθ(z∣x)p_{\theta}(z|x) is intractable, we will approximate it with a variational posterior qϕ(z∣x)=vMF(z;μ,κ)q_{\phi}(z|x)=\text{vMF}(z;\mu,\kappa) where the mean direction μ\mu is the output of encoding neural networks (Figure 1, right side) and κ\kappa is treated as a constant.

Before we can implement a VAE, we need to derive an expression for KL divergence in order to optimize ELBO (Equation 1) and give a sampling algorithm that admits the reparameterization trick Kingma and Welling 2013.

KL divergence

With vMF(⋅,0\cdot,0) as our prior, the KL divergence is: Our KL divergence agrees with that of Davidson et al. 2018 (see their appendix for a derivation), and we have verified it empirically. The equation in Guu et al. 2018 gives slightly different KL values, though differences are small (<<5%) for most κ\kappa and dimension values we encounter.

Critically, this only depends on κ\kappa, not on μ\mu. κ\kappa will be treated as a fixed hyperparameter, so this term will be constant for our model; KL collapse will therefore be rendered impossible.

Figure 2 shows a visualization of the learning trajectories of Gaussian and vMF VAE. For the Gaussian VAE, the KL divergence in the objective function tends to pull the posterior towards the prior centered at the origin and, therefore, make the optimization difficult as mentioned before. For the vMF VAE, given fixed κ\kappa, there is no such vacuous state and μ\mu can vary freely.

Figure 3 shows the KL value and concentration of vMF(μ,κ\mu,\kappa) for two different dimensionalities. KL increases monotonically with κ\kappa, as does concentration measured by cosine similarity. To get a fixed cosine dispersion as dimensionality increases, higher κ\kappa values are needed, resulting in higher KL values.

Sampling from vMF

Following the implementation of Guu et al. 2018, we use the rejection sampling scheme of Wood 1994 to sample a “change magnitude” ww. Our sample is then given by z=wμ+v1−w2z=w\mu+v\sqrt{1-w^{2}}, where vv is a randomly sampled unit vector tangent to the hypersphere at μ\mu. Neither vv nor ww depends on μ\mu, so we can now take gradients of zz with respect to μ\mu as required.

Experiments on Language Modeling

We first evaluate our vMF approach in the NVRNN setting. We will return to this model and analyze its properties further in Sections 6 and 7 after showing experiments on document modeling.

For NVRNN, we use the Penn Treebank Marcus et al. 1993, also used in Bowman et al. 2016, and Yelp 2013 Xu et al. 2016. Examples in the Yelp dataset are much longer and more diverse than those from PTB, requiring more understanding of high-level semantics to generate a coherent sequence. Yelp has a long tail of very long reviews, so we truncate the examples to a maximum length of 50 words; this still gives an average length over twice as long as in the PTB setting. Statistics about all datasets used in this paper are shown in Table 2.

Settings

We evaluate our NVRNN as in Bowman et al. 2016 and explore two different settings. In the Standard setting, the input to the RNN at each time step is the concatenation of the latent code zz and the ground truth word from the last time step, while the Inputless setting does not use the prior word. The more powerful decoder of the Standard setting makes the latent representations inherently less useful. In the Inputless setting, the decoder needs to predict the whole sequence with only the help of given latent code. In this case, a high-quality representation of the sentence is badly needed and the model is driven to learn it.

Our implementation of VAE uses a one layer unidirectional LSTM as both encoder and decoder. We use an embedding size of 100 and hidden units of size 400 in the LSTM. The dimension of the latent code is chosen from {25,50,100}\{25,50,100\} by tuning on the development set. We use SGD to optimize all models with decayed learning rate and gradient clipping. For Yelp, the sentiment bit, which ranges from 1 to 5, is also embedded into a 50 dimension vector and input for every time step of the decoding phase.

Results

Experimental results of the NVRNN are shown in Table 3. We report negative log likelihood (NLL) Reported values are actually a lower bound on the true NLL, computed from ELBO by sampling zz. and perplexity (PPL) on the test set. We follow the implementation reported in Bowman et al. 2016 where the KL term weight is annealed for the Gaussian VAE; vMF VAE works well without weight annealing. The vMF distribution gives a performance boost in all datasets in both the Standard and Inputless settings. Even in the Standard setting, our model is able to successfully use nonzero KL values to achieve better perplexities, and even when KL collapse does not appear to be the case (e.g., G-VAE on the PTB-Standard setting), a Gaussian family of distributions results in lower KLs and worse log likelihoods, possibly due to optimization challenges. In the Inputless setting, we see large gains: vMF VAE reduces PPL from 379 to 262 in PTB, and from 256 to 134 in Yelp compared to Gaussian VAE.

Trade-off Comparison

Besides the overall perplexity, we are also interested in the trade-off between reconstruction loss and KL, and the contribution of KL to the whole objective. Figure 4 shows the ability of our model to explicitly control the balance between the KL and the reconstruction term. First, we “permanently” anneal the Gaussian VAE by setting the weight of the KL term to a constant smaller than 1 (0.2 and 0.5 in our case). We find that this trick does mitigate the KL collapse, but the overall performance is worse. Therefore, this is not only a numerical game about the KL vs. NLL trade-off but a deeper challenge of how to structure models to learn effective latent representations.

For vMF VAE, when we gradually increase the value of κ\kappa, the concentration of the distribution around the mean direction μ\mu is higher and samples from vMF are closer to μ\mu. The model achieves the best perplexity when κ=80\kappa=80. The reconstruction error is bounded around 4.5 due to the difficulty of the task and limited capacity of LSTM decoder. While κ\kappa is a hyperparameter that needs to be tuned, the model is overall not very sensitive to it, and we show in Section 7 that reasonable κ\kappa values transfer across similar tasks.

Experiments on Document Modeling

We also investigate how vMF VAE performs in a different setting, one less plagued by the KL collapse issue. Specifically, the Neural Variational Document Model (NVDM), proposed by Miao et al. 2016, is a VAE-based unsupervised document model. This model follows the VAE framework introduced in Section 2. Our document representation is an indicator vector xx of word presence or absence in the document. Since this is a fixed-size representation, we use 2-layer MLPs with 400 hidden units for both the encoder q(z∣x)q(z|x) and decoder p(x∣z)p(x|z); the decoder places a simple multinomial distribution over words in the vocabulary, and the probability of a document is the product of the probabilities of its words.

For NVDM, we use two standard news corpus, 20 News Groups (20NG) and the Reuters RCV1-v2, which were used in Miao et al. 2016. The preprocessed version can be downloaded from https://github.com/ysmiao/nvdm

Results

Experimental results We do not compare to results from Serban et al. 2017a. Compared to our current results, that work reports very strong performance on 20NG and very weak performance on RCV1. Based on consultation with the authors, they use different preprocessing than Miao et al. 2016. are shown in Table 4. In contrast with NVRNN, the NVDM fully relies on the power of latent code to predict the word distribution, so we never observe a KL collapse, yet vMF still achieves better performance than Gaussian. As shown in Figure 3, in order to keep the same amount of dispersion in samples from the variational posterior, larger latent dimensions need larger κ\kappa values and correspondingly larger KL term values. For 20NG, which is much smaller than RCV1, smaller dimensions therefore give better performance. For both datasets, the settings of κ=100,dim=25\kappa=100,\text{dim}=25 and κ=150,dim∈{50,200}\kappa=150,\text{dim}\in\{50,200\} work well.

What do our VAEs encode?

We design more probing tasks to demonstrate what is encoded in latent representations induced by vMF VAE. One additional model variant we explore here is the NVRNN-BoW model. This is a variant of NVRNN where the decoder additionally conditions on the vector BoW=1n∑i=1ne(xi)BoW=\frac{1}{n}\sum_{i=1}^{n}e(x_{i}), the average word embedding value of the sentence xx. While an artificial setting, this lets us see how effectively the latent code can capture information other than simple word choice by making a form of this information independently available. Table 5 shows results in this setting, where once again we see the KL collapse problem for the Gaussian models and better performance from vMF on perplexity in both the Standard and Inputless settings.

For all of these models, one hypothesis is that the encoder may be learning to memorize the bag of words and then preferentially generate words in that bag from the decoder. To verify this, we investigate whether the BoW representation and the learned latent code can be reconstructed from each other. Specifically, given a sentence xx we can compute BoWBoW as defined above and μ=enc(x)\mu=\textrm{enc}(x), the latent encoding of xx as represented by the mean vector output by the encoder. We can use a simple multilayer perceptron to to try to map from the bag of words to the latent code: μ^=MLP(BoW)\hat{\mu}=MLP(BoW), then learn the parameters of the MLP by minimizing ∥μ^−μ∥2\|\hat{\mu}-\mu\|^{2} on a sample. The same process can be used to learn a mapping from μ\mu back to the bag of words.

Table 6 shows averaged cosine similarities of our reconstructions under both Gaussian and vMF models. For vMF, μ\mu can reconstruct the bag-of-words more accurately than the bag-of-words can reconstruct μ\mu, indicating that the latent code in vMF captures more information beyond the bag of words.

We repeat this experiment in a separate NVRNN model where the decoder can explicitly condition on the BoWBoW vector described above. The results are shown in the right column of Table 6. Our model, v-VAE, achieves a lower cosine similarity than G-VAE (0.23 vs. 0.32), indicating that it capturing less redundant information and using the latent space to more efficiently model other properties of the data.

Sensitivity to word order

Table 6 shows that NVRNN with vMF encodes information beyond the bag of words; a natural hypothesis is that it is encoding word order. We can more directly investigate this in the context of both NVRNN and NVRNN-BoW settings. Inspired by Zhao et al. 2017a, we propose an experiment probing the sensitivity to randomly swapping adjacent pairs of words for the encoding in the Inputless setting on PTB. We vary the probability of swapping each word pair and see how the latent code changes as the number of swaps increases. Ideally, our models should capture ordering information and therefore be sensitive to this change.

Figure 5 shows the results. v-VAE’s representations are more sensitive than those of the G-VAE: they change faster as swaps become more likely. The Gaussian VAE here makes very little use of the latent variable, hence why the representations change very little. In the NVRNN-BoW setting, we see that the models are even more sensitive. vMF enables us to more easily learn this kind of desirable information in our sentence encodings.

Controlling Variance with κ\kappa

A core aspect of our approach so far has been treating κ\kappa as a fixed hyperparameter. Fixing κ\kappa is beneficial from an optimization standpoint: it makes it more difficult for the model to get stuck in local optima. But it also reduces the model’s flexibility, since we can no longer predict per-example κ\kappa values, and it introduces another parameter that the system designer must tune.

Fortunately, a wide range of κ\kappa values appear to work well for the tasks we consider. Figure 6 shows how the concentration parameter κ\kappa changes the results on PTB when the latent dimension and other hyperparameters are held fixed. We have ordered the tasks left-to-right from “hardest” to “easiest” in terms of necessity of latent representation: the Inputless setting needs heavy information from the latent code to reconstruct the sentence, whereas the Standard-BoW setting has an extremely strong decoder to predict the next word. We see that in each of these cases, a wide range of κ\kappa values works, and moreover reasonable κ\kappa values transfer between the two Standard and between the two Inputless settings, indicating that the overall approach is not highly sensitive to these hyperparameter values.

Throughout this work, we have treated κ\kappa as a fixed parameter. However, we can treat κ\kappa in the same way as σ\sigma in the Gaussian case and learn it on a per-instance basis. The KL divergence of vMF is differentiable with respect to κ\kappa given gradients of the modified Bessel function of the first kind, ∇κId(κ)=12(Id−1(κ)+Id+1(κ))\nabla_{\kappa}I_{d}(\kappa)=\frac{1}{2}(I_{d-1}(\kappa)+I_{d+1}(\kappa)) allowing us to change the concentration on a per-instance basis. However, this reintroduces the issue of KL collapse: the KL term will encourage κ\kappa to be as low as possible, potentially making the latent variable vacuous.

In practice, we observe that it is necessary to clip κ\kappa values to a certain range for numerical reasons. Within this range, the model gravitates towards the smallest κ\kappa values and performs substantially worse than models trained with our fixed κ\kappa approach. This indicates that even with the vMF model, the optimization problem posed by ELBO is simply a hard one and the approach of fixing KL divergence is a surprisingly good optimization technique.

Related Work

Deep generative models have achieved impressive successes in domains adjacent to NLP such as image generation Gregor et al. 2015; Oord et al. 2016a and speech generation Chung et al. 2015; Oord et al. 2016b. VAEs specifically Kingma and Welling 2013; Rezende et al. 2014 have been a popular model variant in NLP. They have been applied to tasks including document modeling Miao et al. 2016, language modeling Bowman et al. 2016, and dialogue generation Serban et al. 2017b. VAEs can be also be applied for semi-supervised classification Xu et al. 2017. Recent twists on the standard VAE approach including combining VAE and holistic attribute discriminators for conditional generation Hu et al. 2017 and using a more flexible latent space regularized by an adversarial method Zhao et al. 2017a.

VAE Objective

Several pieces of recent work have highlighted the issues with optimizing the VAE objective. Alemi et al. 2018 shed light on the problem from the perspective of information theory. Zhao et al. 2017b and Higgins et al. 2017 both propose various reweightings of the objective along with theoretical and empirical justification.

Choices of Priors for VAE

Some past work has explored various priors for VAE. Serban et al. 2017a proposed a piecewise constant distribution which deals with multiple modes, but which sacrifices the property of continuous interpolation. Guu et al. 2018 also applied vMF in a VAE model, but used theirs specifically in the sentence-editing case. Davidson et al. 2018 explored vMF in a VAE model for MNIST and a link prediction task. Hasnat et al. 2017 applied the vMF distribution for facial recognition. Other past work has used different decoders, including CNNs Yang et al. 2017 and CNN-RNN hybrids Semeniuta et al. 2017. Changing the decoder is a change largely orthogonal to changing the prior: it can alleviate the KL vanishing issue, but it does not necessarily scale to new settings and does not give explicit control over utilization of the latent code.

Conclusion

In this paper, we propose the use of a von Mises-Fisher VAE to resolve optimization issues in variational autoencoders for text. This choice of distribution allows us to explicitly control the balance between the capacity of the decoder and the utilization of the latent representation in a principled way. Experimental results demonstrate that the proposed model has better performance than a Gaussian VAE across a range of settings. Further analysis shows that vMF VAE is more sensitive to word order information and makes more effective use of the latent code space.

Acknowledgments

This work was partially supported by NSF Grant IIS-1814522, a Bloomberg Data Science Grant, and an equipment grant from NVIDIA. The authors acknowledge the Texas Advanced Computing Center (TACC) at The University of Texas at Austin for providing HPC resources used to conduct this research. Thanks as well to the anonymous reviewers for their helpful comments.

References