Riemannian Normalizing Flow on Variational Wasserstein Autoencoder for Text Modeling

Prince Zizhuang Wang, William Yang Wang

Introduction

Variational Autocoder (VAE) Kingma and Welling (2013); Rezende and Mohamed (2015) is a probabilistic generative model shown to be successful over a wide range of tasks such as image generation Gregor et al. (2015); Yan et al. (2016), dialogue generation Zhao et al. (2017b), transfer learning Shen et al. (2017), and classification Jang et al. (2017). The encoder-decoder architecture of VAE allows it to learn a continuous space of latent representations from high-dimensional data input and makes sampling procedure from such latent space very straightforward. Recent studies also show that VAE learns meaningful representations that encode non-trivial information from input Gao et al. (2018); Zhao et al. (2017a).

Applications of VAE in tasks of Natural Language Processing Bowman et al. (2015); Zhao et al. (2017b, 2018); Miao et al. (2016) is not as successful as those in Computer Vision. With long-short-term-memory network (LSTM) Hochreiter and Schmidhuber (1997) used as encoder-decoder model, the recurrent variational autoencoder Bowman et al. (2015) is the first approach that applies VAE to language modeling tasks. They observe that LSTM decoder in VAE often generates texts without making use of latent representations, rendering the learned codes as useless. This phenomenon is caused by an optimization problem called KL-divergence vanishing when training VAE for text data, where the KL-divergence term in VAE objective collapses to zero. This makes the learned representations meaningless as zero KL-divergence indicates that the latent codes are independent of input texts.

Many recent studies are proposed to address this key issue. Yang et al. (2017); Semeniuta et al. (2017) use convolutional neural network as decoder architecture to limit the expressiveness of decoder model. Xu and Durrett (2018); Zhao et al. (2017a, b, 2018) seek to learn different latent space and modify the learning objective. And, even though not designed to tackle KL vanishing at the beginning, recent studies on Normalizing Flows Rezende and Mohamed (2015); van den Berg et al. (2018) learn meaningful latent space as it helps to transform an over-simplified latent distribution into more flexible distributions.

In this paper, we propose a new type of flow, called Riemannian Normalizing Flow (RNF), together with the recently developed Wasserstein objective Tolstikhin et al. (2018); Arjovsky et al. (2017), to ensure VAE models more robust against the KL vanishing problem. As further explained in later sections, the Wasserstein objective helps to alleviate KL vanishing as it only minimizes the distance between latent marginal distribution and the prior. Moreover, we suspect that the problem also comes from the over-simplified prior assumption about latent space. In most cases, the prior is assumed to be a standard Gaussian, and the posterior is assumed to be a diagonal Gaussian for computational efficiency. These assumptions, however, are not suitable to encode intrinsic characteristics of input into latent codes as in reality the latent space is likely to be far more complex than a diagonal Gaussian.

The RNF model we proposed in this paper thus helps the situation by encouraging the model to learn a latent space that encodes some geometric properties of input space with a well-defined geometric metric called Riemannian metric tensor. This renders the KL vanishing problem as impossible since a latent distribution that respects input space geometry would only collapse to a standard Gaussian when the input also follows a standard Gaussian, which is never the case for texts and sentences datasets. We then empirically evaluate our RNF Variational Wasserstein Autoencoder on standard language modeling datasets and show that our model has achieved state-of-the-art performances. Our major contributions can be summarized as the following:

We propose Riemannian Normalizing Flow, a new type of flow that uses the Riemannian metric to encourage latent codes to respect geometric characteristics of input space.

We introduce a new Wasserstein objective for text modeling, which alleviates KL divergence term vanishing issue, and makes the computation of normalizing flow easier.

Empirical studies show that our model produces state-of-the-art results in language modeling and is able to generate meaningful text sentences that are more diverse.

Related Work

2 KL Divergence Term Vanishing

Background

In this section, we review the basic concepts of Riemannian geometry, normalizing flow, and Wasserstein Autoencoder. We then introduce our new Riemannian normalizing flow in Section 4.

Consider an input space X⊂RD\mathcal{X}\subset R^{D}, a d-dimensional (d<D)(d<D) manifold is a smooth surface of points embedded in X\mathcal{X}. Given a manifold M\mathcal{M}, a Riemannian manifold is a metric space (M,G)(\mathcal{M},G), where GG is the Riemannian metric tensor that assigns an inner product to every point on the manifold. More formally, a Riemannian metric G:Z→Rd×dG:\mathcal{Z}\rightarrow R^{d\times d} is defined as a smooth function such that for any two vectors u,vu,v in the tangent space TzMT_{z}M of each point z∈Mz\in\mathcal{M}, it assigns the following inner product for uu and vv,

The Riemannian metric helps us to characterize many intrinsic properties of a manifold. Consider an arbitrary smooth curve γ(t):[a,b]→M\gamma(t):[a,b]\rightarrow\mathcal{M} on a given manifold M\mathcal{M} with a Riemannian metric tensor GG, the length of this curve is given by

where γt′\gamma^{\prime}_{t} is the curve velocity and lies in the tangent space TγtMT_{\gamma_{t}}M at point γ(t)\gamma(t). When the metric tensor GG is equal to 1 everywhere on the curve, it becomes a metric tensor on Euclidean space, where the length of curve is defined as the integral of the velocity function, L(γ)=∫abγ′tTγt′dt=∫abγt′dt\mathcal{L}(\gamma)=\int_{a}^{b}\sqrt{{\gamma^{\prime}}_{t}^{T}\gamma^{\prime}_{t}}dt=\int_{a}^{b}\gamma^{\prime}_{t}dt. Given the definition of curve length, the geodesic path between any two points can be defined as the curve that minimizes the curve length γ\mathcal{\gamma}. Namely, if γt\gamma_{t} is the geodesic curve connecting γ(a)\gamma(a) and γ(b)\gamma(b), then

Practically, a geodesic line is often found by optimizing the following energy function,

Note that the Euclidean metric is a special case of Riemannian metric. The more general metric tensor GG gives us a sense of how much Riemannian geometry deviates from Euclidean geometry.

2 Review: Normalizing Flow

The powerful inference model of VAE can approximate the true posterior distribution through variational inference. The choice of this approximated posterior is one of the major problems. For computational efficiency, a diagonal Gaussian distribution is often chosen as the form of the posterior. As the covariance matrix is always assumed to be diagonal, the posterior fails to capture dependencies among individual dimensions of latent codes. This poses a difficult problem in variational inference. As it is unlikely that the true posterior has a diagonal form, the approximated diagonal distribution is not flexible enough to match the true posterior even in asymptotic time.

A normalizing flow, developed by Rezende and Mohamed (2015), is then introduced to transform a simple posterior to a more flexible distribution. Formally, a series of normalizing flows is a set of invertible, smooth transformations ft:Rd→Rdf_{t}:R^{d}\rightarrow R^{d}, for t=1,...,Tt=1,...,T, such that given a random variable z0z_{0} with distribution q(z0)q(z_{0}), the resulting random variable zT=(fT∘fT−1∘...∘f1)(z0)z^{T}=(f_{T}\circ f_{T-1}\circ...\circ f_{1})(z_{0}) has the following density function,

Since each transformation fif_{i} for i=1,...,Ti=1,...,T is invertible, its Jacobian determinant exists and can be computed. By optimizing the modified evidence lower bound objective,

the resulting latent codes zTz_{T} will have a more flexible distribution.

Based on how the Jacobian-determinant is computed, there are two main families of normalizing flow Tomczak and Welling (2016); Berg et al. (2018): general normalizing flow and volume preserving flow. While they both search for flexible transformation that has easy-to-compute Jacobian-determinant, the volume-preserving flow aims at finding a specific flow whose Jacobian-determinant equals 1, which simplifies the optimization problem in equation (6). Since we want a normalizing flow that not only gives flexible posterior but also able to uncover the true geometric properties of latent space, we only consider general normalizing flow whose Jacobian-determinant is not a constant as we need it to model the Riemannian metric introduced earlier.

3 Wasserstein Autoencoder

Wasserstien distance has been brought to generative models and is shown to be successful in many image generation tasks Tolstikhin et al. (2018); Arjovsky et al. (2017); Bousquet et al. (2017). Instead of maximizing the evidence lower bound as VAE does, the Wasserstein Autoencoder Tolstikhin et al. (2018) optimizes the optimal transport cost Villani (2008) between the true data distribution PX(x)P_{X}(x) and the generative distribution PG(x)P_{G}(x). This leads to the Wasserstein objective,

where c(⋅)c(\cdot) is the optimal transport cost, G:Z→XG:\mathcal{Z}\rightarrow\mathcal{X} is any generative function, and the coefficient λ\lambda controls the strength of regularization term DZD_{Z}. Given a positive-definite reproducing kernel k:Z×Z→Rk:\mathcal{Z}\times\mathcal{Z}\rightarrow\mathcal{R}, the regularization term DZD_{Z} can be approximated by the Maximum Mean Discrepancy (MMD) Gretton et al. (2012) between the prior PZP_{Z} and the aggregate posterior QZ(z)=∫q(z∣x)p(x)dxQ_{Z}(z)=\int q(z|x)p(x)dx,

Our Approach

In this section we propose our Riemannian Normalizing Flow (RNF). RNF is a new type of flow that makes use of the Riemannian metric tensor introduced earlier in Section 3. This metric enforces stochastic encoder to learn a richer class of approximated posterior distribution in order to follow the true geometry of latent space, which helps to avoid the local optimum in which posterior collapses to a standard prior. We then combine this with WAE and we will explain why and how WAE should be used to train with RNF.

In the context of VAE, learning a latent space that is homeomorphic to input space is often very challenging. Consider a manifold M⊂RD\mathcal{M}\subset R^{D}, a generator model x=f(z):Z→RDx=f(z):\mathcal{Z}\rightarrow R^{D} serves as a low-dimensional parameterization of manifold M\mathcal{M} with respect to z∈Zz\in\mathcal{Z}. For most cases, latent space Z\mathcal{Z} is unlikely to be homeomorphic to M\mathcal{M}, which means that there is no invertible mapping between M\mathcal{M} and Z\mathcal{Z}. And, since the inference model h:M→Zh:\mathcal{M}\rightarrow\mathcal{Z} is nonlinear, the learned latent space often gives a distorted view of input space. Consider the case in Figure 2, where the leftmost graph is the input manifold, and the rightmost graph is the corresponding latent space with curvature reflected by brightness. Let us take two arbitrary points on the manifold and search for the geodesic path connecting these two points. If we consider the distorted latent space as Euclidean, then the geodesic path in latent space does not reflect the true shortest distance between these two points on the manifold, as a straight line in the latent space would cross the whole manifold, while the true geodesic path should circumvent this hole. This distortion is caused by the non-constant curvature of latent space. Hence, the latent space should be considered as a curved space with curvature reflected by the Riemannian metric defined locally around each point. As indicated by the brightness, we see that the central area of latent space is highly curved, and thus has higher energy. The geodesic path connecting the two latent codes minimizes the energy function E(γ)=12∫γ′tTG(γt)γt′dtE(\gamma)=\frac{1}{2}\int{\gamma^{\prime}}_{t}^{T}G(\gamma_{t})\gamma^{\prime}_{t}dt, indicating that it should avoid those regions with high curvature GG.

The question now becomes how to impose this intrinsic metric and curvature into latent space. In this paper, we propose a new form of normalizing flow to incorporate with this geometric characteristic. First, consider a normalizing flow f:Z→Z′f:\mathcal{Z}\rightarrow\mathcal{Z^{\prime}}, we can compute length of a curve in the transformed latent space Z′\mathcal{Z^{\prime}},

In this paper, we introduce Riemannian normalizing flow (RNF) to model curvature. For simplicity, we build our model based on planar flow Rezende and Mohamed (2015). A planar flow is an invertible transformation that retracts and extends the support of original latent space with respect to a plane. Mathematically, a planar flow f:Z→Z′f:\mathcal{Z}\rightarrow\mathcal{Z^{\prime}} has the form,

2 RNF Wasserstein Autoencoder

Here we consider using the Wasserstein objective to model the latent marginal distribution of a curved latent space learned by an RNF. The Wasserstein objective with MMD is appealing in our case for two main reasons.

Second, the MMD regularization in WAE makes it possible to optimize normalizing flow without computing the Jacobian-determinant explicitly. The use of MMD is necessary as getting a closed form KL divergence is no-longer possible after we apply RNF to the posterior. And, since the generative function GG in the reconstruction term of Wasserstein objective can be any function or composition of functions Tolstikhin et al. (2018), we can easily compose an RNF function into GG such that the reconstructed texts are Xˉ=G(f(Z))=G(Z′),Z′∈Z′\bar{X}=G(f(Z))=G(Z^{\prime}),Z^{\prime}\in\mathcal{Z}^{\prime}.

Now, given a series of RNF F=fK∘...∘f1F=f_{K}\circ...\circ f_{1}, and let ZK\mathcal{Z}_{K} be the curved latent space after applying KK flows over the original latent space Z\mathcal{Z}, we optimize the following RNF-Wasserstein objective,

Experimental Results

In this section, we investigate WAE’s performance with Riemannian Normalizing Flow over language and text modeling.

We use Penn Treebank Marcus et al. (1993), Yelp 13 reviews Xu et al. (2016), as in Xu and Durrett (2018); Bowman et al. (2015), and Yahoo Answers used in Xu and Durrett (2018); Yang et al. (2017) to follow and compare with prior studies. We limit the maximum length of a sample from all datasets to 200 words. The datasets statistics is shown in Table 3.

2 Experimental Setup

For each model, we set the maximum vocabulary size to 20K and the maximum length of input to 200 across all data sets. Following Bowman et al. (2015), we use one-layer undirectional LSTM for both encoder-decoder models with hidden size 200. Latent codes dimension is set to 32 for all models. We share Word Embeddings of size 200. For stochastic encoders, both MLPμMLP_{\mu} and MLPσMLP_{\sigma} are two layer fully-connected networks with hidden size 200 and a batch normalizing output layer Ioffe and Szegedy (2015).

We use Adam Kingma and Ba (2015) with learning rate set to 10−310^{-3} to train all models. Dropout is used and is set to 0.2. We train all models for 48 epochs, each of which consists of 2K steps. For models other than WAE, KL-annealing is applied and is scheduled from 0 to 1 at the 21st epoch.

For vmf-VAE Xu and Durrett (2018), we set the word embedding dimension to be 512 and the hidden units to 1024 for Yahoo, and set both of them to 200 for PTB and Yelp. The temperature κ\kappa is set to 80 and is kept constant during training.

3 Language Modeling Results

We show the language modeling results for PTB, Yahoo and Yelp in Table 1. We compare negative log-likelihood (NLL), KL divergence, and perplexity (PPL) with all other existing methods. The negative log-likelihood is approximated by its lower bound.

The numbers show that KL-annealing and dropout used by Bowman et al. (2015) are helpful for PTB, but for complex datasets such as Yahoo and Yelp, the KL divergence still drops to zero due to the over-expressiveness of LSTM. This phenomenon is not alleviated by applying normalizing flow to make the posterior more flexible, as shown in the third row. Part of the reason may be that a simple NF such as a planar flow is not flexible enough and is still dominated by a powerful LSTM decoder.

We find that the KL vanishing is alleviated a little bit if using WAE, which should be the case as WAE objective does not require small KL. We also find that simply applying a planar flow over WAE does not improve the performance that much. On the other hand, using RNF to train WAE dramatically helps the situation which achieves the lowest text perplexity on most conditions except for YAHOO Answers, where He et al. (2019); Yang et al. (2017); Kim et al. (2018) have the current state-of-the-art results. We want to emphasize that CNN-VAE Yang et al. (2017) and SA-VAE Kim et al. (2018) are not directly comparable with other current approaches. Here, we compare with models that use LSTM as encoder-decoder and have similar time complexity, while the use of CNN as decoder in CNN-VAE would dramatically change the model expressiveness, and it is known that SA-VAE’s time complexity Kim et al. (2018); He et al. (2019) is much higher than all other existing approaches.

4 How Good is Riemannian Latent Representation?

Generating Texts from latent spaces

Another way to explore latent space is to look at the quality of generated texts. Here we compare sentences generated from methods that do not use Wasserstein objective and RNF with those generated from curved latent space Z′\mathcal{Z^{\prime}} learned by WAE.

Conclusion

In this paper, we introduced Riemannian Normalizing Flow to train Wasserstein Autoencoder for text modeling. This new model encourages learned latent representation of texts to respect geometric characteristics of input sentences space. Our results show that RNF WAE does significantly improve the language modeling results by modeling the Riemannian geometric space via normalizing flow.

Acknowledgement

We want to thank College of Creative Studies and Gene &\& Lucas Undergraduate Research Fund for providing scholarships and research opportunities for Prince Zizhuang Wang. We also want to thank Yanxin Feng from Wuhan University for helpful discussion about Riemannian Geometry, and Junxian He (CMU), Wenhu Chen (UCSB), and Yijun Xiao (UCSB) for their comments which helped us improve our paper and experiments.

References