Adversarial Symmetric Variational Autoencoder

Yunchen Pu, Weiyao Wang, Ricardo Henao, Liqun Chen, Zhe Gan, Chunyuan Li, Lawrence Carin

Introduction

Recently there has been increasing interest in developing generative models of data, offering the promise of learning based on the often vast quantity of unlabeled data. With such learning, one typically seeks to build rich, hierarchical probabilistic models that are able to fit to the distribution of complex real data, and are also capable of realistic data synthesis.

Generative models are often characterized by latent variables (codes), and the variability in the codes encompasses the variation in the data . The generative adversarial network (GAN) employs a generative model in which the code is drawn from a simple distribution (e.g.e.g., isotropic Gaussian), and then the code is fed through a sophisticated deep neural network (decoder) to manifest the data. In the context of data synthesis, GANs have shown tremendous capabilities in generating realistic, sharp images from models that learn to mimic the structure of real data . The quality of GAN-generated images has been evaluated by somewhat ad hoc metrics like inception score .

However, the original GAN formulation does not allow inference of the underlying code, given observed data. This makes it difficult to quantify the quality of the generative model, as it is not possible to compute the quality of model fit to data. To provide a principled quantitative analysis of model fit, not only should the generative model synthesize realistic-looking data, one also desires the ability to infer the latent code given data (using an encoder). Recent GAN extensions have sought to address this limitation by learning an inverse mapping (encoder) to project data into the latent space, achieving encouraging results on semi-supervised learning. However, these methods still fail to obtain faithful reproductions of the input data, partly due to model underfitting when learning from a fully adversarial objective .

Variational autoencoders (VAEs) are designed to learn both an encoder and decoder, leading to excellent data reconstruction and the ability to quantify a bound on the log-likelihood fit of the model to data . In addition, the inferred latent codes can be utilized in downstream applications, including classification and image captioning . However, new images synthesized by VAEs tend to be unspecific and/or blurry, with relatively low resolution. These limitations of VAEs are becoming increasingly understood. Specifically, the traditional VAE seeks to maximize a lower bound on the log-likelihood of the generative model, and therefore VAEs inherit the limitations of maximum-likelihood (ML) learning . Specifically, in ML-based learning one optimizes the (one-way) Kullback-Leibler (KL) divergence between the distribution of the underlying data and the distribution of the model; such learning does not penalize a model that is capable of generating data that are different from that used for training.

Based on the above observations, it is desirable to build a generative-model learning framework with which one can compute and assess the log-likelihood fit to real (observed) data, while also being capable of generating synthetic samples of high realism. Since GANs and VAEs have complementary strengths, their integration appears desirable, with this a principal contribution of this paper. While integration seems natural, we make important changes to both the VAE and GAN setups, to leverage the best of both. Specifically, we develop a new form of the variational lower bound, manifested jointly for the expected log-likelihood of the observed data and for the latent codes. Optimizing this variational bound involves maximizing the expected log-likelihood of the data and codes, while simultaneously minimizing a symmetric KL divergence involving the joint distribution of data and codes. To compute parts of this variational lower bound, a new form of adversarial learning is invoked. The proposed framework is termed Adversarial Symmetric VAE (AS-VAE), since within the model (ii) the data and codes are treated in a symmetric manner, (iiii) a symmetric form of KL divergence is minimized when learning, and (iiiiii) adversarial training is utilized. To illustrate the utility of AS-VAE, we perform an extensive set of experiments, demonstrating state-of-the-art data reconstruction and generation on several benchmarks datasets.

Background and Foundations

Consider an observed data sample x{\boldsymbol{x}}, modeled as being drawn from pθ(x∣z)p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}}), with model parameters θ{\boldsymbol{\theta}} and latent code z{\boldsymbol{z}}. The prior distribution on the code is denoted p(z)p({\boldsymbol{z}}), typically a distribution that is easy to draw from, such as isotropic Gaussian. The posterior distribution on the code given data x{\boldsymbol{x}} is pθ(z∣x)p_{\boldsymbol{\theta}}({\boldsymbol{z}}|{\boldsymbol{x}}), and since this is typically intractable, it is approximated as qϕ(z∣x)q_{{\boldsymbol{\phi}}}({\boldsymbol{z}}|{\boldsymbol{x}}), parameterized by learned parameters ϕ{\boldsymbol{\phi}}. Conditional distributions qϕ(z∣x)q_{{\boldsymbol{\phi}}}({\boldsymbol{z}}|{\boldsymbol{x}}) and pθ(x∣z)p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}}) are typically designed such that they are easily sampled and, for flexibility, modeled in terms of neural networks . Since z{\boldsymbol{z}} is a latent code for x{\boldsymbol{x}}, qϕ(z∣x)q_{{\boldsymbol{\phi}}}({\boldsymbol{z}}|{\boldsymbol{x}}) is also termed a stochastic encoder, with pθ(x∣z)p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}}) a corresponding stochastic decoder. The observed data are assumed drawn from q(x)q({\boldsymbol{x}}), for which we do not have an explicit form, but from which we have samples, i.e., the ensemble {xi}i=1,N\{{\boldsymbol{x}}_{i}\}_{i=1,N} used for learning.

Our goal is to learn the model pθ(x)=∫pθ(x∣z)p(z)dzp_{\boldsymbol{\theta}}({\boldsymbol{x}})=\int p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}})p({\boldsymbol{z}})d{\boldsymbol{z}} such that it synthesizes samples that are well matched to those drawn from q(x)q({\boldsymbol{x}}). We simultaneously seek to learn a corresponding encoder qϕ(z∣x)q_{{\boldsymbol{\phi}}}({\boldsymbol{z}}|{\boldsymbol{x}}) that is both accurate and efficient to implement. Samples x{\boldsymbol{x}} are synthesized via x∼pθ(x∣z){\boldsymbol{x}}\sim p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}}) with z∼p(z){\boldsymbol{z}}\sim p({\boldsymbol{z}}); z∼qϕ(z∣x){\boldsymbol{z}}\sim q_{\boldsymbol{\phi}}({\boldsymbol{z}}|{\boldsymbol{x}}) provides an efficient coding of observed x{\boldsymbol{x}}, that may be used for other purposes (e.g.e.g., classification or caption generation when x{\boldsymbol{x}} is an image ).

Maximum likelihood (ML) learning of θ{\boldsymbol{\theta}} based on direct evaluation of pθ(x)p_{\boldsymbol{\theta}}({\boldsymbol{x}}) is typically intractable. The VAE seeks to bound pθ(x)p_{\boldsymbol{\theta}}({\boldsymbol{x}}) by maximizing variational expression LVAE(θ,ϕ)\mathcal{L}_{\rm VAE}({\boldsymbol{\theta}},{\boldsymbol{\phi}}), with respect to parameters {θ,ϕ}\{{\boldsymbol{\theta}},{\boldsymbol{\phi}}\}, where

Maximizing LVAE(θ,ϕ)\mathcal{L}_{\rm VAE}({\boldsymbol{\theta}},{\boldsymbol{\phi}}) wrt {θ,ϕ}\{{\boldsymbol{\theta}},{\boldsymbol{\phi}}\} provides a lower bound on 1N∑i=1Nlog⁡pθ(xi)\frac{1}{N}\sum_{i=1}^{N}\log p_{\boldsymbol{\theta}}({\boldsymbol{x}}_{i}), hence the VAE setup is an approximation to ML learning of θ{\boldsymbol{\theta}}. Learning θ{\boldsymbol{\theta}} based on 1N∑i=1Nlog⁡pθ(xi)\frac{1}{N}\sum_{i=1}^{N}\log p_{\boldsymbol{\theta}}({\boldsymbol{x}}_{i}) is equivalent to learning θ{\boldsymbol{\theta}} based on minimizing \mboxKL(q(x)∥pθ(x))\mbox{KL}(q({\boldsymbol{x}})\|p_{\boldsymbol{\theta}}({\boldsymbol{x}})), again implemented in terms of the NN observed samples of q(x)q({\boldsymbol{x}}). As discussed in , such learning does not penalize θ{\boldsymbol{\theta}} severely for yielding x{\boldsymbol{x}} of relatively high probability in pθ(x)p_{\boldsymbol{\theta}}({\boldsymbol{x}}) while being simultaneously of low probability in q(x)q({\boldsymbol{x}}). This means that θ{\boldsymbol{\theta}} seeks to match pθ(x)p_{\boldsymbol{\theta}}({\boldsymbol{x}}) to the properties of the observed data samples, but pθ(x)p_{\boldsymbol{\theta}}({\boldsymbol{x}}) may also have high probability of generating samples that do not look like data drawn from q(x)q({\boldsymbol{x}}). This is a fundamental limitation of ML-based learning , inherited by the traditional VAE in (1).

One reason for the failing of ML-based learning of θ{\boldsymbol{\theta}} is that the cumulative posterior on latent codes ∫pθ(z∣x)q(x)dx≈∫qϕ(z∣x)q(x)dx=qϕ(z)\int p_{\boldsymbol{\theta}}({\boldsymbol{z}}|{\boldsymbol{x}})q({\boldsymbol{x}})d{\boldsymbol{x}}\approx\int q_{\boldsymbol{\phi}}({\boldsymbol{z}}|{\boldsymbol{x}})q({\boldsymbol{x}})d{\boldsymbol{x}}=q_{\boldsymbol{\phi}}({\boldsymbol{z}}) is typically different from p(z)p({\boldsymbol{z}}), which implies that x∼pθ(x∣z){\boldsymbol{x}}\sim p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}}), with z∼p(z){\boldsymbol{z}}\sim p({\boldsymbol{z}}) may yield samples x{\boldsymbol{x}} that are different from those generated from q(x)q({\boldsymbol{x}}). Hence, when learning {θ,ϕ}\{{\boldsymbol{\theta}},{\boldsymbol{\phi}}\} one may seek to match pθ(x)p_{\boldsymbol{\theta}}({\boldsymbol{x}}) to samples of q(x)q({\boldsymbol{x}}), as done in (1), while simultaneously matching qϕ(z)q_{\boldsymbol{\phi}}({\boldsymbol{z}}) to samples of p(z)p({\boldsymbol{z}}). The expression in (1) provides a variational bound for matching pθ(x)p_{\boldsymbol{\theta}}({\boldsymbol{x}}) to samples of q(x)q({\boldsymbol{x}}), thus one may naively think to simultaneously set a similar variational expression for qϕ(z)q_{\boldsymbol{\phi}}({\boldsymbol{z}}), with these two variational expressions optimized jointly. However, to compute this additional variational expression we require an analytic expression for qϕ(x,z)=qϕ(z∣x)q(x)q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}})=q_{\boldsymbol{\phi}}({\boldsymbol{z}}|{\boldsymbol{x}})q({\boldsymbol{x}}), which also means we need an analytic expression for q(x)q({\boldsymbol{x}}), which we do not have.

Examining (2), we also note that LVAE(θ,ϕ)\mathcal{L}_{\rm VAE}({\boldsymbol{\theta}},{\boldsymbol{\phi}}) approximates −\mboxKL(qϕ(x,z)∥pθ(x,z))-\mbox{KL}(q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}})\|p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})), which has limitations aligned with those discussed above for ML-based learning of θ{\boldsymbol{\theta}}. Analogous to the above discussion, we would also like to consider −\mboxKL(pθ(x,z)∥qϕ(x,z))-\mbox{KL}(p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})\|q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}})). So motivated, in Section 3 we develop a new form of variational lower bound, applicable to maximizing 1N∑i=1Nlog⁡pθ(xi)\frac{1}{N}\sum_{i=1}^{N}\log p_{\boldsymbol{\theta}}({\boldsymbol{x}}_{i}) and 1M∑j=1Mlog⁡qϕ(zj)\frac{1}{M}\sum_{j=1}^{M}\log q_{\boldsymbol{\phi}}({\boldsymbol{z}}_{j}), where zj∼p(z){\boldsymbol{z}}_{j}\sim p({\boldsymbol{z}}) is the jj-th of MM samples from p(z)p({\boldsymbol{z}}). We demonstrate that this new framework leverages both \mboxKL(pθ(x,z)∥qϕ(x,z))\mbox{KL}(p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})\|q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}})) and \mboxKL(qϕ(x,z)∥pθ(x,z))\mbox{KL}(q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}})\|p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})), by extending ideas from adversarial networks.

2 Adversarial Learning

The original idea of GAN was to build an effective generative model pθ(x∣z)p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}}), with z∼p(z){\boldsymbol{z}}\sim p({\boldsymbol{z}}), as discussed above. There was no desire to simultaneously design an inference network qϕ(z∣x)q_{\boldsymbol{\phi}}({\boldsymbol{z}}|{\boldsymbol{x}}). More recently, authors have devised adversarial networks that seek both pθ(x∣z)p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}}) and qϕ(z∣x)q_{\boldsymbol{\phi}}({\boldsymbol{z}}|{\boldsymbol{x}}). As an important example, Adversarial Learned Inference (ALI) considers the following objective function:

where the expectations are approximated with samples, as in (1). The function fψ(x,z)f_{\boldsymbol{\psi}}({\boldsymbol{x}},{\boldsymbol{z}}), termed a discriminator, is typically implemented using a neural network with parameters ψ{\boldsymbol{\psi}} . Note that in (3) we need only sample from pθ(x,z)=pθ(x∣z)p(z)p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})=p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}})p({\boldsymbol{z}}) and qϕ(x,z)=qϕ(z∣x)q(x)q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}})=q_{\boldsymbol{\phi}}({\boldsymbol{z}}|{\boldsymbol{x}})q({\boldsymbol{x}}), avoiding the need for an explicit form for q(x)q({\boldsymbol{x}}).

The framework in (3) can, in theory, match pθ(x,z)p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}}) and qϕ(x,z)q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}}), by finding a Nash equilibrium of their respective non-convex objectives . However, training of such adversarial networks is typically based on stochastic gradient descent, which is designed to find a local mode of a cost function, rather than locating an equilibrium . This objective mismatch may lead to the well-known instability issues associated with GAN training .

To alleviate this problem, some researchers add a regularization term, such as reconstruction loss or mutual information , to the GAN objective, to restrict the space of suitable mapping functions, thus avoiding some of the failure modes of GANs, i.e., mode collapsing. Below we will formally match the joint distributions as in (3), and reconstruction-based regularization will be manifested by generalizing the VAE setup via adversarial learning. Toward this goal we consider the following lemma, which is analogous to Proposition 1 in .

Consider Random Variables (RVs) x{\boldsymbol{x}} and z{\boldsymbol{z}} with joint distributions, p(x,z)p({\boldsymbol{x}},{\boldsymbol{z}}) and q(x,z)q({\boldsymbol{x}},{\boldsymbol{z}}). The optimal discriminator D∗(x,z)=σ(f∗(x,z))D^{*}({\boldsymbol{x}},{\boldsymbol{z}})=\sigma(f^{*}({\boldsymbol{x}},{\boldsymbol{z}})) for the following objective

is f∗(x,z)=log⁡p(x,z)−log⁡q(x,z)f^{*}({\boldsymbol{x}},{\boldsymbol{z}})=\log p({\boldsymbol{x}},{\boldsymbol{z}})-\log q({\boldsymbol{x}},{\boldsymbol{z}}).

Under Lemma 1, we are able to estimate the log⁡qϕ(x,z)−log⁡pθ(x)p(z)\log q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}})-\log p_{\boldsymbol{\theta}}({\boldsymbol{x}})p({\boldsymbol{z}}) and log⁡pθ(x,z)−log⁡q(x)qϕ(z)\log p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})-\log q({\boldsymbol{x}})q_{\boldsymbol{\phi}}({\boldsymbol{z}}) using the following corollary.

For RVs x{\boldsymbol{x}} and z{\boldsymbol{z}} with encoder joint distribution qϕ(x,z)=q(x)qϕ(z∣x)q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}})=q({\boldsymbol{x}})q_{\boldsymbol{\phi}}({\boldsymbol{z}}|{\boldsymbol{x}}) and decoder joint distribution pθ(x,z)=p(z)pθ(x∣z)p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})=p({\boldsymbol{z}})p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}}), consider the following objectives:

If the parameters ϕ{\boldsymbol{\phi}} and θ{\boldsymbol{\theta}} are fixed, with fψ1∗f_{{\boldsymbol{\psi}}_{1}^{*}} the optimal discriminator for (5) and fψ2∗f_{{\boldsymbol{\psi}}_{2}^{*}} the optimal discriminator for (6), then

The proof is provided in the Appendix A. We also assume in Corollary 1.1 that fψ1(x,z)f_{{\boldsymbol{\psi}}_{1}}({\boldsymbol{x}},{\boldsymbol{z}}) and fψ2(x,z)f_{{\boldsymbol{\psi}}_{2}}({\boldsymbol{x}},{\boldsymbol{z}}) are sufficiently flexible such that there are parameters ψ1∗{{\boldsymbol{\psi}}_{1}^{*}} and ψ2∗{{\boldsymbol{\psi}}_{2}^{*}} capable of achieving the equalities in (7). Toward that end, fψ1f_{{\boldsymbol{\psi}}_{1}} and fψ2f_{{\boldsymbol{\psi}}_{2}} are implemented as ψ1{\boldsymbol{\psi}}_{1}- and ψ2{\boldsymbol{\psi}}_{2}-parameterized neural networks (details below), to encourage universal approximation .

Adversarial Symmetric Variational Auto-Encoder (AS-VAE)

A similar expression holds for LVAEz(θ,ϕ)\mathcal{L}_{\rm VAEz}({\boldsymbol{\theta}},{\boldsymbol{\phi}}), in terms of fψ2∗(x,z)f_{{\boldsymbol{\psi}}_{2}^{*}}({\boldsymbol{x}},{\boldsymbol{z}}). This naturally suggests the cumulative variational expression

where ψ1{\boldsymbol{\psi}}_{1} and ψ2{\boldsymbol{\psi}}_{2} are updated using the adversarial objectives in (5) and (6), respectively.

In the original VAE, in which (1) was optimized, the reparametrization trick was invoked wrt qϕ(z∣x)q_{\boldsymbol{\phi}}({\boldsymbol{z}}|{\boldsymbol{x}}), with samples zϕ(x,ϵ){\boldsymbol{z}}_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{\epsilon}}) and ϵ∼N(0,I){\boldsymbol{\epsilon}}\sim\mathcal{N}(0,{\bf I}), as the expectation was performed wrt this distribution; this reparametrization is convenient for computing gradients wrt ϕ{\boldsymbol{\phi}}. In the AS-VAE in (10), expectations are also needed wrt pθ(x∣z)p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}}). Hence, to implement gradients wrt θ{\boldsymbol{\theta}}, we also constitute a reparametrization of pθ(x∣z)p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}}). Specifically, we consider samples xθ(z,ξ){\boldsymbol{x}}_{\boldsymbol{\theta}}({\boldsymbol{z}},{\boldsymbol{\xi}}) with ξ∼N(0,I){\boldsymbol{\xi}}\sim\mathcal{N}(0,{\bf I}). LVAExz(θ,ϕ,ψ1,ψ2)\mathcal{L}_{\rm VAExz}({\boldsymbol{\theta}},{\boldsymbol{\phi}},{\boldsymbol{\psi}}_{1},{\boldsymbol{\psi}}_{2}) in (10) is re-expressed as

The expectations in (11) are approximated via samples drawn from q(x)q({\boldsymbol{x}}) and p(z)p({\boldsymbol{z}}), as well as samples of ϵ{\boldsymbol{\epsilon}} and ξ{\boldsymbol{\xi}}. xθ(z,ξ){\boldsymbol{x}}_{\boldsymbol{\theta}}({\boldsymbol{z}},{\boldsymbol{\xi}}) and zϕ(x,ϵ){\boldsymbol{z}}_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{\epsilon}}) can be implemented with a Gaussian assumption or via density transformation , detailed when presenting experiments in Section 5.

The complete objective of the proposed Adversarial Symmetric VAE (AS-VAE) requires the cumulative variational in (11), which we maximize wrt ψ1{\boldsymbol{\psi}}_{1} and ψ1{\boldsymbol{\psi}}_{1} as in (5) and (6), using the results in (7). Hence, we write

The following proposition characterizes the solutions of (12) in terms of the joint distributions of x{\boldsymbol{x}} and z{\boldsymbol{z}}.

The equilibrium for the min-max objective in (12) is achieved by specification {θ∗,ϕ∗,ψ1∗,ψ2∗}\{{\boldsymbol{\theta}}^{*},{\boldsymbol{\phi}}^{*},{\boldsymbol{\psi}}_{1}^{*},{\boldsymbol{\psi}}_{2}^{*}\} if and only if (7) holds, and pθ∗(x,z)=qϕ∗(x,z)p_{{\boldsymbol{\theta}}^{*}}({\boldsymbol{x}},{\boldsymbol{z}})=q_{{\boldsymbol{\phi}}^{*}}({\boldsymbol{x}},{\boldsymbol{z}}).

The proof is provided in the Appendix A. This theoretical result implies that (ii) θ∗{\boldsymbol{\theta}}^{*} is an estimator that yields good reconstruction, and (iiii) ϕ∗{\boldsymbol{\phi}}^{*} matches the aggregated posterior qϕ(z)q_{\boldsymbol{\phi}}({\boldsymbol{z}}) to prior distribution p(z)p({\boldsymbol{z}}).

Related Work

VAEs represent one of the most successful deep generative models developed recently. Aided by the reparameterization trick, VAEs can be trained with stochastic gradient descent. The original VAEs implement a Gaussian assumption for the encoder. More recently, there has been a desire to remove this Gaussian assumption. Normalizing flow employs a sequence of invertible transformation to make the distribution of the latent codes arbitrarily flexible. This work was followed by inverse auto-regressive flow , which uses recurrent neural networks to make the latent codes more expressive. More recently, SteinVAE applies Stein variational gradient descent to infer the distribution of latent codes, discarding the assumption of a parametric form of posterior distribution for the latent code. However, these methods are not able to address the fundamental limitation of ML-based models, as they are all based on the variational formulation in (1).

Experiments

We evaluate our model on three datasets: MNIST, CIFAR-10 and ImageNet. To balance performance and computational cost, pθ(x∣z)p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}}) and qϕ(z∣x)q_{\boldsymbol{\phi}}({\boldsymbol{z}}|{\boldsymbol{x}}) are approximated with a normalizing flow of length 80 for the MNIST dataset, and a Gaussian approximation for CIFAR-10 and ImageNet data. All network architectures are provided in the Appendix B. All parameters were initialized with Xavier , and optimized via Adam with learning rate 0.0001. We do not perform any dataset-specific tuning or regularization other than dropout . Early stopping is employed based on the average reconstruction loss of x{\boldsymbol{x}} and z{\boldsymbol{z}} on validation sets.

We show three types of results, using part of or all of our model to illustrate each component. (ii) AS-VAE-r: This model is trained with the first half of the objective in (11) to maximize LVAEx(θ,ϕ)\mathcal{L}_{\rm VAEx}({\boldsymbol{\theta}},{\boldsymbol{\phi}}) in (8); it is an ML-based method which focuses on reconstruction. (iiii) AS-VAE-g: This model is trained with the second half of the objective in (11) to maximize LVAEz(θ,ϕ)\mathcal{L}_{\rm VAEz}({\boldsymbol{\theta}},{\boldsymbol{\phi}}) in (9); it can be considered as maximizing the likelihood of qϕ(z)q_{\boldsymbol{\phi}}({\boldsymbol{z}}), and designed for generation. (iiiiii) AS-VAE: This is our proposed model, developed in Section 3.

To the authors’ knowledge, we are the first to report both inception score (IS) and NLL for natural images from a single model. For comparison, we implemented DCGAN and PixelCNN++ as baselines. The implementation of DCGAN is based on a similar network architecture as our model. Note that for NLL a lower value is better, whereas for IS a higher value is better.

Appendix A Proof

We start from a simple observation pθ(x)=∫zpθ(x,z)dz=∫zp(z)pθ(x∣z)dzp_{\boldsymbol{\theta}}({\boldsymbol{x}})=\int_{\boldsymbol{z}}p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})d{\boldsymbol{z}}=\int_{\boldsymbol{z}}p({\boldsymbol{z}})p_{\boldsymbol{\theta}}({\boldsymbol{x}}|{\boldsymbol{z}})d{\boldsymbol{z}}. The second term in (5) of the main paper can be rewritten as

Therefore, the objective function LA1(ψ1)\mathcal{L}_{\rm A1}({{\boldsymbol{\psi}}_{1}}) in (5) can be expressed as

This integral of (17) is maximal as a function of fψ1(x,z)f_{{\boldsymbol{\psi}}_{1}}({\boldsymbol{x}},{\boldsymbol{z}}) if and only if the integrand is maximal for every (x,z)({\boldsymbol{x}},{\boldsymbol{z}}). Note that the problem max⁡xalog⁡x+blog⁡(1−x)\max_{x}a\log x+b\log(1-x) achieves maximum at x=a/(a+b)x=a/(a+b) and σ(x)=1/(1+e−x)\sigma(x)=1/(1+e^{-x}). Hence, we have the optimal function of fψ1f_{{\boldsymbol{\psi}}_{1}} at

Similarly, we have fψ2∗(x,z)=log⁡pθ(x,z)−log⁡qϕ(z)q(x)f_{{\boldsymbol{\psi}}_{2}^{*}}({\boldsymbol{x}},{\boldsymbol{z}})=\log p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})-\log q_{\boldsymbol{\phi}}({\boldsymbol{z}})q({\boldsymbol{x}}).

Proof of Proposition 1

Assume {θ∗,ϕ∗,ψ1∗,ψ2∗}\{{\boldsymbol{\theta}}^{*},{\boldsymbol{\phi}}^{*},{\boldsymbol{\psi}}_{1}^{*},{\boldsymbol{\psi}}_{2}^{*}\} achieves an equilibrium of (12) in the main paper. The Corollary 1.1 indicates that fψ1∗=log⁡qϕ(x,z)−log⁡pθ(x)p(z)f_{{\boldsymbol{\psi}}_{1}^{*}}=\log q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}})-\log p_{\boldsymbol{\theta}}({\boldsymbol{x}})p({\boldsymbol{z}}) and fψ2∗(x,z)=log⁡pθ(x,z)−log⁡qϕ(z)q(x)f_{{\boldsymbol{\psi}}_{2}^{*}}({\boldsymbol{x}},{\boldsymbol{z}})=\log p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})-\log q_{\boldsymbol{\phi}}({\boldsymbol{z}})q({\boldsymbol{x}}).

The minimum of the first two terms is achieved if and only if pθ(x,z)=qϕ(x,z)p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})=q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}}), while the minimum of the last two terms is achieved at pθ(x)=q(x)p_{\boldsymbol{\theta}}({\boldsymbol{x}})=q({\boldsymbol{x}}) and p(z)=qϕ(z)p({\boldsymbol{z}})=q_{\boldsymbol{\phi}}({\boldsymbol{z}}), respectively. Note that if the joint match pθ(x,z)=qϕ(x,z)p_{\boldsymbol{\theta}}({\boldsymbol{x}},{\boldsymbol{z}})=q_{\boldsymbol{\phi}}({\boldsymbol{x}},{\boldsymbol{z}}) is achieved, the marginals will also match, which indicates that the optimal (θ∗,ϕ∗)({\boldsymbol{\theta}}^{*},{\boldsymbol{\phi}}^{*}) is achieved if and only if pθ∗(x,z)=qϕ∗(x,z)p_{{\boldsymbol{\theta}}^{*}}({\boldsymbol{x}},{\boldsymbol{z}})=q_{{\boldsymbol{\phi}}^{*}}({\boldsymbol{x}},{\boldsymbol{z}}).

Appendix B Model Architecture

The model architectures are shown as following. For fψ1(x,z)f_{{\boldsymbol{\psi}}_{1}}({\boldsymbol{x}},{\boldsymbol{z}}) and fψ2(x,z)f_{{\boldsymbol{\psi}}_{2}}({\boldsymbol{x}},{\boldsymbol{z}}), we use the same architecture but the parameters are not shared.

Appendix C Additional Results