Your GAN is Secretly an Energy-based Model and You Should use Discriminator Driven Latent Sampling

Tong Che, Ruixiang Zhang, Jascha Sohl-Dickstein, Hugo Larochelle, Liam Paull, Yuan Cao, Yoshua Bengio

Introduction

Generative Adversarial Networks (GANs) are state-of-the-art models for a large variety of tasks such as image generation , semi-supervised learning , image editing , image translation , and imitation learning . The GAN framework consists of two neural networks, the generator GG and the discriminator DD. The optimization process is formulated as an adversarial game, with the generator trying to fool the discriminator and the discriminator trying to better classify samples as real or fake.

Despite the ability of GANs to generate high-resolution, sharp samples, the samples of GAN models sometimes contain bad artifacts or are even not recognizable . It is conjectured that this is due to the inherent difficulty of generating high dimensional complex data, such as natural images, and the optimization challenge of the adversarial formulation. In order to improve sample quality, conventional sampling techniques, such as increasing the temperature, are commonly adopted for GAN models . Recently, new sampling methods such as Discriminator Rejection Sampling (DRS) , Metropolis-Hastings Generative Adversarial Network (MH-GAN) , and Discriminator Optimal Transport (DOT) have shown promising results by utilizing the information provided by both the generator and the discriminator. However, these sampling techniques are either inefficient or lack theoretical guarantees, possibly reducing the sample diversity and making the mode dropping problem more severe.

In this paper, we show that GANs can be better understood through the lens of Energy-Based Models (EBM). In our formulation, GAN generators and discriminators collaboratively learn an “implicit” energy-based model. However, efficient sampling from this energy based model directly in pixel space is extremely challenging for several reasons. One is that there is no tractable closed form for the implicit energy function in pixel space. This motivates an intriguing possibility: that Markov Chain Monte Carlo (MCMC) sampling may prove more tractable in the GAN’s latent space.

Surprisingly, we find that the implicit energy based model defined jointly by a GAN generator and discriminator takes on a simpler, tractable form when it is written as an energy-based model over the generator’s latent space. In this way, we propose a theoretically grounded way of generating high quality samples from GANs through what we call Discriminator Driven Latent Sampling (DDLS). DDLS leverages the information contained in the discriminator to re-weight and correct the biases and errors in the generator. Through experiments, we show that our proposed method is highly efficient in terms of mixing time, is generally applicable to a variety of GAN models (e.g. Minimax, Non-Saturating, and Wasserstein GANs), and is robust across a wide range of hyper-parameters. An energy-based model similar to our work is also obtained simultaneously in independent work in the form of an approximate MLE lower bound.

We highlight our main contributions as follows:

We provide more evidence that it is beneficial to sample from the energy-based model defined both by the generator and the discriminator instead of from the generator only.

We derive an equivalent formulation of the pixel-space energy-based model in the latent space, and show that sampling is much more efficient in the latent space.

We show experimentally that samples from this energy-based model are of higher quality than samples from the generator alone.

We show that our method can approximately extend to other GAN formulations, such as Wasserstein GANs.

Background

GANs are a powerful class of generative models defined through an adversarial minimax game between a generator network GG and a discriminator network DD. The generator GG takes a latent code zz from a prior distribution p(z)p(z) and produces a sample G(z)∈XG(z)\in X. The discriminator takes a sample x∈Xx\in X as input and aims to classify real data from fake samples produced by the generator, while the generator is asked to fool the discriminator as much as possible. We use pdp_{d} to denote the true data-generating distribution and pgp_{g} to denote the implicit distribution induced by the prior and the generator network. The standard non-saturating training objective for the generator and discriminator is defined as:

2 Energy-Based Models and Langevin Dynamics

One solution to the problem of slow-sampling Markov Chains is to perform sampling using a carfefully crafted latent space . Our method shows how one can execute such latent space MCMC in GAN models.

Methodology

Suppose we have a GAN model trained on a data distribution pdp_{d} with a generator G(z)G(z) with generator distribution pgp_{g} and a discriminator D(x)D(x). We assume that pgp_{g} and pdp_{d} have the same support. This can be guaranteed by adding small Gaussian noise to these two distributions.

The training of GANs is an adversarial game which generally does not converge to the optimal generator, so usually pdp_{d} and pgp_{g} do not match perfectly at the end of training. However, the discriminator provides a quantitative estimate for how much these two distributions (mis)match. Let’s assume the discriminator is near optimality, namely D(x)≈pd(x)pd(x)+pg(x)D(x)\approx\frac{p_{d}(x)}{p_{d}(x)+p_{g}(x)}. From this equation, let d(x)d(x) be the logit of D(x)D(x), in which casepd(x)pd(x)+pg(x)=11+pg(x)pd(x)≈11+exp⁡(−d(x))\frac{p_{d}(x)}{p_{d}(x)+p_{g}(x)}=\frac{1}{1+\frac{p_{g}(x)}{p_{d}(x)}}\approx\frac{1}{1+\exp\left(-d(x)\right)}, and we have ed(x)≈pd/pge^{d(x)}\approx p_{d}/p_{g}, and pd(x)≈pg(x)ed(x)p_{d}(x)\approx p_{g}(x)e^{d(x)}. Normalization of pg(x)ed(x)p_{g}(x)e^{d(x)} is not guaranteed, and it will not typically be a valid probabilistic model. We therefore consider the energy-based model pd∗=pg(x)ed(x)/Z0p_{d}^{*}=p_{g}(x)e^{d(x)}/Z_{0}, where Z0Z_{0} is a normalization constant. Intuitively, this formulation has two desirable properties. First, as we elaborate later, if D=D∗D=D^{*} where D∗D^{*} is the optimal discriminator, then pd∗=pdp_{d}^{*}=p_{d}. Secondly, it corrects the bias in the generator via weighting and normalization. If we can sample from this distribution, it should improve our samples.

There are two difficulties in sampling efficiently from pd∗p_{d}^{*}:

Doing MCMC in pixel space to sample from the model is impractical due to the high dimensionality and long mixing time.

pg(x)p_{g}(x) is implicitly defined and its density cannot be computed directly.

In the next section we show how to overcome these two difficulties.

2 Rejection Sampling and MCMC in Latent Space

Our approach to the above two problems is to formulate an equivalent energy-based model in the latent space. To derive this formulation, we first review rejection sampling . With pgp_{g} as the proposal distribution, we have ed(x)/ Z0=pd∗(x) / pg(x)e^{d(x)}/\,Z_{0}=p^{*}_{d}(x)\,/\,p_{g}(x). Denote M=max⁡xpd∗(x)/ pg(x)M=\max_{x}p^{*}_{d}(x)/\,p_{g}(x) (this is well-defined if we add a Gaussian noise to the output of the generator and xx is in a compact space). If we accept samples from proposal distribution pgp_{g} with probability pd∗ / (Mpg)p^{*}_{d}\,/\,(Mp_{g}), then the samples we produce have the distribution pd∗p^{*}_{d}.

We can alternatively interpret the rejection sampling procedure above as occurring in the latent space zz. In this interpretation, we first sample zz from p(z)p(z), and then perform rejection sampling on zz with acceptance probability ed(G(z)) / (MZ0)e^{d(G(z))}\,/\,(MZ_{0}). Only once a latent sample zz has been accepted do we generate the pixel level sample x=G(z)x=G(z).

This rejection sampling procedure on zz induces a new probability distribution pt(z)p_{t}(z). To explicitly compute this distribution we need to conceptually reverse the definition of rejection sampling. We formally write down the “reverse” lemma of rejection sampling as Lemma 1, to be used in our main theorem.

Namely, we have the prior proposal distribution p0(z)p_{0}(z) and an acceptance probability r(z)=ed(G(z))/ (MZ0)r(z)=e^{d(G(z))}/\,(MZ_{0}). We want to compute the distribution after the rejection sampling procedure with r(z)r(z). With Lemma 1, we can see that pt(z)=p0(z)r(z) / Z′p_{t}(z)=p_{0}(z)r(z)\,/\,Z^{\prime}. We expand on the details in our main theorem.

3 Main Theorem

Assume pdp_{d} is the data generating distribution, and pgp_{g} is the generator distribution induced by the generator G:Z→XG:\mathcal{Z}\rightarrow\mathcal{X}, where Z\mathcal{Z} is the latent space with prior distribution p0(z)p_{0}(z). Define Boltzmann distribution pd∗=elog⁡pg(x)+d(x)/ Z0p_{d}^{*}=e^{\log p_{g}(x)+d(x)}/\,Z_{0}, where Z0Z_{0} is the normalization constant.

Assume pgp_{g} and pdp_{d} have the same support. We address the case when this assumption does not hold in Corollary 2. Further, let D(x)D(x) be the discriminator, and d(x)d(x) be the logit of DD, namely D(x)=σ(d(x))D(x)=\sigma\left(d(x)\right). We define the energy function E(z)=−log⁡p0(z)−d(G(z))E(z)=-\log p_{0}(z)-d(G(z)), and its Boltzmann distribution pt(z)=e−E(z)/ Zp_{t}(z)=e^{-E(z)}/\,Z. Then we have:

pd∗=pdp^{*}_{d}=p_{d} when DD is the optimal discriminator.

If we sample z∼ptz\sim p_{t}, and x=G(z)x=G(z), then we have x∼pd∗x\sim p^{*}_{d}. Namely, the induced probability measure G∘pt=pd∗G\circ p_{t}=p^{*}_{d}.

Interestingly, pt(z)p_{t}(z) has the form of an energy-based model, pt(z)=e−E(z)/ Z′p_{t}(z)=e^{-E(z)}/\,Z^{\prime}, with tractable energy function E(z)=−log⁡p0(z)−d(G(z))E(z)=-\log p_{0}(z)-d(G(z)). In order to sample from this Boltzmann distribution, one can use an MCMC sampler, such as Langevin dynamics or Hamiltonian Monte Carlo. We defer the proofs and the MCMC algorithm to our Supplemental Material.

4 Sampling Wasserstein GANs with Langevin Dynamics

Wasserstein GANs are different from original GANs in that they target the Wassertein loss. Although when the discriminator is trained to optimality, the discriminator can recover the Kantorovich dual of the optimal transport between pgp_{g} and pdp_{d}, the target distribution pdp_{d} cannot be exactly recovered using the information in pgp_{g} and DDIn Tanaka , the authors claim that it is possible to recover pdp_{d} with DD and pgp_{g} in WGAN in certain metrics, but we show in the Appendix that their assumptions don’t hold and in the L1L^{1} metric, which WGAN uses, it is not possible to recover pdp_{d}. . However, in the following we show that in practice, the optimization of WGAN can be viewed as an approximation of an energy-based model, which can also be sampled with our method.

The objectives of Wasserstein GANs can be summarized as:

where DD is restricted to be a KK-Lipschitz function.

On the other hand, consider a new energy-based generative model (which also has a generator and a discriminator) trained with the following objectives (for detailed algorithm, please refer to our Supplemental Material):

Discriminator training phase (D-phase). Unlike GANs, our energy-based model tries to match the distribution pt(x)=pg(x)eDϕ(x)/ Zp_{t}(x)=p_{g}(x)e^{D_{\phi}(x)}/\,Z with the data distribution pdp_{d}, where pt(x)p_{t}(x) can be interpreted as an EBM with energy Dϕ(x)−log⁡pg(x)D_{\phi}(x)-\log p_{g}(x). In this phase, the generator is kept fixed, and the discriminator is trained.

Generator training phase (G-phase). The generator is trained such that pg(x)p_{g}(x) matches pt(x)p_{t}(x), in this phase we treat DD as fixed and train GG.

In the D-phase, we are training an EBM with data from pdp_{d}. The gradient of the KL-divergence (which is our loss function for D-phase) can be written as :

Namely we are trying to maximize DD on real data and trying to minimize it on fake data. Note that the fake data distribution ptp_{t} is a function of both the generator and discriminator, and cannot be sampled directly. As with other energy-based models, we can use an MCMC procedure such as Langevin dynamics to generate samples from ptp_{t} .

In the G-phase, we can train the model with the gradient of KL-divergence KL(pg∣∣pt′)\text{KL}(p_{g}\mid\mid p_{t}^{\prime}) as our loss. Let pt′p_{t}^{\prime} be a fixed copy of ptp_{t}, we can compute the gradient as (see the Appendix for more details):

Note that the losses above coincide with what we are optimizing in WGANs, with two differences:

In WGAN, we optimize DD on pgp_{g} instead of ptp_{t}. This may not be a big difference in practice, since as training progresses ptp_{t} is expected to approach pgp_{g}, as the optimizing loss for the generator explicitly acts to bring pgp_{g} closer to ptp_{t} (Equation 4). Moreover, it has recently been found in LOGAN that optimizing DD on ptp_{t} rather than pgp_{g} can lead to better performance.

In WGAN, we impose a Lipschitz constraint on DD. This constraint can be viewed as a smoothness regularizer. Intuitively it will make the distribution pt(x)=pg(x)e−Dϕ(x)/ Zp_{t}(x)=p_{g}(x)e^{-D_{\phi}(x)}/\,Z more “flat” than pdp_{d}, but pt(x)p_{t}(x) (which lies in a distribution family parameterized by DD) remains an approximator to pdp_{d} subject to this constraint.

Thus, we can conclude that for a Wasserstein GAN with discriminator DD, WGAN approximately optimizes the KL divergence of pt=pg(x)e−D(x)/ Zp_{t}=p_{g}(x)e^{-D(x)}/\,Z with pdp_{d}, with the constraint that DD is KK-Lipschitz. This suggests that one can also perform DDLS on the WGAN latent space to generate improved samples, using an energy function E(z)=−log⁡p0(z)−D(G(z))E(z)=-\log p_{0}(z)-D(G(z)).

5 Practical Issues and the Mode Dropping Problem

Mode dropping is a major problem in training GANs. In our main theorem it is assumed that pgp_{g} and pdp_{d} have the same support. We also assumed that G:Z→XG:\mathcal{Z}\rightarrow\mathcal{X} is a deterministic function. Thus, if GG cannot recover some of the modes in pdp_{d}, pd∗p_{d}^{*} also cannot recover these modes.

However, we can partially solve the mode dropping problem by introducing an additional Gaussian noise z′∼N(0,1;z′)=p1(z′)z^{\prime}\sim N(0,1;z^{\prime})=p_{1}(z^{\prime}) to the output of the generator, namely we define the new deterministic generator G∗(z,z′)=G(z)+ϵz′G^{*}(z,z^{\prime})=G(z)+\epsilon z^{\prime}. We treat z′z^{\prime} as a part of the generator, and do DDLS on joint latent variables (z,z′)(z,z^{\prime}). The Langevin dynamics on this joint energy will help the model to move data points that are a little bit off-mode to the data manifold, and we have the follwing Corollary:

Assume pdp_{d} is the data generating distribution with small Gaussian noise added. The generator G:Z→XG:\mathcal{Z}\rightarrow\mathcal{X} is a deterministic function, where Z\mathcal{Z} is the latent space endowed with prior distribution p0(z)p_{0}(z). Assume z′∼p1(z′)=N(0,1;z)z^{\prime}\sim p_{1}(z^{\prime})=N(0,1;z) is an additional Gaussian noise variable with dim⁡z′=dim⁡X\dim z^{\prime}=\dim\mathcal{X}. Let ϵ>0\epsilon>0, denote the distribution of the extended generator G∗(z,z′)=G(z)+ϵz′G^{*}(z,z^{\prime})=G(z)+\epsilon z^{\prime} as pgp_{g}. D(x)D(x) is the discriminator trained between pgp_{g} and pdp_{d}. Let d(x)d(x) be the logit of DD, namely D(x)=σ(d(x))D(x)=\sigma\left(d(x)\right). Define pd∗=elog⁡pg(x)+d(x)/ Z0p_{d}^{*}=e^{\log p_{g}(x)+d(x)}/\,Z_{0}, where Z0Z_{0} is the normalization constant. We define the energy function in the extended latent space E(z,z′)=−log⁡p0(z)−log⁡p1(z′)−d(G∗(z,z′))E(z,z^{\prime})=-\log p_{0}(z)-\log p_{1}(z^{\prime})-d(G^{*}(z,z^{\prime})), and its Boltzmann distribution pt(z,z′)=e−E(z,z′)/ Zp_{t}(z,z^{\prime})=e^{-E(z,z^{\prime})}/\,Z. Then we have:

pd∗=pdp^{*}_{d}=p_{d} when DD is the optimal discriminator.

If we sample (z,z′)∼pt(z,z^{\prime})\sim p_{t}, and x=G∗(z,z′)x=G^{*}(z,z^{\prime}), then we have x∼pd∗x\sim p^{*}_{d}. Namely, the induced probability measure G∗∘pt=pd∗G^{*}\circ p_{t}=p^{*}_{d}.

Related Work

Previous work has considered utilizing the discriminator to achieve better sampling for GANs. Discriminator rejection sampling and Metropolis-Hastings GANs use pgp_{g} as the proposal distribution and DD as the criterion of acceptance or rejection. However, these methods are inefficient as they may need to reject a wechlot of samples. Intuitively, one major drawback of these methods is that since they operate in the pixel space, their algorithm can use discriminators to reject samples when they are bad, but cannot easily guide latent space updates using the discriminator which would improve these samples. The advantage of DDLS over DRS or MH-GAN is similar to the advantage of SGD over zero-th order optimization algorithms.

Trained classifiers have similarly been used to correct probabilistic models in other contexts . Discriminator optimal transport (DOT) is another way of sampling GANs. They use deterministic gradient descent in the latent space to get samples with higher DD-values, However, since pgp_{g} and DD cannot recover the data distribution exactly, DOT has to make the optimization local in a small neighborhood of generated samples, which hurts the sample performance. Also, DOT is not guaranteed to converge to the data distribution even under ideal assumptions (DD is optimal).

Other previous work considered the usage of probabilistic models defined jointly by the generator and discriminator. In , the authors use the idea of training an EBM defined jointly by a generator and an additional critic function in the text generation setting. uses an additional discriminator as a bias corrector for generative models via importance weighting. considered rejection sampling in latent space in encoder-decoder models.

Energy-based models have gained significant attention in recent years. Most work focuses on the maximum likelihood learning of energy-based models . Other work has built new connections between energy based models and classifiers . The primary difficulty in training energy-based models comes from effectively estimating and sampling the partition function. The contribution to training from the partition function can be estimated via MCMC , via training another generator network , or via surrogate objectives to maximum likelihood . The connection between GANs and EBMs has been studied by many authors . Our paper can be viewed as establishing a new connection between GANs and EBMs which allows efficient latent MCMC sampling.

Experimental results

In this section we present a set of experiments demonstrating the effectiveness of our method on both synthetic and real-world datasets. In section 5.1 we illustrate how the proposed method, DDLS, can improve the distribution modeling of a trained GAN and compare with other baseline methods. In section 5.2 we show that DDLS can improve the sample quality on real world datasets, both qualitatively and quantitatively.

Following the same setting used in , we apply DDLS to a WGAN model trained on two synthetic datasets, 25-gaussians and Swiss Roll, and investigate the effect and performance of the proposed sampling method.

Implementation details We follow the same synthetic experiments design as in DOT , while parameterizing the prior with a standard normal distribution instead of a uniform distribution. Please refer to C.1 for more details.

Qualitative results With the trained generator and discriminator, we generate 50005000 samples from the generator, then apply DDLS in latent space to obtain enhanced samples. We also apply the DOT method as a baseline. All results are depicted in Figure 2 and Figure 2 together with the target dataset samples. For the 25-Gaussian dataset we can see that DDLS recovered and preserved all modes while significantly eliminating spurious modes compared to a vanilla generator and DOT. For the Swiss Roll dataset we can also observe that DDLS successfully improved the distribution and recovered the underlying low-dimensional manifold of the data distribution. This qualitative evidence supports the hypothesis that our GANs as energy based model formulation outperforms the noisy implicit distribution induced by the generator alone.

Quantitative results We first examine the performance of DDLS quantitavely by using the metrics proposed by DRS . We generate 10,00010,000 samples with the DDLS algorithm, and each sample is assigned to its closest mixture component. A sample is of “high quality” if it is within four standard deviations of its assigned mixture component, and a mode is successfully “recovered” if at least one high-quality sample is assigned to it.

As shown in Table 1, our proposed model achieves a higher “high-quality” ratio. We also investigate the distance between the distribution induced by our GAN as EBM formulation and the true data distribution. We use the Earth Mover’s Distance (EMD) between the two corresponding empirical distributions as a surrogate, as proposed in DOT . As shown in Table 2, the EMD between our sampling distribution and the ground-truth distribution is significantly below the baselines. Note that we use our own re-implementation, and numbers differ slightly from those previously published.

2 CIFAR-10 and CelebA

In this section we evaluate the performance of the proposed DDLS method on the CIFAR-10 dataset and CelebA dataset.

Implementation details We provide detailed description of baseline models, DDLS hyper-parameters and evaluation protocol in C.2.

Quantitative results We evaluate the quality and diversity of generated samples via the Inception Score and Fréchet Inception Distance (FID) . In Table 3, we show the Inception score improvements from DDLS on CIFAR-10 and CelebA, compared to MH-GAN and DRS , following the same evaluation protocol and using the same baseline models (DCGAN and WGAN) in . On CIFAR-10, we applied DDLS to the unconditional generator of SN-GAN to generate 5000050000 samples and report all results in Table 3. We found that the proposed method significantly improves the Inception Score of the baseline SN-GAN model from 8.228.22 to 9.099.09 and reduces the FID from 21.721.7 to 15.4215.42. Our unconditional model outperforms previous state-of-the-art GANs and other sampling-enhanced GANs and even approaches the performance of conditional BigGANs which achieves an Inception Score 9.229.22 and an FID of 14.7314.73, without the need of additional class information, training and parameters.

Qualitative results We illustrate the process of Langevin dynamics sampling in latent space in Figure 4 by generating samples for every 1010 iterations. We find that our method helps correct the errors in the original generated image, and makes changes towards more semantically meaningful and sharp output by leveraging the pre-trained discriminator. We include more generated samples for visualizing the Langevin dynamics in the appendix. To demonstrate that our model is not simply memorizing the CIFAR-10 dataset, we find the nearest neighbors of generated samples in the training dataset and show the results in Figure 4.

Mixing time evaluation MCMC sampling methods often suffer from extremely long mixing times, especially for high-dimensional multi-modal data. For example, more than 600600 MCMC iterations are need to obtain the most performance gain in MH-GAN on real data. We demonstrate the sampling efficiency of our method by showing that we can expect a much shorter time to achieve competitive performance by migrating the Langevin sampling process to the latent space, compared to sampling in high-dimensional multi-modal pixel space. We evaluate the Inception Score and the energy function for every 1010 iterations of Langevin dynamics and depict the results in Figure 5 in Appendix.

3 ImageNet

In this section we evaluate the performance of the proposed DDLS method on the ImageNet dataset.

Implementation details As with CIFAR-10, we adopt the Spectral Normalization GAN (SN-GAN) as our baseline GAN model. We take the publicly available pre-trained models of SN-GAN and apply DDLS. Fine tuning is performed on the discriminator, as described in Section 5.2. Implementation choices are otherwise the same as for CIFAR-10, with additional details in the appendix. We show the quantitative results in Table 4, where we substantially outperform the baseline.

Conclusion and Future Work

In this paper, we have shown that a GAN’s discriminator can enable better modeling of the data distribution with Discriminator Driven Latent Sampling (DDLS). The intuition behind our model is that learning a generative model to do structured generative prediction is usually more difficult than learning a classifier, so the errors made by the generator can be significantly corrected by the discriminator. The major advantage of DDLS is that it allows MCMC sampling in the latent space, which enables efficient sampling and better mixing.

For future work, we are exploring the inclusion of additional Gaussian noise variables in each layer of the generator, treated as latent variables, such that DDLS can be used to provide a correcting signal for each layer of the generator. We believe that this will lead to further sampling improvements. Also, prior work on VAEs has shown that learned re-sampling or energy-based transformation of their priors can be effective . It would thus be particularly interesting to explore whether VAE-based models can be improved by constructing an energy based model for their prior based on an auxiliary discriminator.

Broader Impact

The development of powerful generative models which can generate fake images, audio, and video which appears realistic is a source of acute concern , and can enable fake news or propaganda. On the other hand these same technologies also enable assistive technologies like text to speech , new art forms , and new design technologies .

Our work enables the creation of more powerful generative models. As such it is multi-use, and may result in both positive and negative societal consequences. However, we believe that improving scientific understanding tends also to improve the human condition – so in the absence of a reason to expect specific harm, we believe that in expectation our work will have a positive impact on the world.

One potential benefit of our work is that it leads to better calibrated GANs, which are less likely to simply drop under-represented sample classes from their generated output. The tendency of machine learning models to produce worse outcomes for groups which are underrepresented in their training data (in terms of race, gender, or otherwise) is well documented . The use of our technique should produce generative models which are slightly less prone to this type of bias.

Acknowledgments and Disclosure of Funding

The authors are grateful to Ben Poole and the anonymous reviewers for proof-reading the paper and suggesting improvements. This work was supported by CIFAR and Compute Canada.

References

Appendix A Proofs

On space XX there is a probability distribution p(x)p(x). r(x):X→r(x):X\rightarrow is a measurable function on XX. We consider sampling from pp, accepting with probability r(x)r(x), and repeating this procedure until a sample is accepted. We denote the resulting probability measure of the accepted samples q(x)q(x). Then we have:

From the definition of rejection sampling, we can see that in order to get the distribution q(x)q(x), we can sample xx from p(x)p(x) and do rejection sampling with probability r′(x)=q(x) / (Mp(x))r^{\prime}(x)=q(x)\,/\,(Mp(x)), where M≥q(x)/ p(x)M\geq q(x)/\,p(x) for all xx. So we have r′(x)=r(x) / (ZM)r^{\prime}(x)=r(x)\,/\,(ZM). If we choose M=1/ ZM=1/\,Z, then from r(x)≤1r(x)\leq 1 for all xx, we can see that MM satisfies M≥q(x)/ p(x)=r(x) / ZM\geq q(x)/\,p(x)=r(x)\,/\,Z, for all xx. So we can choose M=1 / ZM=1\,/\,Z, resulting in r(x)=r′(x)r(x)=r^{\prime}(x). ∎

Assume pdp_{d} is the data generating distribution, and pgp_{g} is the generator distribution induced by the generator G:Z→XG:\mathcal{Z}\rightarrow\mathcal{X}, where Z\mathcal{Z} is the latent space with prior distribution p0(z)p_{0}(z). Define pd∗=elog⁡pg(x)+d(x)/ Z0p_{d}^{*}=e^{\log p_{g}(x)+d(x)}/\,Z_{0}, where Z0Z_{0} is the normalization constant.

Assume pgp_{g} and pdp_{d} have the same support. This assumption is typically satisfied when dim⁡(z)≥dim⁡(x)\operatorname{dim}(z)\geq\operatorname{dim}(x). We address the case that dim⁡(z)<dim⁡(x)\operatorname{dim}(z)<\operatorname{dim}(x) in Corollary 2. Further, let D(x)D(x) be the discriminator, and d(x)d(x) be the logit of DD, namely D(x)=σ(d(x))D(x)=\sigma\left(d(x)\right). We define the energy function E(z)=−log⁡p0(z)−d(G(z))E(z)=-\log p_{0}(z)-d(G(z)), and its Boltzmann distribution pt(z)=e−E(z)/ Zp_{t}(z)=e^{-E(z)}/\,Z. Then we have:

pd∗=pdp^{*}_{d}=p_{d} when DD is the optimal discriminator.

If we sample z∼ptz\sim p_{t}, and x=G(z)x=G(z), then we have x∼pd∗x\sim p^{*}_{d}. Namely, the induced probability measure G∘pt=pd∗G\circ p_{t}=p^{*}_{d}.

(1) follows from the fact that when DD is optimal, D(x)=pgpd+pgD(x)=\frac{p_{g}}{p_{d}+p_{g}}, so D(x)=σ(log⁡pd−log⁡pg)D(x)=\sigma(\log p_{d}-\log p_{g}), which implies that d(x)=log⁡pd−log⁡pgd(x)=\log p_{d}-\log p_{g} (which is finite on the support of pgp_{g} due to the fact that they have the same support). Thus, pd∗(x)=pd(x)/ Z0p_{d}^{*}(x)=p_{d}(x)/\,Z_{0}, we must have Z0=1Z_{0}=1 for normalization, so pd∗=pdp_{d}^{*}=p_{d}.

For (2), for samples x∼pgx\sim p_{g}, if we do rejection sampling with probability pd∗(x)/ (Mpg(x))=ed(x)/ (MZ0)p^{*}_{d}(x)/\,(Mp_{g}(x))=e^{d(x)}/\,(MZ_{0}) (where MM is a constant with M≥pd∗(x)/ pg(x)M\geq p^{*}_{d}(x)/\,p_{g}(x)), we get samples from the distribution pd∗p_{d}^{*}. We can view this rejection sampling as a rejection sampling in the latent space Z\mathcal{Z}, where we perform rejection sampling on p0(z)p_{0}(z) with acceptance probability r(z)=pd∗(G(z))/ (Mpg(G(z)))=ed(G(z))/ Mr(z)=p_{d}^{*}(G(z))/\,(Mp_{g}(G(z)))=e^{d(G(z))}/\,M. Applying lemma 1, we see that this rejection sampling procedure induces a probability distribution pt(z)=p0(z)r(z)/Cp_{t}(z)=p_{0}(z)r(z)/C on the latent space Z\mathcal{Z}. CC is the normalization constant. Thus sampling from pd∗(x)p_{d}^{*}(x) is equivalent to sampling from pt(z)p_{t}(z) and generating with G(z)G(z). ∎

Assume pdp_{d} is the data generating distribution with small Gaussian noise added. The generator G:Z→XG:\mathcal{Z}\rightarrow\mathcal{X} is a deterministic function, where Z\mathcal{Z} is the latent space endowed with prior distribution p0(z)p_{0}(z). Assume z′∼p1(z′)=N(0,1;z)z^{\prime}\sim p_{1}(z^{\prime})=N(0,1;z) is an additional Gaussian noise variable with dim⁡z′=dim⁡X\dim z^{\prime}=\dim\mathcal{X}. Let ϵ>0\epsilon>0, denote the distribution of the extended generator G∗(z,z′)=G(z)+ϵz′G^{*}(z,z^{\prime})=G(z)+\epsilon z^{\prime} as pgp_{g}. D(x)D(x) is the discriminator trained between pgp_{g} and pdp_{d}. Let d(x)d(x) be the logit of DD, namely D(x)=σ(d(x))D(x)=\sigma\left(d(x)\right). Define pd∗=elog⁡pg(x)+d(x)/ Z0p_{d}^{*}=e^{\log p_{g}(x)+d(x)}/\,Z_{0}, where Z0Z_{0} is the normalization constant. We define the energy function in the extended latent space E(z,z′)=−log⁡p0(z)−log⁡p1(z′)−d(G∗(z,z′))E(z,z^{\prime})=-\log p_{0}(z)-\log p_{1}(z^{\prime})-d(G^{*}(z,z^{\prime})), and its Boltzmann distribution pt(z,z′)=e−E(z,z′)/ Zp_{t}(z,z^{\prime})=e^{-E(z,z^{\prime})}/\,Z. Then we have:

pd∗=pdp^{*}_{d}=p_{d} when DD is the optimal discriminator.

If we sample (z,z′)∼pt(z,z^{\prime})\sim p_{t}, and x=G∗(z,z′)x=G^{*}(z,z^{\prime}), then we have x∼pd∗x\sim p^{*}_{d}. Namely, the induced probability measure G∗∘pt=pd∗G^{*}\circ p_{t}=p^{*}_{d}.

Let G∗(z,z′)G^{*}(z,z^{\prime}) be the generator GG defined in Theorem 1, we can see that pdp_{d} and pgp_{g} have the same support. Apply Theorem 1 and we deduce the corollary. ∎

Appendix B An Analysis of WGAN

In this section, we first give an example that in WGAN, given the optimal discriminator DD and pgp_{g}, it is not possible to recover pdp_{d}.

Consider the following case: the underlying space is one dimensional space of real numbers R\mathbf{R}. pgp_{g} is the Dirac δ\delta-distribution δ−1\delta_{-1} and data distribution pdp_{d} is the Dirac δ\delta-distribution δa\delta_{a}, where a>0a>0 is a constant.

We can easily identity function f(x)=xf(x)=x is the optimal 11-Lipschitz function which separates pgp_{g} and pdp_{d}. Namely, we let D(x)=xD(x)=x is the optimal discriminator.

However, DD is not a function of aa. Namely, we cannot recover pd=δap_{d}=\delta_{a} with information provided by DD and pgp_{g}. This is the main reason that collaborative sampling algorithms based on W-GAN formulation such as DOT could not provide exact theoretical guarantee, even if the discriminator is optimal.

B.2 Mathematical Details of Approximating WGAN with EBMs

In the paper, we show that the optimization of WGAN can be viewed as an approximation of an energy-based model. We present more details here.

Appendix C Experimental details

Source code of all experiments of this work is included in the supplemental material , where all detailed hyper-parameters can be found.

The 25-Gaussians dataset is generated by a mixture of twenty-five two-dimensional isotropic Gaussian distributions with variance 0.01, and means separated by 1, arranged in a grid. The Swiss Roll dataset is a standard dataset for testing dimensionality reduction algorithms. We use the implementation from scikit-learn, and rescale the coordinates as suggested by . We train a Wasserstein GAN model with the standard WGAN-GP objective. Both the generator and discriminator are fully connected neural networks with ReLU nonlinearities, and we follow the same architecture design as in DOT , while parameterizing the prior with a standard normal distribution instead of a uniform distribution. We optimize the model using the Adam optimizer, with α=0.0001,β1=0.5,β2=0.9\alpha=0.0001,\beta_{1}=0.5,\beta_{2}=0.9.

C.2 CIFAR-10 and CelebA

For CIFAR-10 dataset, we adopt the Spectral Normalization GAN (SN-GAN) as our baseline GAN model. We take the publicly available pre-trained models of unconditional SN-GAN and apply DDLS. For CelebA dataset, we adopt DCGAN and WGAN as the baseline model following the same setting in . We first sample latent codes from the prior distribution, then run the Langevin dynamics procedure with an initial step size 0.010.01 up to 10001000 iterations to generate enhanced samples. Following the practice in we separately set the standard deviation of the Gaussian noise as 0.10.1. We optionally fine-tune the pre-trained discriminator with an additional fully-connected layer and a logistic output layer using the binary cross-entropy loss to calibrate the discriminator as suggested by .

We show more generated samples of DDLS during langevin dynamics in Fig. 6. We run 10001000 steps of Langevin dynamics and plot generated sample for every 1010 iterations. We include 1000010000 more randomly generated samples in the supplemental material.

C.3 Imagenet

We introduce more details of the preliminary experimental results on Imagenet dataset here. We run the Langevin dynamics sampling algorithm with an initial step size 0.010.01 up to 10001000 iterations. We decay the step size with a factor 0.10.1 for every 200200 iterations. The standard deviation of Gaussian noise is annealed simultaneously with the step size. The discriminator is not yet calibrated in this preliminary experiment.

Appendix D DDLS Algorithm

We show the detailed algorithm using Langevin dynamics in Alg. 1.

Appendix E Hybrid WGAN-EBM Training Algorithm

In Sec. 3.4, we described an EBM algorithm which WGAN is approximately optimizing. Here we detail this algorithm in Alg. 2.