Score-based Generative Modeling in Latent Space
Arash Vahdat, Karsten Kreis, Jan Kautz
Introduction
The long-standing goal of likelihood-based generative learning is to faithfully learn a data distribution, while also generating high-quality samples. Achieving these two goals simultaneously is a tremendous challenge, which has led to the development of a plethora of different generative models. Recently, score-based generative models (SGMs) demonstrated astonishing results in terms of both high sample quality and likelihood . These models define a forward diffusion process that maps data to noise by gradually perturbing the input data. Generation corresponds to a reverse process that synthesizes novel data via iterative denoising, starting from random noise. The problem then reduces to learning the score function—the gradient of the log-density—of the perturbed data . In a seminal work, Song et al. show how this modeling approach is described with a stochastic differential equation (SDE) framework which can be converted to maximum likelihood training . Variants of SGMs have been applied to images , audio , graphs and point clouds .
Albeit high quality, sampling from SGMs is computationally expensive. This is because generation amounts to solving a complex SDE, or equivalently ordinary differential equation (ODE) (denoted as the probability flow ODE in ), that maps a simple base distribution to the complex data distribution. The resulting differential equations are typically complex and solving them accurately requires numerical integration with very small step sizes, which results in thousands of neural network evaluations . Furthermore, generation complexity is uniquely defined by the underlying data distribution and the forward SDE for data perturbation, implying that synthesis speed cannot be increased easily without sacrifices. Moreover, SDE-based generative models are currently defined for continuous data and cannot be applied effortlessly to binary, categorical, or graph-structured data.
Here, we propose the Latent Score-based Generative Model (LSGM), a new approach for learning SGMs in latent space, leveraging a variational autoencoder (VAE) framework . We map the input data to latent space and apply the score-based generative model there. The score-based model is then tasked with modeling the distribution over the embeddings of the data set. Novel data synthesis is achieved by first generating embeddings via drawing from a simple base distribution followed by iterative denoising, and then transforming this embedding via a decoder to data space (see Fig. 1). We can consider this model a VAE with an SGM prior. Our approach has several key advantages:
Synthesis Speed: By pretraining the VAE with a Normal prior first, we can bring the marginal distribution over encodings (the aggregate posterior) close to the Normal prior, which is also the SGM’s base distribution. Consequently, the SGM only needs to model the remaining mismatch, resulting in a less complex model from which sampling becomes easier. Furthermore, we can tailor the latent space according to our needs. For example, we can use hierarchical latent variables and apply the diffusion model only over a subset of them, further improving synthesis speed.
Expressivity: Training a regular SGM can be considered as training a neural ODE directly on the data . However, previous works found that augmenting neural ODEs and more generally generative models with latent variables improves their expressivity. Consequently, we expect similar performance gains from combining SGMs with a latent variable framework.
Tailored Encoders and Decoders: Since we use the SGM in latent space, we can utilize carefully designed encoders and decoders mapping between latent and data space, further improving expressivity. Additionally, the LSGM method can therefore be naturally applied to non-continuous data.
LSGMs can be trained end-to-end by maximizing the variational lower bound on the data likelihood. Compared to regular score matching, our approach comes with additional challenges, since both the score-based denoising model and its target distribution, formed by the latent space encodings, are learnt simultaneously. To this end, we make the following technical contributions: (i) We derive a new denoising score matching objective that allows us to efficiently learn the VAE model and the latent SGM prior at the same time. (ii) We introduce a new parameterization of the latent space score function, which mixes a Normal distribution with a learnable SGM, allowing the SGM to model only the mismatch between the distribution of latent variables and the Normal prior. (iii) We propose techniques for variance reduction of the training objective by designing a new SDE and by analytically deriving importance sampling schemes, allowing us to stably train deep LSGMs. Experimentally, we achieve state-of-the-art 2.10 FID on CIFAR-10 and 7.22 FID on CelebA-HQ-256, and significantly improve upon likelihoods of previous SGMs. On CelebA-HQ-256, we outperform previous SGMs in synthesis speed by two orders of magnitude. We also model binarized images, MNIST and OMNIGLOT, achieving state-of-the-art likelihood on the latter.
Background
Here, we review continuous-time score-based generative models (see for an in-depth discussion). Consider a forward diffusion process for continuous time variable , where is the starting variable and its perturbation at time . The diffusion process is defined by an Itô SDE:
that trains the parameteric score function at time for a given weighting coefficient . is the -generating distribution and is the diffusion kernel, which is available in closed form for certain and . Since is not analytically available, Song et al. rely on denoising score matching that converts the objective in Eq. 2 to:
which can again be transformed into denoising score matching (Eq. 3) following Vincent .
Score-based Generative Modeling in Latent Space
The LSGM framework in Fig. 1 consists of the encoder , SGM prior , and decoder . The SGM prior leverages a diffusion process as defined in Eq. 1 and diffuses samples in latent space to the standard Normal distribution . Generation uses the reverse SDE to sample from with time-dependent score function , and the decoder to map the synthesized encodings to data space. Formally, the generative process is written as . The goal of training is to learn , the parameters of the encoder , score function , and decoder , respectively.
We train LSGM by minimizing the variational upper bound on negative data log-likelihood :
following a VAE approach , where approximates the true posterior .
In this paper, we use Eq. 6 with decomposed KL divergence into its entropy and cross entropy terms. The reconstruction and entropy terms are estimated easily for any explicit encoder as long as the reparameterization trick is available . The challenging part in training LSGM is to train the cross entropy term that involves the SGM prior. We motivate and present our expression for the cross-entropy term in Sec. 3.1, the parameterization of the SGM prior in Sec. 3.2, different weighting mechanisms for the training objective in Sec. 3.3, and variance reduction techniques in Sec. 3.4.
One may ask, why not train LSGM with Eq. 5 and rely on the KL in Eq. 4. Directly using the KL expression in Eq. 4 is not possible, as it involves the marginal score , which is unavailable analytically for common non-Normal distributions such as Normalizing flows. Transforming into denoising score matching does not help either, since in that case the problematic term appears in the term (see Eq. 3). In contrast to previous works , we cannot simply drop , since it is, in fact, not constant but depends on , which is trainable in our setup.
To circumvent this problem, we instead decompose the KL in Eq. 5 and rather work directly with the cross entropy between the encoder distribution and the SGM prior . We show:
with and a Normal transition kernel , where and are obtained from and for a fixed initial variance at .
A proof with generic expressions for and as well as an intuitive interpretation are in App. A.
Importantly, unlike for the KL objective of Eq. 4, no problematic terms depending on the marginal score arise. This allows us to use this denoising score matching objective for the cross entropy term in Theorem 1 not only for optimizing (which is commonly done in the score matching literature), but also for the encoding distribution. It can be used even with complex distributions, defined, for example, in a hierarchical fashion or via Normalizing flows . Our novel analysis shows that, for diffusion SDEs following Eq. 1, only the cross entropy can be expressed purely with . Neither KL nor entropy in can be expressed without the problematic term (details in the Appendix).
Note that in Theorem 1, the term in the score matching expression corresponds to the score that originates from diffusing an initial distribution. In practice, we use the expression to learn an SGM prior , which models by a neural network. With the learnt score (here we explicitly indicate the parameters to clarify that this is the learnt model), the actual SGM prior is defined via the generative reverse-time SDE (or, alternatively, a closely-connected ODE, see Sec. 2 and App. D), which generally defines its own, separate marginal distribution at . Importantly, the learnt, approximate score is not necessarily the same as one would obtain when diffusing . Hence, when considering the learnt score , the score matching expression in our Theorem only corresponds to an upper bound on the cross entropy between and defined by the generative reverse-time SDE. This is discussed in detail in concurrent works . Hence, from the perspective of the learnt SGM prior, we are training with an upper bound on the cross entropy (similar to the bound on the KL in Eq. 4), which can also be considered as the continuous version of the discretized variational objective derived by Ho et al. .
2 Mixing Normal and Neural Score Functions
In VAEs , is often chosen as a standard Normal . For recent hierarchical VAEs , using the reparameterization trick, the prior can be converted to (App. E).
Considering a single dimensional latent space, we can assume that the prior at time is in the form of a geometric mixture where is a trainable SGM prior and is a learnable scalar mixing coefficient. Formulating the prior this way has crucial advantages: (i) We can pretrain LSGM’s autoencoder networks assuming , which corresponds to training the VAE with a standard Normal prior. This pretraining step will bring the distribution of latent variable close to , allowing the SGM prior to learn a much simpler distribution in the following end-to-end training stage. (ii) The score function for this mixture is of the form . When the score function is dominated by the linear term, we expect that the reverse SDE can be solved faster, as its drift is dominated by this linear term.
For our multivariate latent space, we obtain diffused samples at time by sampling with , where . Since we have , similar to , we parameterize the score function by , where is defined by our mixed score parameterization that is applied elementwise to the components of the score. With this, we simplify the cross entropy expression to:
where is a time-dependent weighting scalar.
3 Training with Different Weighting Mechanisms
The weighting term in Eq. 7 trains the prior with maximum likelihood. Similar to , we observe that when is dropped while training the SGM prior (i.e., ), LSGM often yields higher quality samples at a small cost in likelihood. However, in our case, we can only drop the weighting when training the prior. When updating the encoder parameters, we still need to use the maximum likelihood weighting to ensure that the encoder is brought closer to the true posterior Minimizing w.r.t is equivalent to minimizing w.r.t .. Tab. 1 summarizes three weighting mechanisms we consider in this paper: corresponds to maximum likelihood, is the unweighted objective used by , and is a variant obtained by dropping only . This weighting mechanism has a similar affect on the sample quality as ; however, in Sec. 3.4, we show that it is easier to define a variance reduction scheme for this weighting mechanism.
The following summarizes our training objectives (with and ):
where Eq. 8 trains the VAE encoder and decoder parameters using the variational bound from Eq. 6. Eq. 9 trains the prior with one of the three weighting mechanisms. Since the SGM prior participates in the objective only in the cross entropy term, we only consider this term when training the prior. Efficient algorithms for training with the objectives are presented in App. G.
4 Variance Reduction
The objectives in Eqs. 8 and 9 involve sampling of the time variable , which has high variance . We introduce several techniques for reducing this variance for all three objective weightings. We focus on the “variance preserving” SDEs (VPSDEs) , defined by where linearly interpolates in (other SDEs discussed in App. B).
(1) Geometric VPSDE: To reduce the variance sampling uniformly from , we can design the SDE such that is constant for . We show in App. B that a with geometric variance satisfies this condition. We call a VPSDE with this a geometric VPSDE. and are the hyperparameters of the SDE, with . Although our geometric VPSDE has a geometric variance progression similar to the “variance exploding” SDE (VESDE) , it still enjoys the “variance preserving” property of the VPSDE. In App. B, we show that the VESDE does not come with a reduced variance for -sampling by default.
(2) Importance sampling (IS): We can keep and unchanged for the original linear VPSDE, and instead use IS to minimize variance. The theory of IS shows that the proposal has minimum variance . In App. B, we show that we can sample from using inverse transform sampling where is the inverse of and . This variance reduction technique is available for any VPSDE with arbitrary .
In Fig. 3, we train a small LSGM on CIFAR-10 with weighting using (i) the original VPSDE with uniform sampling, (ii) the same SDE but with our IS from , and (iii) the proposed geometric VPSDE. Note how both (ii) and (iii) significantly reduce the variance and allow us to monitor the progress of the training objective. In this case, (i) has difficulty minimizing the objective due to the high variance. In App. B, we show how IS proposals can be formed for other SDEs, including the VESDE and Sub-VPSDE from .
Variance reduction for unweighted and reweighted objectives: When training with , analytically deriving IS proposal distributions for arbitrary is challenging. For linear VPSDEs, we provide a derivation in App. B to obtain the optimal IS distribution. In contrast, defining IS proposal distributions is easier when training with . In App. B, we show that the optimal distribution is in the form which is sampled by with . In Fig. 3, we visualize the IS distributions for the three weighting mechanisms for the linear VPSDE with the original parameters from . for the likelihood weighting is more tilted towards due to the term in .
When using differently weighted objectives for training, we can either sample separate with different IS distributions for each objective, or use IS for the SGM objective (Eq. 9) and reweight the samples according to the likelihood objective for encoder training (Eq. 8). See App. G for details.
Related Work
Our work builds on score-matching , specifically denoising score matching , which makes our work related to recent generative models using denoising score matching- and denoising diffusion-based objectives . Among those, use a discretized diffusion process with many noise scales, building on , while Song et al. introduce the continuous time framework using SDEs. Experimentally, these works focus on image modeling and, contrary to us, work directly in pixel space. Various works recently tried to address the slow sampling of these types of models and further improve output quality. add an adversarial objective, introduce non-Markovian diffusion processes that allow to trade off synthesis speed, quality, and sample diversity, learn a sequence of conditional energy-based models for denoising, distill the iterative sampling process into single shot synthesis, and learn an adaptive noise schedule, which is adjusted during synthesis to accelerate sampling. Further, propose empirical variance reduction techniques for discretized diffusions and introduce a new, heuristically motivated, noise schedule. In contrast, our proposed noise schedule and our variance reduction techniques are analytically derived and directly tailored to our learning setting in the continuous time setup.
Recently, presented a method to generate graphs using score-based models, relaxing the entries of adjacency matrices to continuous values. LSGM would allow to model graph data more naturally using encoders and decoders tailored to graphs .
Since our model can be considered a VAE with score-based prior, it is related to approaches that improve VAE priors. For example, Normalizing flows and hierarchical distributions , as well as energy-based models have been proposed as VAE priors. Furthermore, classifiers , adversarial methods , and other techniques have been used to define prior distributions implicitly. In two-stage training, a separate generative model is trained in latent space as a new prior after training the VAE itself . Our work also bears a resemblance to recent methods on improving the sampling quality in generative adversarial networks using gradient flows in the latent space , with the main difference that these prior works use a discriminator to update the latent variables, whereas we train an SGM.
Concurrent works: proposed to learn a denoising diffusion model in the latent space of a VAE for symbolic music generation. This work does not introduce an end-to-end training framework of the combined VAE and denoising diffusion model and instead trains them in two separate stages. In contrast, concurrently with us proposed an end-to-end training approach, and combines contrastive learning with diffusion models in the latent space of VAEs for controllable generation. However, consider the discretized diffusion objective , while we build on the continuous time framework. Also, these models are not equipped with the mixed score parameterization and variance reduction techniques, which we found crucial for the successful training of SGM priors.
Additionally, concurrently with us proposed likelihood-based training of SGMs in data spaceWe build on the V1 version of , which was substantially updated after the NeurIPS submission deadline.. developed a bound for the data likelihood in their Theorem 3 of their second version, using a denoising score matching objective, closely related to our cross entropy expression. However, our cross entropy expression is much simpler as we show how several terms can be marginalized out analytically for the diffusion SDEs employed by us (see our proof in App. A). The same marginalization can be applied to Theorem 3 in when the drift coefficient takes a special affine form (i.e., ). Moreover, discusses the likelihood-based training of SGMs from a fundamental perspective and shows how several score matching objectives become a variational bound on the data likelihood. introduced a notion of signal-to-noise ratio (SNR) that results in a noise-invariant parameterization of time that depends only on the initial and final noise. Interestingly, our importance sampling distribution in Sec. 3.4 has a similar noise-invariant parameterization of time via , which also depends only on the initial and final diffusion process variances. We additionally show that this time parameterization results in the optimal minimum-variance objective, if the distribution of latent variables follows a standard Normal distribution. Finally, proposed a modified time parameterization that allows modeling unbounded data scores.
Experiments
Here, we examine the efficacy of LSGM in learning generative models for images.
Implementation details: We implement LSGM using the NVAE architecture as VAE backbone and NCSN++ as SGM backbone. NVAE has a hierarchical latent structure. The diffusion process input is constructed by concatenating the latent variables from all groups in the channel dimension. For NVAEs with multiple spatial resolutions in latent groups, we only feed the smallest resolution groups to the SGM prior and assume that the remaining groups have a standard Normal distribution.
Sampling: To generate samples from LSGM at test time, we use a black-box ODE solver to sample from the prior. Prior samples are then passed to the decoder to generate samples in data space.
Evaluation: We measure NELBO, an upper bound on negative log-likelihood (NLL), using Eq. 6. For estimating , we rely on the probability flow ODE , which provides an unbiased but stochastic estimation of . This stochasticity prevents us from performing an importance weighted estimation of NLL (see App. F for details). For measuring sample quality, Fréchet inception distance (FID) is evaluated with 50K samples. Implementation details in App. G.
Unconditional color image generation: Here, we present our main results for unconditional image generation on CIFAR-10 (Tab. 3) and CelebA-HQ-256 (5-bit quantized) (Tab. 3). For CIFAR-10, we train 3 different models: LSGM (FID) and LSGM (balanced) both use the VPSDE with linear and -weighting for the SGM prior in Eq. 9, while performing IS as derived in Sec. 3.4. They only differ in how the backbone VAE is trained. LSGM (NLL) is a model that is trained with our novel geometric VPSDE, using -weighting in the prior objective (further details in App. G). When set up for high image quality, LSGM achieves a new state-of-the-art FID of 2.10. When tuned towards NLL, we achieve a NELBO of , which is significantly better than previous score-based models. Only autoregressive models, which come with very slow synthesis, and VDVAE reach similar or higher likelihoods, but they usually have much poorer image quality.
For CelebA-HQ-256, we observe that when LSGM is trained with different SDE types and weighting mechanisms, it often obtains similar NELBO potentially due to applying the SGM prior only to small latent variable groups and using Normal priors at the larger groups. With -weighting and linear VPSDE, LSGM obtains the state-of-the-art FID score of 7.22 on a par with the original SGM .
For both datasets, we also report results for the VAE backbone used in our LSGM. Although this baseline achieves competitive NLL, its sample quality is behind our LSGM and the original SGM.
Modeling binarized images: Next, we examine LSGM on dynamically binarized MNIST and OMNIGLOT . We apply LSGM to binary images using a decoder with pixel-wise independent Bernoulli distributions. For these datasets, we report both NELBO and NLL in nats in Tab. 5 and Tab. 5. On OMNIGLOT, LSGM achieves state-of-the-art likelihood of 87.79 nat, outperforming previous models including VAEs with autoregressive decoders, and even when comparing its NELBO against importance weighted estimation of NLL for other methods. On MNIST, LSGM outperforms previous VAEs in NELBO, reaching a NELBO 1.09 nat lower than the state-of-the-art NVAE.
Qualitative results: We visualize qualitative results for all datasets in Fig. 5. On the complex multimodal CIFAR-10 dataset, LSGM generates sharp and high-quality images. On CelebA-HQ-256, LSGM generates diverse samples from different ethnicity and age groups with varying head poses and facial expressions. On MNIST and OMNIGLOT, the generated characters are sharp and high-contrast.
Sampling time: We compare LSGM against the original SGM trained on the CelebA-HQ-256 dataset in terms of sampling time and number of function evaluations (NFEs) of the ODE solver. Song et al. propose two main sampling techniques including predictor-corrector (PC) and probability flow ODE. PC sampling involves 4000 NFEs and takes 44.6 min. on a Titan V for a batch of 16 images. It yields 7.23 FID score (see Tab. 3). ODE-based sampling from SGM takes 3.91 min. with 335 NFEs, but it obtains a poor FID score of 128.13 with as ODE solver error toleranceWe use the VESDE checkpoint at https://github.com/yang-song/score_sde_pytorch. Song et al. report that ODE-based sampling yields worse FID scores for their models (see D.4 in ). The problem is more severe for VESDEs. Unfortunately, at submission time only a VESDE model was released..
In a stark contrast, ODE-based sampling from our LSGM takes 0.07 min. with average of 23 NFEs, yielding 7.22 FID score. LSGM is 637 and 56 faster than original SGM’s PC and ODE sampling, respectively. In Fig. 4, we visualize FID scores and NFEs for different ODE solver error tolerances. Our LSGM achieves low FID scores for relatively large error tolerances.
We identify three main reasons for this significantly faster sampling from LSGM: (i) The SGM prior in our LSGM models latent variables with 3232 spatial dim., whereas the original SGM directly models 256256 images. The larger spatial dimensions require a deeper network to achieve a large receptive field. (ii) Inspecting the SGM prior in our model suggests that the score function is heavily dominated by the linear term at the end of training, as the mixing coefficients are all . This makes our SGM prior smooth and numerically faster to solve. (iii) Since SGM is formed in the latent space in our model, errors from solving the ODE can be corrected to some degree using the VAE decoder, while in the original SGM errors directly translate to artifacts in pixel space.
2 Ablation Studies
SDEs, objective weighting mechanisms and variance reduction. In Tab. 6, we analyze the different weighting mechanisms and variance reduction techniques and compare the geometric VPSDE with the regular VPSDE with linear . In the table, SGM-obj.-weighting denotes the weighting mechanism used when training the SGM prior (via Eq. 9). -sampling (SGM-obj.) indicates the sampling approach for , where , and denote the IS distributions for the weighted (likelihood), the unweighted, and the reweighted objective, respectively. For training the VAE encoder (last term in Eq. 8), we either sample a separate batch with importance sampling following (only necessary when the SGM prior is not trained with itself), or we reweight the samples drawn for training the prior according to the likelihood objective (denoted by rew.). n/a indicates fields that do not apply: The geometric VPSDE has optimal variance for the weighted (likelihood) objective already with uniform sampling; there is no additional IS distribution. Also, we did not derive IS distributions for the geometric VPSDE for . NaN indicates experiments that failed due to training instabilities. Previous work have reported instability in training large VAEs. We find that our method inherits similar instabilities from VAEs; however, importance sampling often stabilizes training our LSGM. As expected, we obtain the best NELBOs (red) when training with the weighted, maximum likelihood objective (). Importantly, our new geometric VPSDE achieves the best NELBO. Furthermore, the best FIDs (blue) are obtained either by unweighted () or reweighted () SGM prior training, with only slightly worse NELBOs. These experiments were run on the CIFAR10 dataset, using a smaller model than for our main results above (details in App. G).
End-to-end training. We proposed to train LSGM end-to-end, in contrast to . Using a similar setup as above we compare end-to-end training of LSGM during the second stage with freezing the VAE encoder and decoder and only training the SGM prior in latent space during the second stage. When training the model end-to-end, we achieve an FID of and NELBO of ; when freezing the VAE networks during the second stage, we only get an FID of and NELBO of . These results clearly motivate our end-to-end training strategy.
Mixing Normal and neural score functions. We generally found training LSGM without our proposed “mixed score” formulation (Sec. 3.2) to be unstable during end-to-end training, highlighting its importance. To quantify the contribution of the mixed score parametrization for a stable model, we train a small LSGM with only one latent variable group. In this case, without the mixed score, we reached an FID of and NELBO of ; with it, we got an FID of and NELBO of . Without the inductive bias provided by the mixed score, learning that the marginal distribution is close to a Normal one for large purely from samples can be very hard in the high-dimensional latent space, where our diffusion is run. Furthermore, due to our importance sampling schemes, we tend to oversample small, rather than large . However, synthesizing high-quality images requires an accurate score function estimate for all . On the other hand, the log-likelihood of samples is highly sensitive to local image statistics and primarily determined at small . It is plausible that we are still able to learn a reasonable estimate of the score function for these small even without the mixed score formulation. That may explain why log-likelihood suffers much less than sample quality, as estimated by FID, when we remove the mixed score parameterization.
Additional experiments and model samples are presented in App. H.
Conclusions
We proposed the Latent Score-based Generative Model, a novel framework for end-to-end training of score-based generative models in the latent space of a variational autoencoder. Moving from data to latent space allows us to form more expressive generative models, model non-continuous data, and reduce sampling time using smoother SGMs. To enable training latent SGMs, we made three core contributions: (i) we derived a simple expression for the cross entropy term in the variational objective, (ii) we parameterized the SGM prior by mixing Normal and neural score functions, and (iii) we proposed several techniques for variance reduction in the estimation of the training objective. Experimental results show that latent SGMs outperform recent pixel-space SGMs in terms of both data likelihood and sample quality, and they can also be applied to binary datasets. In large image generation, LSGM generates data several orders of magnitude faster than recent SGMs. Nevertheless, LSGM’s synthesis speed does not yet permit sampling at interactive rates, and our implementation of LSGM is currently limited to image generation. Therefore, future work includes further accelerating sampling, applying LSGMs to other data types, and designing efficient networks for LSGMs.
Broader Impact
Generating high-quality samples while fully covering the data distribution has been a long-standing challenge in generative learning. A solution to this problem will likely help reduce biases in generative models and lead to improving overall representation of minorities in the data distribution. SGMs are perhaps one of the first deep models that excel at both sample quality and distribution coverage. However, the high computational cost of sampling limits their widespread use. Our proposed LSGM reduces the sampling complexity of SGMs by a large margin and improves their expressivity further. Thus, in the long term, it can enable the usage of SGMs in practical applications.
Here, LSGM is examined on the image generation task which has potential benefits and risks discussed in . However, LSGM can be considered a generic framework that extends SGMs to non-continuous data types. In principle LSGM could be used to model, for example, language , music , or molecules . Furthermore, like other deep generative models, it can potentially be used also for non-generative tasks such as semi-supervised and representation learning . This makes the long-term social impacts of LSGM dependent on the downstream applications.
Funding Statement
All authors were funded by NVIDIA through full-time employment.
References
Appendix A Proof for Theorem 1
Without loss of generality, we state the theorem in general form without conditioning on .
with and a Normal transition kernel where and are obtained from and for a fixed initial variance at .
Theorem 1 amounts to estimating the cross entropy between and with denoising score matching and can be understood intuitively in the context of LSGM: We are drawing samples from a potentially complex encoding distribution , add Gaussian noise with small initial variance to obtain a well-defined initial distribution, and then smoothly perturb the sampled encodings using a diffusion process, while learning a denoising model, the SGM prior. Note that from the perspective of the learnt SGM prior, which is defined by the separate reverse-time generative SDE with the learnt score function model (see Sec. 2), the expression in our theorem becomes an upper bound (see discussion in Sec. 3.1).
The first part of our proof follows a similar proof strategy as was used by Song et al. . We start the proof with a more generic diffusion process in the form:
and analogously for .
since , as assumed in the Theorem (in practice, the used SDEs are designed such that ).
where inserts the Fokker Planck equations for and , respectively. Furthermore, is integration by parts assuming similar limiting behavior of and at as Song et al. . Specifically, we know that and must decay towards zero at to be normalized. Furthermore, we assumed and to have at most polynomial growth (or decay, when looking at it from the other direction) at , which implies faster exponential growth/decay of and . Also, and grow/decay at most polynomially, too, since the gradient of a polynomial is still a polynomial. Hence, one can work out that all terms to be evaluated at after integration by parts vanish. Finally, uses the log derivative trick and some rearrangements, and is obtained by inserting and .
which we can interpret as a general score matching-based expression for calculating the cross entropy, analogous to the expressions for the Kullback-Leibler divergence and entropy derived by Song et al. .
However, as discussed in the main paper, dealing with the marginal score is problematic for complex “input” distributions . Hence, we further transform the cross entropy expression into a denoising score matching-based expression:
with and where in we have used the following identity from Vincent :
In , we have added and subtracted and in we rearrange the terms into denoising score matching. In the following, we show that the term marked by (I) depends only on the diffusion parameters and does not depend on when takes a special affine (linear) form , which is often used for training SGMs and which we assume in our Theorem.
Note that for linear , we can derive the mean and variance (there are no “off-diagonal” co-variance terms here, since all dimensions undergo diffusion independently) of the distribution at any time in closed form, essentially solving the Fokker-Planck equation for this special case analytically. In that case, if the initial distribution at is Normal then the distribution stays Normal and the mean and variance completely describe the distribution, i.e. . The mean and variance are given by the differential equations and their solutions :
Here, denotes the mean of the distribution at and the component-wise variance at . After transforming into the denoising score matching expression above, what we are doing is essentially drawing samples from the potentially complex , then placing simple Normal distributions with variance at those samples, and then letting those distributions evolve according to the SDE. acts as a hyperparameter of the model.
In this case, i.e. when the distribution is Normal at all , we can represent samples from the intermediate distributions in reparameterized from where . We also know that With this we can write down (i) as:
Furthermore, since at , its entropy is . With this, we get the following simple expression for the cross-entropy:
Expressing the integral as an expectation completes the proof:
The expression in Theorem 1 measures the cross entropy between and at . However, one should consider practical implications of the choice of initial variance when estimating the cross entropy between two distributions using our expression, as we discuss below.
Consider two arbitrary distributions and . If the forward diffusion process has a non-zero initial variance (i.e., ), the actual distributions and at in the score matching expression are defined by and , which correspond to convolving and each with a Normal distribution with variance . In this case, and are not identical to and , respectively, in general. However, we can approximate and using and , respectively, when is small. That is why our expression in Theorem 1 that measures , can be considered as an approximation of when takes a positive small value. Note that in practice, our is indeed generally very small (see Tab. 7).
On the other hand, when (e.g., when using the VPSDE from Song et al. ), we know that and are identical to and . However, in this case, the initial distribution at is essentially an infinitely sharp Normal and we cannot evaluate the integral over the full interval . Hence, we limit its range to , where is another hyperparameter. In this case, we can approximate the cross entropy using:
Appendix B Variance Reduction
In order to study the variance of the training objective, we derive analytically, assuming that both . This is a reasonable simplification for our analysis because pretraining our LSGM model with a prior brings close to and our SGM prior is often dominated by the fixed Normal mixture component. Nevertheless, we empirically observe that the variance reduction techniques developed with this simplification still work well when and are not exactly .
In this section, we start with presenting the mixed score parameterization for generic SDEs in App. B.1. Then, we discuss variance reduction with importance sampling for these generic SDEs in App. B.2. Finally, in App. B.3 and App. B.4, we focus on variance reduction of the VPSDEs and VESDEs, respectively, and we briefly discuss the Sub-VPSDE in App. B.5.
The mixed score parameterization uses the score that is obtained when dealing with Normal input data and just predicts an additional residual score. In the main text, we assume that the variance of the standard Normal data stays the same throughout the diffusion process, which is the case for VPSDEs. But the way Normal data diffuses depends generally on the underlying SDE and generic SDEs behave differently than the regular VPSDE in that regard.
Consider the generic forward SDEs in the form:
If our data distribution is standard Normal, i.e. , using Eq. 13, we have
In the case of VPSDEs, we have which corresponds to the mixed score introduced in the main text.
B.2 Variance Reduction of Cross Entropy with Importance Sampling for Generic SDEs
Let’s consider the cross entropy expression for and where we have scaled down the variance of to to accommodate the fact that the diffusion process with initial variance applies a perturbation with variance in its initial step (hence, the marginal distribution at is and we know that the optimal score is , i.e., the Normal component).
The cross entropy with the optimal score is:
Therefore, the IW distribution with minimum variance for is
Finally, the cross entropy objective with importance weighting becomes
B.3 VPSDE
Consider the simple forward diffision process in the form:
which corresponds to the VPSDE from Song et al. . The appealing characteristic of this diffusion model is that if , intermediate will also have a standard Normal distribution and its variance is constant (i.e., ). In the original VPSDE, is defined by a linear function that interpolates between .
Our analysis in App. B.2, Eq. 30 shows that the cross entropy can be expressed as:
where for the VPSDE we have used .
A sample-based estimation of this expectation has a low variance if is constant for all . By solving the ODE , we can see that a log-linear noise schedule of the form satisfies this condition, with , , and .
Using Eq. 13, we can find an expression for that generates such noise schedule:
We call a VPSDE with defined as above a geometric VPSDE. For small and close to , all inputs diffuse closely towards the standard Normal prior at . In that regard, notice that our geometric VPSDE is well-behaved with positive only within the relevant interval and for . These conditions also imply for all . This is expected for any VPSDE. We can approach unit variance arbitrarily closely but not reach it exactly.
Importantly, our geometric VPSDE is different from the “variance-exploding” SDE (VESDE), proposed by Song et al. (also see App. C). The VESDE leverages a SDE in which the variance grows in an almost unbounded way, while the mean of the input distribution stays constant. Because of this, the hyperparameters of the VESDE must be chosen carefully in a data-dependent manner , which can be problematic in our case (see discussion in App. B.4). Furthermore, Song et al. also found that the VESDE does not perform well when used with probability flow-based sampling . In contrast, our geometric VPSDE combines the variance preserving behavior (i.e. standard Normal input data remains standard Normal throughout the diffusion process; all individual inputs diffuse towards standard Normal prior) of the VPSDE with the geometric growth of the variance in the diffusion process, which was first used in the VESDE.
Finally, for the geometric VPSDE we also have that for Normal input data. Hence, data is encoded “as continuously as possible” throughout the diffusion process. This is in line with the arguments made by Song et al. in . We hypothesize that this is particularly beneficial towards learning models with strong likelihood or NELBO performance. Indeed, in our experiments we observe the geometric VPSDE to perform best on this metric.
B.3.2 Variance Reduction for Likelihood Weighting (Importance Sampling)
Above, we have assumed that we sample from a uniform distribution for and we have defined and such that the variance of a Monte-Carlo estimation of the expectation is minimum. Another approach for improving the sample-based estimate of the expectation is to keep and unchanged and to use importance sampling such that the variance of the estimate is minimum.
Using importance sampling, we can rewrite the expectation in Eq. 42 as:
where is a proposal distribution. The theory of importance sampling shows that will have the smallest variance. In order to use this proposal distribution, we require (i) sampling from and (ii) evaluating the objective using this importance sampling technique.
Sampling from by inverse transform sampling: It’s easy to see that the normalization constant for is . Thus, the PDF is:
We can derive inverse transform sampling by deriving the inverse CDF:
where is the inverse of .
Importance Weighted Objective: The cross entropy is then written as (ignoring the constants here):
B.3.3 Variance Reduction for Unweighted Objective
Using a similar derivation as in App. B.2, we can show that for the unweighted objective for and , we have
with proposal distribution . Recall that in the VPSDE with linear , we have
Hence, the normalization constant of is
Similarly, we can write the CDF of as
solving for then results in
B.3.4 Variance Reduction for Reweighted Objective
For the reweighted mechanism, we drop only from the cross entropy objective but we keep . Using a similar derivation in App. B.2, we can show that unweighted objective for and , we have
with proposal distribution .
In this case, we have the following proposal , its CDF and inverse CDF :
Note that usually and . In that case, the inverse CDF can be thought of as .
Remark: It is worth noting that the derivation of the importance sampling distribution for the reweighted objective does not make any assumption on the form of . Thus, the IS distribution can be formed for any VPSDE when training with the reweighted objective, including the original VPSDE with linear and also our new geometric VPSDE.
B.4 VESDE
with .
Solving the Fokker-Planck equation for input distribution results in
Typical values for and are and (CIFAR10). Usually, we use .
Note that when the input data is distributed as , the variance at time in VESDE is given by:
Note that is typically very large and chosen empirically based on the scale of the data . However, this is tricky in our case, as the role of the data is played by the latent space encodings, which themselves are changing during training. We did briefly experiment with the VESDE and calculated as suggested in using the encodings after the VAE pre-training stage. However, these experiments were not successful and we suffered from significant training instabilities, even with variance reduction techniques. Therefore, we did not further explore this direction.
Nevertheless, our proposed variance reduction techniques via importance sampling can be derived also for the VESDE. Hence, for completeness, they are shown below.
Let’s have a closer look at the likelihood objective when using the VESDE for modeling the standard Normal data. Following similar arguments as in previous sections, we have . With the optimal score (i.e., the Normal component), we have the following expression for from Eq. 30:
Since the term inside the expectation is not constant in , the VESDE does not result in an objective with naturally minimal variance, opposed to our proposed geometric VPSDE.
We derive an importance sampling scheme with a proposal distribution
where is the inverse of .
So, the objective with importance sampling is then:
In contrast to the VESDE, the geometric VPSDE combines the geometric progression in diffusion variance directly with minimal variance in the objective by design. Furthermore, it is simpler to set up, because we can always choose for the geometric VPSDE and do not have to use a data-specific as proposed by .
B.4.2 Variance Reduction for Unweighted Objective
When we drop all “prefactors” in the objective, the importance sampling distribution stays the same as above, since is constant in . The objective becomes:
B.4.3 Variance Reduction for Reweighted Objective
To define the importance sampling for the reweighted objective by , we use the fact that in VESDEs. Using a similar derivation as in App. B.2, we show:
Thus, the optimal proposal for reweighted objective and the inverse CDF are:
So, the reweighted objective with importance sampling is:
Note that in practice, we can safely set as initial is non-zero in the VESDE.
B.5 Sub-VPSDE
Song et al. also proposed the Sub-VPSDE . It is defined as:
with the same linear as for the regular VPSDE.
Solving the Fokker-Planck equation for input distribution at results in
Deriving importance sampling distributions for variance reduction for the Sub-VPSDE can be more complicated than for the VPSDE, Geometric VPSDE, and VESDE and we did not investigate this in detail. However, for the same linear the Sub-VPSDE is close to the VPSDE, only slightly reducing the the variance of the diffusion process distribution for small . This suggests that the IS distribution derived using the regular VPSDE will likely also significantly reduce the variance of the objective due to -sampling of the Sub-VPSDE, just not as optimally as theoretically possible. In Fig. 6, we show the training NELBO of an LSGM trained on CIFAR-10 with -weighting using the Sub-VPSDE. We show the NELBO both for uniform sampling as well as for sampling from the IS distribution that was originally derived for the regular VPSDE with the same (the experiment and model setup is otherwise the same as the one for the ablation study on SDEs, weighting mechanisms and variance reduction). We indeed observe a significantly reduced training objective variance. We were consequently able to train large LSGM models in a stable manner using the Sub-VPSDE with VPSDE-based IS. However, the strongest generative performance in either NLL or FID was not achieved using the Sub-VPSDE, but with the Geometric VPSDE or regular VPSDE. For that reason, we did not focus on the Sub-VPSDE in our main experiments. However, a generative modeling performance comparison of the VPSDE vs. Sub-VPSDE in a smaller LSGM model is presented in App. H.4.
Appendix C Expressions for the Normal Transition Kernel
In our derivations of the Normal transition kernel , we only considered the general case in Eq. 12 and Eq. 13. However, the expression for can be further simplified for different SDEs that are considered in this paper. For completeness, we provide the expressions for below:
In both VESDE and Geometric VPSDE, the initial variance is denoted by . These diffusion processes start from a slightly perturbed version of the data at . In VESDE, by definition is large (as the name variance exploding SDE suggests) and it is set based on the scale of the data . In contrast, in the Geometric VPSDE does not depend on the scale of the data and it is set to . In the VPSDE, the initial variance is denoted by the hyperparameter . In contrast to VESDE and Geometric VPSDE, we often set the initial variance to zero in VPSDE, meaning that the diffusion process models the data distribution exactly at . However, using the VPSDE with comes at the cost of not being able to sample in the full interval $$ during training and also prevents us from solving the probability flow ODE all the way to zero during sampling .
Appendix D Probability Flow ODE
In LSGM, to sample from our SGM prior in latent space and to estimate NELBOs, we follow Song et al. and build on the connection between SDEs and ODEs. We use black-box ODE solvers to solve the probability flow ODE. Here, we briefly recap this approach.
All SDEs used in this paper can be written in the general form
The reverse of this diffusion process is also a diffusion process running backwards in time , defined by
where denotes a standard Wiener process going backwards in time, now represents a negative infinitesimal time increment, and is the score function of the diffusion process distribution at time . Interestingly, Song et al. have shown that there is a corresponding ODE that generates the same marginal probability distributions when acting upon the same prior distribution . It is given by
and usually called the probability flow ODE. This connects score-based generative models using diffusion processes to continuous Normalizing flows, which are based on ODEs . Note that in practice is approximated by a learnt model. Therefore, the generative distributions defined by the ODE and SDE above are formally not exactly equivalent when inserting this learnt model for the score function expression. Nevertheless, they often achieve quite similar performance in practice . This aspect is discussed in detail in concurrent work by Song et al. .
We can use the above ODE for efficient sampling of the model via black-box ODE solvers. Specifically, we can draw samples from the standard Normal prior distribution at and then solve this ODE towards . In fact, this is how we perform sampling from the latent SGM prior in our paper. Similarly, we can also use this ODE to calculate the probability of samples under this generative process using the instantaneous change of variables formula (see for details). We rely on this for calculating the probability of latent space samples under the score-based prior in LSGM. Note that this involves calculating the trace of the Jacobian of the ODE function. This is usually approximated via Hutchinson’s trace estimator, which is unbiased but has a certain variance (also see discussion in Sec. F).
This approach is applicable similarly for all diffusion processes and SDEs considered in this paper.
Appendix E Converting VAE with Hierarchical Normal Prior to Standard Normal Prior
Converting a VAE with hierarchical prior to a standard Normal prior can be done using a simple change of variables. Consider a VAE with hierarchical encoder and hierarchical prior where represent all latent variables and:
where for simplicity we have assumed that the variance is shared for all the components. We can reparameterize the latent variables by introducing . With this reparameterization, the equivalent VAE is:
where . In this equivalent parameterization, we can consider as latent variables with a standard Normal prior.
In NVAE , the prior has the same hierarchical form as in Eq. 78. However, the authors observe that the residual parameterization of the encoder often improves the generative performance. In this parameterization, with a small modification, the encoder is defined by:
where the encoder is tasked to predict the residual parameters and . Using the same reparameterization as above (), we have the equivalent VAE in the form:
where . In other words, the residual parameterization of encoder, introduced in NVAE, predicts the mean and variance for the distributions directly.
Appendix F Bias in Importance Weighted Estimation of Log-Likelihood
A common approach for estimating test log-likelihood in VAEs is to use the importance weighted bound on log-likelihood . In LSGM, we have access to an unbiased but stochastic estimation of the prior likelihood which we obtain using the probability flow ODE . The stochasticity in the estimation comes from Hutchinson’s trick . In VAEs, the test log-likelihood is estimated using importance weighted (IW) estimation :
which is a statistical lower bound on .
In this section, we provide an informal analysis that shows that IW estimation with can overestimate the log-likelihood when is measured with an unbiased estimator with variance . In our analysis we assume that is small and we use Taylor expansion to study how the IW bound varies. Under our analysis, we observe that the bias has and it can be minimized by ensuring that is sufficiently small.
where H is the Hessian matrix for the function at . Note that the gradient is the softmax function and . Thus, we have:
So, when the importance weights are estimated with sufficiently small variance , the bias is proportional to the variance of this estimate.
In our experiments, we observe that the variance of the estimate is not small enough to obtain a reliable estimate of test likelihood using the importance weighted bound. One way to reduce the variance is to use many randomly sampled noise vectors in Hutchinson’s trick. However, this makes NLL estimation computationally too expensive. Fortunately, when evaluating NELBO (which corresponds to here), the NELBO estimate is unbiased and its variance is small because of averaging across big test datasets (with often 10k samples). For example, on MNIST the standard deviation of our estimate is 0.36 nat, while the standard deviation of NELBO is 0.07 nat.
Appendix G Additional Implementation Details
All hyperparameters for our main models are provided in Tab. 7.
The VAE backbone for all LSGM models is NVAE https://github.com/NVlabs/NVAE (NVIDIA Source Code License), one of the best-performing VAEs in the literature. It has a hierarchical latent space with group-wise autoregressive latent variable dependencies and it leverages residual neural networks (for architecture details see ). It uses depth-wise separable convolutions in the decoder. Although both the approximate posterior and the prior are hierarchical in its original version, we can reparametrize the prior and write it as a product of independent Normal distributions (see Sec. E).
The VAE’s most important hyperparameters include the number of latent variable groups and their spatial resolution, the channel depth of the latent variables, the number of residual cells per group, and the number of channels in the convolutions in the residual cells. Furthermore, when training the VAE during the first stage we are using KL annealing and KL balancing, as described in . For some models, we complete KL annealing during the pre-training stage, while for other models we found it beneficial to anneal only up to a KL-weight in the ELBO during the first stage and complete KL annealing during the main end-to-end LSGM training stage. This provides additional flexibility in learning an expressive distribution in latent space during the second training stage, as it prevents more latent variables from becoming inactive while the prior is being trained gradually. However, when using a very large backbone VAE together with an SGM objective that does not correspond to maximum likelihood training, i.e. - or -weighting, we empirically observe that this approach can also hurt NLL, while slightly improving FID (see CIFAR10 (best FID) model).
Note that the VAE Backbone performance for CIFAR10 reported in Tab. 2 in the main paper corresponds to the 20-group backbone VAE (trained to full KL-weight ) from the CIFAR10 (balanced) LSGM model (see hyperparameter Tab. 7).
Image Decoders: Since SGMs assume that the data is continuous, they rely on uniform dequantization when measuring data likelihood. However, in LSGM, we rely on decoders designed specifically for images with discrete intensity values. On color images, we use mixtures of discretized logistics , and on binary images, we use Bernoulli distributions. These decoder distributions are both available from the NVAE implementation.
G.2 Latent SGM Prior
Our denoising networks for the latent SGM prior are based on the NCSN++ architecture from Song et al. , adapted such that the model ingests and predicts tensors according to the VAE’s latent variable dimensions. We vary hyperparameters such as the number of residual cells per spatial resolution level and the number of channels in convolutions. Note that all our models use dropout in the SGM prior. Some of our models use upsampling and downsampling operations with anti-aliasing based on Finite Impulse Response (FIR) , following Song et al. .
NVAE has a hierarchical latent structure. For small image datasets including CIFAR-10, MNIST and OMNIGLOT all the latent variables have the same spatial dimensions. Thus, the diffusion process input is constructed by concatenating the latent variables from all groups in the channel dimension. Our NVAE backbone on the CelebA-HQ-256 dataset comes with multiple spatial resolutions in latent groups. In this case, we only feed the smallest resolution groups to the SGM prior and assume that the remaining groups have a standard Normal distribution.
G.3 Training Details
To optimize our models, we are mostly following the previous literature. The VAE’s encoder and decoder networks are trained using an Adamax optimizer , following NVAE . In the second stage, the whole model is trained with an Adam optimizer and we perform learning rate annealing for the VAE network optimization, while we keep the learning rate constant when optimizing the SGM prior parameters. At test time, we use an exponential moving average (EMA) of the parameters of the SGM prior with EMA decay rate, following . Note that, when using the VPSDE with linear , we are also generally following and use and . We did not observe any benefits in using the EMA parameters for the VAE networks.
G.4 Evaluation Details
For evaluation, we are drawing samples and calculating log-likelihoods using the probability flow ODE, leveraging black-box ODE solvers, following . Similar to , we are using an RK45 ODE solver , based on scipy, using the torchdiffeq interface https://github.com/rtqichen/torchdiffeq (MIT License). Integration cutoffs close to zero and ODE solver error tolerances used for evaluation are indicated in Tab. 7 (for example, for the VPSDE with linear we usually use and therefore have that goes to at , hence preventing us from integrating the probability flow ODE all the way to exactly . This was handled similarly by Song et al. ).
Following the conventions established by previous work , when evaluating our main models we compute FID at frequent intervals during training and report FID and NLL at the minimum observed FID.
Vahdat and Kautz in NVAE observe that setting the batch normalization (BN) layers to train mode during sampling (i.e., using batch statistics for normalization instead of moving average statistics) improves sample quality. We similarly observe that setting BN layers to train mode improves sample quality by about 1 FID score on the CelebA-HQ-256 dataset, but it does not affect performance on the CIFAR-10 dataset. In contrast to NVAE, we do not change the temperature of the prior during sampling, as we observe that it hurts generation quality.
G.5 Ablation Experiments
Here we provide additional details and discussions about the ablation experiments performed in the paper.
The models that were used for the ablation experiment on SDEs, objective weighting mechanisms and variance reduction and produced the results in Tab. 6 in the main paper use an overall similar setup as the CIFAR10 (best NLL) one, with a few exceptions: They are trained only for 1000 epochs and evaluation always happens using the checkpoint at the end of training. Furthermore, the total batchsize over all GPUs is reduced from 256 to 128. Additionally, only 2 instead of 8 cells per residual are used in the latent SGM prior networks. Finally, the VAE’s KL term is annealed all the way to during the first training stage for these experiments. All other hyperparameters correspond to the CIFAR10 (best NLL) setup, except those that are explicitly varied as part of the ablation study and mentioned in Tab. 6 in the paper.
As discussed in the main paper, the results of this ablation study overall validate that importance sampling is important to stabilize training, that the -weighting mechanism as well as our novel geometric VPSDE are well suited for training towards strong likelihood, and that the - and -weighting mechanisms tend to produce better FIDs. Although these trends generally hold, it is noteworthy that not all results translate perfectly to our large models that we used to produce our main results. For instance, the setting with -weighting and no importance sampling for the SGM objective, which produced the best FID in Tab. 6 (main paper), is generally unstable for our bigger models, in line with our observation that IS is usually necessary to stabilize training. The stable training run for this setting in Tab. 6 can be considered an outlier.
Furthermore, for CIFAR10 we obtained our very best FID results using the VPSDE, -weighting, IS, and sample reweighting for the -objective, while for the slightly smaller models used for the results in Tab. 6, there is no difference between using sample reweighting and drawing a separate batch with for training for this case (see Tab. 6 main paper, VPSDE, , fields). Also, CelebA-HQ-256 behaves slightly different for the large models in that the VPSDE with -weighting and sampling a separate batch with for -training performed best by a small margin (see hyperparameter Tab. 7).
G.5.2 Ablation: End-to-End Training
The model used for the results on the ablation study regarding end-to-end training vs. fully separate VAE and SGM prior training is the same one as used for the ablation study on SDEs, objective weighting mechanisms and variance reduction above, evaluated in a similar way. For this experiment, we used the VPSDE, -objective weighting, IS for with when training the SGM prior, and we did draw a second batch with for training (only relevant for the end-to-end training setup).
G.5.3 Ablation: Mixing Normal and Neural Score Functions
The model used for the ablation study on mixing Normal and neural score functions is again similar to the one used for the other ablations with the exception that the underlying VAE has only a single latent variable group, which makes it much smaller and removes all hierarchical dependencies between latent variables. We tried training multiple models with larger backbone VAEs, but they were generally unstable when trained without our mixed score parametrization, which only hightlights its importance. As for the previous ablation, for this experiment we used the VPSDE, -objective weighting, IS for with when training the SGM prior, and we did draw a second batch with for training .
G.6 Training Algorithms
To unambiguously clarify how we train our LSGMs, we summarized the training procedures in three different algorithms for different situations:
Likelihood training with IS. In this case, the SGM prior and the encoder share the same weighted likelihood objective and do not need to be updated separately.
Un/Reweighted training with separate IS of for SGM-objective and -objective. Here, the SGM prior and the encoder need to be updated with different weightings, because the encoder always needs to be trained using the weighted (maximum likelihood) objective. We draw separate batches using separate IS distribution for the two differently weighted objectives (i.e. last term in Eq. 8 from main paper vs. Eq. 9).
Un/Reweighted training with IS of for the SGM-objective and reweighting for the -objective. What this means is that when training the encoder with the score-based cross entropy term (last term in Eq. 8 from main paper), we are using an importance sampling distribution that was actually tailored to un- or reweighted training for the SGM objective (Eq. 9 from main paper) and therefore isn’t optimal for the weighted (maximum likelihood) objective necessary for encoder training. However, if we nevertheless use the same importance sampling distribution, we do not need to draw a second batch of for encoder training. In practice, this boils down to different (re-)weighting factors in the cross entropy term (see Algorithm 3).
For efficiency comparison between approaches (2) and (3), we observe that (3) consumes more memory than (2) in general but it can be faster due to the shared computation for the denoising step. Due to the memory limitations, we use (2) on large image datasets. Note that the choice between (2) and (3) may affect generative performance as we empirically observed in our experiments.
G.7 Computational Resources
In total, the research project consumed GPU hours, which translates to an electricity consumption of about MWh. We used an in-house GPU cluster of V100 NVIDIA GPUs.
Appendix H Additional Experiments
In this section, we provide additional samples generated by our models for CIFAR-10 in Fig. 7, and CelebA-256-HQ in Fig. 8.
H.2 MNIST: Small VAE Experiment
Here, we examine our LSGM on a small VAE architecture. We specifically follow and build a small VAE in the NVAE codebase. In particular, the model does not have hierarchical latent variables, but only a single latent variable group with a total of 64 latent variables. Encoder and decoder consist of small ResNets with 6 residual cells in total (every two cells there is a down- or up-sampling operation, so we have 3 blocks with 2 residual cells per block). The experiments are done on dynamically binarized MNIST. As we can see in Table 8, our implementation of the VAE obtains a similar test NELBO as . However, our LSGM improves the NELBO by almost 4.6 nats. This simple experiment shows that we can even obtain good generative performance with our LSGM using small VAE architectures.
H.3 CIFAR-10: Neural Network Evaluations during Sampling
In Tab. 9, we report the number of neural network evaluations performed by the ODE solver during sampling from our CIFAR-10 models. ODE solver error tolerance is and time integration cutoff is . CIFAR-10 is a highly diverse and more multimodal dataset, compared to CelebA-HQ-256. Because of that, the latent SGM prior that is learnt is more complex, requiring more function evaluations.
H.4 CIFAR-10: Sub-VPSDE vs. VPSDE
In App. B.5 we discussed how variance reduction techniques derived based on the VPSDE can also help reducing the variance of the sample-based estimate of the training objective when using the Sub-VPSDE in the latent space SGM. Here, we perform a quantitative comparison between the VPSDE and the Sub-VPSDE, following the same experimental setup and using the same models as for the ablation study on SDEs, objective weighting mechanisms, and variance reduction (experiment details in App. G.5.1). The results are reported in Tab. 10. We find that the VPSDE generally performs slightly better in FID, while we observed little difference in NELBO in these experiments. Importantly, the Sub-VPSDE also did not outperform our novel geometric VPSDE in NELBO. We also see that the combination of Sub-VPSDE with -weighting performs poorly. Consequently, we did not explore the Sub-VPSDE further in our main experiments.
H.5 CelebA-HQ-256: Different ODE Solver Error Tolerances
In Fig. 9, we visualize CelebA-HQ-256 samples from our LSGM model for varying ODE solver error tolerances.
H.6 CelebA-HQ-256: Ancestral Sampling
For our experiments in this paper, we use the probability flow ODE to sample from the model. However, on CelebA-HQ-256, we observe that ancestral sampling from the prior instead of solving the probability flow ODE often generates much higher quality samples. However, the FID score is slightly worse for this approach. In Fig. 10, Fig. 11, and Fig. 12, we visualize samples generated with different numbers of steps in ancestral sampling.
H.7 CelebA-HQ-256: Sampling from VAE Backbone vs. LSGM
For the quantitative results on the CelebA-HQ-256 dataset in the main text, we use an LSGM with spatial dimension of 3232 for the latent variables in the SGM prior. However, for the qualitative results we used an LSGM with the prior spatial dimension of 6464. The 3232 dimensional model achieves a better FID score compared to the 6464 dimensional model (FID 7.22 vs. 8.53) and sampling from it is much faster (2.7 sec. vs. 39.9 sec.). However, the visual quality of the samples is slightly worse. In this section, we visualize samples generated by the 3232 dimensional model as well as the VAE backbone for this model. In this experiment, the VAE backbone is fully trained. Samples from our VAE backbone are visualized in Fig. 13 and for our 3232 dimensional LSGM in Fig. 14.
H.8 Evolution Samples on the ODE and SDE Reverse Generative Process
In Fig. 15, we visualize the evolution of the latent variables under both the reverse generative SDE and also the probability flow ODE. We are decoding the intermediate latent samples along the reverse-time generative process via the decoder to pixel space.