Diagnosing and Enhancing VAE Models
Bin Dai, David Wipf
Introduction
Ideally we might like to minimize the negative log-likelihood -\log p_{\theta}(\mbox{\boldmathx}) averaged across the ground-truth measure , i.e., solve \min_{\theta}\int_{\mbox{\boldmath\chi}}-\log p_{\theta}(\mbox{\boldmathx})\mu_{gt}(d\mbox{\boldmathx}). Unfortunately though, the required marginalization over is generally infeasible. Instead the VAE model relies on tractable encoder q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx}) and decoder p_{\theta}(\mbox{\boldmathx}|\mbox{\boldmathz}) distributions, where represents additional trainable parameters. The canonical VAE cost is a bound on the average negative log-likelihood given by
where the inequality follows directly from the non-negativity of the KL-divergence. Here can be viewed as tuning the tightness of bound, while dictates the actual estimation of . Using a few standard manipulations, this bound can also be expressed as
which explicitly involves the encoder/decoder distributions and is conveniently amenable to SGD optimization of via a reparameterization trick (Kingma and Welling, 2014; Rezende et al., 2014). The first term in (2) can be viewed as a reconstruction cost (or a stochastic analog of a traditional autoencoder), while the second penalizes posterior deviations from the prior p(\mbox{\boldmathz}). Additionally, for any realizable implementation via SGD, the integration over must be approximated via a finite sum across training samples \{\mbox{\boldmathx}^{(i)}\}_{i=1}^{n} drawn from . Nonetheless, examining the true objective can lead to important, practically-relevant insights.
At least in principle, q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx}) and p_{\theta}(\mbox{\boldmathx}|\mbox{\boldmathz}) can be arbitrary distributions, in which case we could simply enforce q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx})=p_{\theta}(\mbox{\boldmathz}|\mbox{\boldmathx})\propto p_{\theta}(\mbox{\boldmathx}|\mbox{\boldmathz})p(\mbox{\boldmathz}) such that the bound from (1) is tight. Unfortunately though, this is essentially always an intractable undertaking. Consequently, largely to facilitate practical implementation, a commonly adopted distributional assumption for continuous data is that both q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx}) and p_{\theta}(\mbox{\boldmathx}|\mbox{\boldmathz}) are Gaussian. This design choice has previously been cited as a key limitation of VAEs (Burda et al., 2015; Kingma et al., 2016), and existing quantitative tests of generative modeling quality thus far dramatically favor contemporary alternatives such as generative adversarial networks (GAN) (Goodfellow et al., 2014). Regardless, because the VAE possesses certain desirable properties relative to GAN models (e.g., stable training (Tolstikhin et al., 2018), interpretable encoder/inference network (Brock et al., 2016), outlier-robustness (Dai et al., 2018), etc.), it remains a highly influential paradigm worthy of examination and enhancement.
In Section 2 we closely investigate the implications of VAE Gaussian assumptions leading to a number of interesting diagnostic conclusions. In particular, we differentiate the situation where , in which case we prove that recovering the ground-truth distribution is actually possible iff the VAE global optimum is reached, and , in which case the VAE global optimum can be reached by solutions that reflect the ground-truth distribution almost everywhere, but not necessarily uniquely so. In other words, there could exist alternative solutions that both reach the global optimum and yet do not assign the same probability measure as .
Section 3 then further probes this non-uniqueness issue by inspecting necessary conditions of global optima when . This analysis reveals that an optimal VAE parameterization will provide an encoder/decoder pair capable of perfectly reconstructing all \mbox{\boldmathx}\in\mbox{\boldmath\chi} using any drawn from q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx}). Moreover, we demonstrate that the VAE accomplishes this using a degenerate latent code whereby only dimensions are effectively active. Collectively, these results indicate that the VAE global optimum can in fact uniquely learn a mapping to the correct ground-truth manifold when , but not necessarily the correct probability measure within this manifold, a critical distinction.
Next we leverage these analytical results in Section 4 to motivate an almost trivially-simple, two-stage VAE enhancement for addressing typical regimes when . In brief, the first stage just learns the manifold per the allowances from Section 3, and in doing so, provides a mapping to a lower dimensional intermediate representation with no degenerate dimensions that mirrors the regime. The second (much smaller) stage then only needs to learn the correct probability measure on this intermediate representation, which is possible per the analysis from Section 2. Experiments from Section 5 reveal that this procedure can generate high-quality crisp samples, avoiding the blurriness often attributed to VAE models in the past (Dosovitskiy and Brox, 2016; Larsen et al., 2015). And to the best of our knowledge, this is the first demonstration of a VAE pipeline that can produce stable FID scores, an influential recent metric for evaluating generated sample quality (Heusel et al., 2017), that are comparable to at least some popular GAN models under neutral testing conditions. Moreover, this is accomplished without additional penalty functions, cost function modifications, or sensitive tuning parameters. Finally, Section 6 provides concluding thoughts and a discussion of broader VAE modeling paradigms, such as those involving normalizing flows, parameterized families for p(\mbox{\boldmathz}), or modifications to encourage disentangled representations.
High-Level Impact of VAE Gaussian Assumptions
Conventional wisdom suggests that VAE Gaussian assumptions will introduce a gap between and the ideal negative log-likelihood \int_{\mbox{\boldmath\chi}}-\log p_{\theta}(\mbox{\boldmathx})\mu_{gt}(d\mbox{\boldmathx}), compromising efforts to learn the ground-truth measure. However, we will now argue that this pessimism is in some sense premature. In fact, we will demonstrate that, even with the stated Gaussian distributions, there exist parameters and that can simultaneously: (i) Globally optimize the VAE objective and, (ii) Recover the ground-truth probability measure in a certain sense described below. This is possible because, at least for some coordinated values of and , q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx}) and p_{\theta}(\mbox{\boldmathz}|\mbox{\boldmathx}) can indeed become arbitrarily close. Before presenting the details, we first formalize a -simple VAE, which is merely a VAE model with explicit Gaussian assumptions and parameterizations:
A -simple VAE is defined as a VAE model with \mbox{dim}[\mbox{\boldmathz}]=\kappa latent dimensions, the Gaussian encoder q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx})=\mathcal{N}(\mbox{\boldmathz}|\mbox{\boldmath\mu}_{z},\mbox{\boldmath\Sigma}_{z}), and the Gaussian decoder p_{\theta}(\mbox{\boldmathx}|\mbox{\boldmathz})=\mathcal{N}(\mbox{\boldmathx}|\mbox{\boldmath\mu}_{x},\mbox{\boldmath\Sigma}_{x}). Moreover, the encoder moments are defined as \mbox{\boldmath\mu}_{z}=f_{\mu_{z}}(\mbox{\boldmathx};\phi) and \mbox{\boldmath\Sigma}_{z}=\mbox{\boldmathS}_{z}\mbox{\boldmathS}_{z}^{\top} with \mbox{\boldmathS}_{z}=f_{S_{z}}(\mbox{\boldmathx};\phi). Likewise, the decoder moments are \mbox{\boldmath\mu}_{x}=f_{\mu_{x}}(\mbox{\boldmathz};\theta) and \mbox{\boldmath\Sigma}_{x}=\gamma\mbox{\boldmathI}. Here is a tunable scalar, while , and specify parameterized differentiable functional forms that can be arbitrarily complex, e.g., a deep neural network.
Equipped with these definitions, we will now demonstrate that a -simple VAE, with , can achieve the optimality criteria (i) and (ii) from above. In doing so, we first consider the simpler case where , followed by the extended scenario with . The distinction between these two cases turns out to be significant, with practical implications to be explored in Section 4.
This follows because by VAE design it must be that \mathcal{L}(\theta,\phi)\geq-\int p_{gt}(\mbox{\boldmathx})\log p_{gt}(\mbox{\boldmathx})d\mbox{\boldmathx}, and in the present context, this lower bound is achievable iff the conditions from (3) hold. Collectively, this implies that the approximate posterior produced by the encoder q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx}) is in fact perfectly matched to the actual posterior p_{\theta}(\mbox{\boldmathz}|\mbox{\boldmathx}), while the corresponding marginalized data distribution p_{\theta}(\mbox{\boldmathx}) is perfectly matched the ground-truth density p_{gt}(\mbox{\boldmathx}) as desired. Perhaps surprisingly, a -simple VAE can actually achieve such a solution:
All the proofs can be found in the appendices. So at least when , the VAE Gaussian assumptions need not actually prevent the optimal ground-truth probability measure from being recovered, as long as the latent dimension is sufficiently large (i.e., ). And contrary to popular notions, a richer class of distributions is not required to achieve this. Of course Theorem 2 only applies to a restricted case that excludes ; however, later we will demonstrate that a key consequence of this result can nonetheless be leveraged to dramatically enhance VAE performance.
2 Manifold Dimension Less Than Ambient Space Dimension (r<d𝑟𝑑r<d)
When , additional subtleties are introduced that will be unpacked both here and in the sequel. To begin, if both q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx}) and p_{\theta}(\mbox{\boldmathx}|\mbox{\boldmathz}) are arbitrary/unconstrained (i.e., not necessarily Gaussian), then . To achieve this global optimum, we need only choose such that q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx})=p_{\theta}(\mbox{\boldmathz}|\mbox{\boldmathx}) (minimizing the KL term from (1)) while selecting such that all probability mass collapses to the correct manifold . In this scenario the density p_{\theta}(\mbox{\boldmathx}) will become unbounded on and zero elsewhere, such that \int_{\mbox{\boldmath\chi}}-\log p_{\theta}(\mbox{\boldmathx})\mu_{gt}(d\mbox{\boldmathx}) will approach negative infinity.
Regardless, there is an absolutely crucial distinction between Theorem 3 and the simpler case quantified by Theorem 2. Although both describe conditions whereby the -simple VAE can achieve the minimal possible objective, in the case achieving the lower bound (whether the specific parameterization for doing so is unique or not) necessitates that the ground-truth probability measure has been recovered almost everywhere. But the situation is quite different because we have not ruled out the possibility that a different set of parameters could push to and yet not achieve (6). In other words, the VAE could reach the lower bound but fail to closely approximate . And we stress that this uniqueness issue is not a consequence of the VAE Gaussian assumptions per se; even if q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx}) were unconstrained the same lack of uniqueness can persist.
To conclude, the key take-home message from this section is that, at least in principle, VAE Gaussian assumptions need not actually be the root cause of any failure to recover ground-truth distributions. Instead we expose a structural deficiency that lies elsewhere, namely, the non-uniqueness of solutions that can optimize the VAE objective without necessarily learning a close approximation to . But to probe this issue further and motivate possible workarounds, it is critical to further disambiguate these optimal solutions and their relationship with ground-truth manifolds. This will be the task of Section 3, where we will explicitly differentiate the problem of locating the correct ground-truth manifold, from the task of learning the correct probability measure within the manifold.
Note that the only comparable prior work we are aware of related to the results in this section comes from Doersch (2016), where the implications of adopting Gaussian encoder/decoder pairs in the specialized case of are briefly considered. Moreover, the analysis there requires additional much stronger assumptions than ours, namely, that p_{gt}(\mbox{\boldmathx}) should be nonzero and infinitely differentiable everywhere in the requisite 1D ambient space. These requirements of course exclude essentially all practical usage regimes where or , or when ground-truth densities are not sufficiently smooth.
Optimal Solutions and the Ground Truth Manifold
We will now more closely examine the properties of optimal -simple VAE solutions, and in particular, the degree to which we might expect them to at least reflect the true , even if perhaps not the correct probability measure defined within . To do so, we must first consider some necessary conditions for VAE optima:
Let denote an optimal -simple VAE solution (with ) where the decoder variance is fixed (i.e., it is the sole unoptimized parameter). Moreover, we assume that is not a Gaussian distribution when .This requirement is only included to avoid a practically irrelevant form of non-uniqueness that exists with full, non-degenerate Gaussian distributions. Then for any , there exists a such that .
This result implies that we can always reduce the VAE cost by choosing a smaller value of , and hence, if is not constrained, it must be that if we wish to minimize (2). Despite this necessary optimality condition, in existing practical VAE applications, it is standard to fix during training. This is equivalent to simply adopting a non-adaptive squared-error loss for the decoder and, at least in part, likely contributes to unrealistic/blurry VAE-generated samples (Bousquet et al., 2017). Regardless, there are more significant consequences of this intrinsic favoritism for , in particular as related to reconstructing data drawn from the ground-truth manifold :
Applying the same conditions and definitions as in Theorem 4, then for all drawn from , we also have that
By design any random draw \mbox{\boldmathz}\sim q_{\phi_{\gamma}^{*}}\left(\mbox{\boldmathz}|\mbox{\boldmathx}\right) can be expressed as f_{\mu_{z}}(\mbox{\boldmathx};\phi_{\gamma}^{*})+f_{S_{z}}(\mbox{\boldmathx};\phi_{\gamma}^{*})\mbox{\boldmath\varepsilon} for some \mbox{\boldmath\varepsilon}\sim{\mathcal{N}}(\mbox{\boldmath\varepsilon}|{\bf 0},\mbox{\boldmathI}). From this vantage point then, (7) effectively indicates that any \mbox{\boldmathx}\in\mbox{\boldmath\chi} will be perfectly reconstructed by the VAE encoder/decoder pair at globally optimal solutions, achieving this necessary condition despite any possible stochastic corrupting factor f_{S_{z}}(\mbox{\boldmathx};\phi_{\gamma}^{*})\mbox{\boldmath\varepsilon}.
This observation is significant when we consider the inclusion of addition latent dimensions by allowing . Clearly based on the analysis above, adding dimensions to cannot improve the value of the VAE data term in any meaningful way. However, it can have a detrimental impact on the the KL regularization factor in the regime, where
But we must emphasize that the VAE can learn independently of the actual distribution within . Addressing the latter is a completely separate issue from achieving the perfect reconstruction error defined by Theorem 5. This fact can be understood within the context of a traditional PCA-like model, which is perfectly capable of learning a low-dimensional subspace containing some training data without actually learning the distribution of the data within this subspace. The central issue is that there exists an intrinsic bias associated with the VAE objective such that fitting the distribution within the manifold will be completely neglected whenever there exists the chance for even an infinitesimally better approximation of the manifold itself.
Stated differently, if VAE model parameters have learned a near optimal, parsimonious latent mapping onto using , then the VAE cost will scale as regardless of . Hence there remains a huge incentive to reduce the reconstruction error still further, allowing to push even closer to zero and the cost closer to . And if we constrain to be sufficiently large so as to prevent this from happening, then we risk degrading/blurring the reconstructions and widening the gap between q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx}) and p_{\theta}(\mbox{\boldmathz}|\mbox{\boldmathx}), which can also compromise estimation of . Fortunately though, as will be discussed next there is a convenient way around this dilemma by exploiting the fact that this dominanting factor goes away when .
From Theory to Practical VAE Enhancements
Sections 2 and 3 have exposed a collection of VAE properties with useful diagnostic value in and of themselves. But the practical utility of these results, beyond the underappreciated benefit of learning , warrant further exploration. In this regard, suppose we wish to develop a generative model of high-dimensional data \mbox{\boldmathx}\in\mbox{\boldmath\chi} where unknown low-dimensional structure is significant (i.e., the case with unknown). The results from Section 3 indicate that the VAE can partially handle this situation by learning a parsimonious representation of low-dimensional manifolds, but not necessarily the correct probability measure within such a manifold. In quantitative terms, this means that a decoder p_{\theta}(\mbox{\boldmathx}|\mbox{\boldmathz}) will map all samples from an encoder q_{\phi}(\mbox{\boldmathz}|\mbox{\boldmathx}) to the correct manifold such that the reconstruction error is negligible for any \mbox{\boldmathx}\in\mbox{\boldmath\chi}. But if the measure on has not been accurately estimated, then
where q_{\phi}(\mbox{\boldmathz}) is sometimes referred to as the aggregated posterior (Makhzani et al., 2016). In other words, the distribution of the latent samples drawn from the encoder distribution, when averaged across the training data, will have lingering latent structure that is errantly incongruous with the original isotropic Gaussian prior. This then disrupts the pivotal ancestral sampling capability of the VAE, implying that samples drawn from \mathcal{N}(\mbox{\boldmathz}|0,I) and then passed through the decoder p_{\theta}(\mbox{\boldmathx}|\mbox{\boldmathz}) will not closely approximate . Fortunately, our analysis suggests the following two-stage remedy:
Train a second -simple VAE, with independent parameters and latent representation , to learn the unknown distribution q_{\phi}(\mbox{\boldmathz}), i.e., treat q_{\phi}(\mbox{\boldmathz}) as a new ground-truth distribution and use samples \{\mbox{\boldmathz}^{(i)}\}_{i=1}^{n} to learn it.
Samples approximating the original ground-truth can then be formed via the extended ancestral process \mbox{\boldmathu}\sim\mathcal{N}(\mbox{\boldmathu}|{\bf 0},\mbox{\boldmathI}), \mbox{\boldmathz}\sim p_{\theta^{\prime}}(\mbox{\boldmathz}|\mbox{\boldmathu}), and finally \mbox{\boldmathx}\sim p_{\theta}(\mbox{\boldmathx}|\mbox{\boldmathz}).
Consequently, as long as we set , the operational regime of the second-stage VAE is effectively equivalent to the situation described in Section 2.1 where the manifold dimension is equal to the ambient dimension.Note that if a regular autoencoder were used to replace the first-stage VAE, then this would no longer be the case, so indeed a VAE is required for both stages unless we have strong prior knowledge such that we may confidently set . And as we have already shown there via Theorem 2, the VAE can readily handle this situation, since in the narrow context of the second-stage VAE, , the troublesome factor becomes zero, and any globally minimizing solution is uniquely matched to the new ground-truth distribution q_{\phi}(\mbox{\boldmathz}). Consequently, the revised aggregated posterior q_{\phi^{\prime}}(\mbox{\boldmathu}) produced by the second-stage VAE should now closely resemble {\mathcal{N}}(\mbox{\boldmathu}|{\bf 0},\mbox{\boldmathI}). And importantly, because we generally assume that , we have found that the second-stage VAE can be quite small.
It should also be emphasized that concatenating the two VAE stages and jointly training does not generally improve the performance. If trained jointly the few extra second-stage parameters can simply be hijacked by the dominant influence of the first stage reconstruction term and forced to work on an incrementally better fit of the manifold rather than addressing the critical mismatch between q_{\phi}(\mbox{\boldmathz}) and {\mathcal{N}}(\mbox{\boldmathu}|{\bf 0},\mbox{\boldmathI}). This observation can be empirically tested, which we have done in multiple ways. For example, we have tried fusing the respective encoders and decoders from the first and second stages to train what amounts to a slightly more complex single VAE model. We have also tried merging the two stages including the associated penalty terms. In both cases, joint training does not help at all as expected, with average performance no better than the first stage VAE (which contains the vast majority of parameters). Consequently, although perhaps counterintuitive, separate training of these two VAE stages is actually critical to achieving high quality results as will be demonstrated next.
Empirical Evaluation of VAE Two-Stage Enhancement
In this section we first present quantitative evaluations of the proposed two-stage VAE against various GAN and VAE baselines. We then describe experiments explicitly designed to corroborate our theoretical findings from Sections 2 and 3, including a demonstration of robustness to the choice of . Finally, we include representative samples generated from our model.
We first present quantitative evaluation of novel generated samples using the large-scale testing protocol of GAN models from (Lucic et al., 2018). In this regard, GANs are well-known to dramatically outperform existing VAE approaches in terms of the Fréchet Inception Distance (FID) score (Heusel et al., 2017) and related quantitative metrics. For fair comparison, (Lucic et al., 2018) adopted a common neutral architecture for all models, with generator and discriminator networks based on InfoGAN (Chen et al., 2016a); the point here is standardized comparisons, not tuning arbitrarily-large networks to achieve the lowest possible absolute FID values. We applied the same architecture to our first-stage VAE decoder and encoder networks respectively for direct comparison. For the low-dimensional second-stage VAE we used small, 3-layer networks contributing negligible additional parameters beyond the first stage (see the appendices for further design details).
We evaluated our proposed VAE pipeline, henceforth denoted as 2-Stage VAE, against three baseline VAE models differing only in the decoder output layer: a Gaussian layer with fixed , a Gaussian layer with a learned , and a cross-entropy layer as has been adopted in several previous applications involving images (Chen et al., 2016b). We also tested the Gaussian decoder VAE model (with learned ) combined with an encoder augmented with normalizing flows (Rezende and Mohamed, 2015), as well as the recently proposed Wasserstein autoencoder (WAE) (Tolstikhin et al., 2018) which maintains a VAE-like structure. All of these models were adapted to use the same neutral architecture from (Lucic et al., 2018). Note also that the WAE includes two variants, referred to as WAE-MMD and WAE-GAN because different Maximum Mean Discrepancy (MMD) and GAN regularization factors are involved. We conduct experiments using the former because it does not involve potentially-unstable adversarial training, consistent with the other VAE baselines.Later we compare against both WAE-MMD and WAE-GAN using the setup from (Tolstikhin et al., 2018). Additionally, we present results from (Lucic et al., 2018) involving numerous competing GAN models, including MM GAN (Goodfellow et al., 2014), WGAN (Arjovsky et al., 2017), WGAN-GP (Gulrajani et al., 2017), NS GAN (Fedus et al., 2017), DRAGAN (Kodali et al., 2017), LS GAN (Mao et al., 2017) and BEGAN (Berthelot et al., 2017). Testing is conducted across four significantly different datasets: MNIST (LeCun et al., 1998), Fashion MNIST (Xiao et al., 2017), CIFAR-10 (Krizhevsky and Hinton, 2009) and CelebA (Liu et al., 2015).
For each dataset we executed independent trials and report the mean and standard deviation of the FID scores in Table 1.All reported FID scores for VAE and GAN models were computed using TensorFlow (https://github.com/bioinf-jku/TTUR). We have found that alternative PyTorch implementations (https://github.com/mseitzer/pytorch-fid) can produce different values in some circumstances. This seems to be due, at least in part, to subtle differences in the underlying inception models being used for computing the scores. Either way, a consistent implementation is essential for calibrating results across different scenarios. No effort was made to tune VAE training hyperparameters (e.g., learning rates, etc.); rather a single generic setting was first agnostically selected and then applied to all VAE-like models (including the WAE-MMD). As an analogous baseline, we also report the value of the best GAN model (in terms of average FID across all four datasets) when trained without data-dependent tuning using suggested settings from the authors; see (Lucic et al., 2018)[Figure 4] for details. Under these conditions, our single 2-Stage VAE is quite competitive with the best individual GAN model, while the other VAE baselines were not competitive.
Note that the relatively poor performance of the WAE-MMD on MNIST and Fashion MNIST data can be attributed to the sensitivity of this approach to the value of , which for consistency with other models was fixed at for all experiments. This value is likely much larger than actually needed for these simpler data types (meaning ), and the WAE-MMD model can potentially be more reliant on having some . We will return to this issue in Section 5.2 below, where among other things, we empirically examine how performance varies with across different modeling paradigms.
Table 1 also displays FID scores from GAN models evaluated using hyperparameters obtained from a large-scale search executed independently across each dataset to achieve the best results; 100 settings per model per dataset, plus an optimal, data-dependent stopping criteria as described in (Lucic et al., 2018). Within this broader paradigm, cases of severe mode collapse were omitted when computing final GAN FID averages. Despite these considerable GAN-specific advantages, the FID performance of the default 2-Stage VAE is within the range of the heavily-optimized GAN models for each dataset unlike the other VAE baselines. Overall then, these results represent the first demonstration of a VAE pipeline capable of competing with GANs in the arena of generated sample quality using a shared/neutral architecture.
Of course there are still other ways of evaluating sample quality. For example, although the FID score is a widely-adopted measure that significantly improves upon the earlier Inception Score (Salimans et al., 2016), it has been shown to exhibit bias in certain circumstances (Bińkowski et al., 2018). To address this issue, the recently-proposed Kernel Inception Distance (KID) applies a polynomial-kernel MMD measure to estimate the inception distance and is believed to enjoy better statistical properties (Bińkowski et al., 2018). Note that we cannot evaluate all of the GAN baselines using the KID score; only the authors of (Lucic et al., 2018) could easily do this given the huge number of trained models involved that are not publicly available, and the need to retrain selected models multiple times to produce new average scores at optimal hyperparameter settings. However, we can at least compare our trained 2-Stage VAE to other VAE/AE-based models. Table 2 presents these results, where we observe that the same improvement patterns reported with respect to FID are preserved when we apply KID instead, providing further confidence in our approach.
Thus far we have presented performance evaluations of VAE models without any special tuning on a neutral testing platform; however, we have not as of yet compared our approach against state-of-the-art VAE-like competition benefitting from an encoder-decoder architecture and training protocol adjusted for a specific dataset. To the best of our knowledge, the WAE model involves the only published evaluation of this kind (at least at the time of our original posting of this work), where different architectures are designed for use with MNIST and CelebA datasets (Tolstikhin et al., 2018). However, because actual FID scores are only reported for CelebA (which is far more complex than MNIST anyway), we focus our attention on head-to-head testing with this data. In particular, we adopt the exact same encoder-decoder networks as the WAE models, and train using the same number of epochs. We do not tune any hyperparameters whatsoever, and apply the same small second-stage VAE as used in previous experiments. As before, the second-stage size is a small fraction of the first stage, so any benefit is not simply the consequence of a larger network structure. Results are reported in Table 3, where the 2-Stage VAE even outperforms the WAE-GAN model, which has the advantage of adversarial training tuned for this combination of data and network architecture.
2 Experimental Corroboration of Theoretical Results
The true test of any theoretical contribution is the degree to which it leads to useful, empirically-testable predictions about behavior in real-world settings. In the present context, although our theory from Sections 2 and 3 involves some unavoidable simplifying assumptions, it nonetheless makes predictions that can be tested under practically-relevant conditions where these assumptions may not strictly hold. We now present the results of such tests, which provide strong confirmation of our previous analysis. In particular, after providing validation of Theorems 4 and 5, we explicitly demonstrate that the second stage of our 2-Stage VAE model can reduce the gap between q(\mbox{\boldmathz}) and p(\mbox{\boldmathz}), and that the combined VAE stages are quite robust to the value of as predicted.
Validation of Theorem 4: This theorem implies that will converge to zero at any global minimum of the stated VAE objective under consideration. Figure 1(a) presents empirical support for this result, where indeed the decoder variance does tend towards zero during training (red line). This then allows for tighter image reconstructions (dark blue curve) with lower average squared error, i.e., a better manifold fit as expected. Additionally, Figure 1(b) compares the FID score computed using reconstructed training images from different VAE models (these are not new generated samples). The VAE with a learnable achieves the lowest FID on all four datasets, implying that it also produces more realistic overall reconstructions as goes to zero. As will be discussed further in Section 6, high-quality image reconstructions are half the battle in ultimately achieving good generative modeling performance.
Validation of Theorem 5: Figure 2 bolsters this theorem, and the attendant analysis which follows in Section 3, by showcasing the dissimilar impact of noise factors applied to different directions in the latent space before passage through the decoder mean network . In a direction where an eigenvalue of \mbox{\boldmath\Sigma}_{z} is large (i.e., a superfluous dimension), a random perturbation is completely muted by the decoder as predicted. In contrast, in directions where such eigenvalues are small (i.e., needed for representing the manifold), varying the input causes large changes in the image space reflecting reasonable movement along the correct manifold.
Reduced Mismatch between q_{\phi}(\mbox{\boldmathz}) and p(\mbox{\boldmathz}): Although the VAE with a learnable can achieve high-quality reconstructions, the associated aggregated posterior is still likely not close to a standard Gaussian distribution as implied by (10). This mismatch then disrupts the critical ancestral sampling process. As we have previously argued, the proposed 2-Stage VAE has the ability to overcome this issue and achieve a standard Gaussian aggregated posterior, or at least nearly so. As empirical evidence for this claim, Figure 4 displays the singular value spectrum of latent sample matrices \mbox{\boldmathZ}=\{\mbox{\boldmathz}^{(i)}\}_{i=1}^{n} drawn from q_{\phi}(\mbox{\boldmathz}) (first stage), and \mbox{\boldmathU}=\{\mbox{\boldmathu}^{(i)}\}_{i=1}^{n} drawn from q_{\phi^{\prime}}(\mbox{\boldmathu}) (enhanced second stage). As expected, the latter is much closer to the spectrum from an analogous i.i.d. \mathcal{N}(0,\mbox{\boldmathI}) matrix. We also used these same sample matrices to estimate the MMD metric (Gretton et al., 2007) between \mathcal{N}(0,\mbox{\boldmathI}) and the aggregated posterior distributions from the first and second stages in Table 4. Clearly the second stage has dramatically reduced the difference from \mathcal{N}(0,\mbox{\boldmathI}) as quantified by the MMD. Overall, these results indicate a superior latent representation, providing high-level support for our 2-Stage VAE proposal.
Robustness to Latent-Space Dimensionality: According to the analysis from Section 3, the 2-Stage VAE should be relatively insensitive to having a good estimate of the ground-truth manifold dimension . As long as we choose , then our approach should in principle be able to compensate for any dimensionality mismatch. To recap, the first VAE stage should fill in useless dimensions with random noise, such that q_{\phi}(\mbox{\boldmathz}) does not lie on any lower-dimensional manifold. The second stage then operates within the regime arbitrated by Theorem 2 such that we obtain a tractable means for sampling from q_{\phi}(\mbox{\boldmathz}). This robustness to need not be shared by alternative approaches, such as those predicated upon deterministic autoencoders or single-stage VAE models. For example, both the WAE-MMD and WAE-GAN models are dependent on having a reasonable estimate for , at least for the deterministic encoder-decoder structures that were empirically tested in (Tolstikhin et al., 2018). This is because, if , reconstructing the data on the manifold is not possible, and if , then p(\mbox{\boldmathz})=\mathcal{N}(\mbox{\boldmathz}|0,I) and q_{\phi}(\mbox{\boldmathz}) cannot be matched since the latter is necessarily confined by the training data to an -dimensional manifold within -dimensional space. This likely explains the poor performance of WAE-MMD on MNIST and Fashion MNIST data reported in Table 1.
To further probe these issues, we again adopt the neutral testing framework from (Lucic et al., 2018), and retrain each model as is varied. We conduct this experiment using Fashion MNIST (relatively small/simple) and CelebA (more complex). FID scores from both reconstructions of the training data, and novel generated samples are shown in Figure 5. From these plots (left column) we observe that both the 2-Stage VAE and WAE-MMD have a similar reconstruction FID values that become smaller as increases (reconstructions should generally improve with increasing ). Note that the WAE-MMD relies on a Gaussian output layer with a small value,Unlike standard VAEs that are capable of self-calibration in some sense, with and elements of \mbox{\boldmath\Sigma}_{z} jointly pushing towards small values, may be difficult to learn for WAE models because there is also a required weighting factor balancing the MMD (or GAN) loss. Additionally, the WAE code posted in association with (Tolstikhin et al., 2018) does not attempt to learn . and hence the reconstructions are unlikely to be significantly different from the 2-Stage VAE. In contrast, the other baselines are considerably worse, especially when is fixed as is often done in practice.
The situation changes considerably however when we examine the FID scores obtained from generated samples. While the VAE models remain relatively stable over a significant range of sufficiently large values, some of which are likely to be considerably larger than necessary (e.g., on Fashion MNIST data), the WAE-MMD performance is much more sensitive. Of course we readily concede that compensatory, data-dependent tuning of WAE-MMD hyperparameters might allow for some additional improvements in these curves, but this then further highlights the larger point: the VAE models were not tuned in any way for these experiments and yet still maintain stable behavior, with our 2-Stage variant providing a sizeable advantage.
Of course obviously in practice if we set to be far too large, then the training will likely become much more difficult, since in addition to learning the correct ground-truth manifold, we are also burdening the model to detect a much larger number of unnecessary dimensions. But even so, the 2-Stage VAE is arguably quite robust to within these experimental settings, and certainly we need not set to achieve good results; a reasonable appears to be sufficient. Still, as a final caveat, we should mention that the VAE cannot automatically manage all forms of excess model capacity. As discussed in (Dai et al., 2018)[Section 4], if the decoder mean network becomes inordinately complex/deep, then a useless degenerate solution involving the implicit memorization of training data can break even the natural VAE regularization mechanisms we have described herein (in principle, this can happen even with for a finite training dataset). Obviously though, GAN models and virtually all other deep generative pipelines share a similar vulnerability.
3 Qualitative Evaluation of Generated Samples
Finally, we qualitatively evaluate samples generated via our 2-Stage VAE using a simple, convenient residual network structure (with fewer parameters than the InfoGAN architecture). Details of this network are shown in the appendices. Randomly generated samples from our 2-Stage VAE are shown in Figure 6 for MNIST and CelebA data. Additional samples can be found in the appendices.
Discussion
It is often assumed that there exists an unavoidable trade-off between the stable training, valuable attendant encoder network, and resistance to mode collapse of VAEs, versus the impressive visual quality of images produced by GANs. While we certainly are not claiming that our two-stage VAE model is superior to the latest and greatest GAN-based architectures in terms of the realism of generated samples, we do strongly believe that this work at least narrows that gap substantially such that VAEs are worth considering in a broader range of applications. We now close by situating our work within the context of existing VAE enhancements, as well as recent efforts to learn so-called disentangled representations.
Although a variety of alternative VAE renovations have been proposed, unlike our work, nearly all of these have focused on improving the log-likelihood scores assigned by the model to test data. In particular, multiple elegant approaches involve replacing the Gaussian encoder network with a richer class of distributions instantiated through normalizing flows or related (Burda et al., 2015; Kingma et al., 2016; Rezende and Mohamed, 2015; van den Berg et al., 2018). While impressive log-likelihood gains have been demonstrated, this achievement is largely orthogonal to the goal of improving quantitative measures of visual quality (Theis et al., 2016), which has been our focus herein. Additionally, improving the VAE encoder does not address the uniqueness issue raised in Section 2, and therefore, a second stage could potentially benefit these models too under the right circumstances.
Broadly speaking, if the overriding objective is generating realistic samples using an encoder-decoder-based architecture (VAE or otherwise), two important, well-known criteria must be satisfied:In presenting these criteria, we are implicitly excluding models that employ powerful non-Gaussian decoders that can parameterize complex distributions even while excluding any signal from , e.g., PixelCNN-based decoders or related (van den Oord et al., 2016). Such models are considerably different than a canonical AE and do not generally provide a way of accurately reconstructing the training data using a low-dimensional representation.
Small reconstruction error when passing through the encoder-decoder networks, and
An aggregate posterior q_{\phi}(\mbox{\boldmathz}) that is close to some known distribution like p(\mbox{\boldmathz})=\mathcal{N}(\mbox{\boldmathz}|0,I) that is easy to sample from.
The first criteria can be naturally enforced by a deterministic AE, but also for the VAE as becomes small as quantified by Theorem 5. Of course the second criteria is equally important. Without it, we have no tractable way of generating random inputs that, when passed through the learned decoder, produce realistic output samples resembling the training data distribution.
Criteria (i) and (ii) can be addressed in multiple different ways. For example, (Tomczak and Welling, 2018; Zhao et al., 2018) replace \mathcal{N}(\mbox{\boldmathz}|0,I) with a parameterized class of prior distributions such that there exist more flexible pathways for pushing p(\mbox{\boldmathz}) and q_{\phi}(\mbox{\boldmathz}) closer together; a VAE-like objective is used for this purpose in (Tomczak and Welling, 2018), while (Zhao et al., 2018) employs adversarial training. Consequently, even if q_{\phi}(\mbox{\boldmathz}) is not Gaussian, we can nonetheless sample from a known non-Gaussian alternative using either of these approaches. This is certainly an interesting idea, but it has not as of yet been shown to improve FID scores. For example, only log-likelihood values on relatively small black-and-white images are reported in (Tomczak and Welling, 2018), while discrete data and associated evaluation metrics are the focus of (Zhao et al., 2018).
In fact, prior to our work the only competing encoder-decoder-based architecture that explicitly attempts to improve FID scores is the WAE model from (Tolstikhin et al., 2018), which can be viewed as a generalization of the adversarial autoencoder (Makhzani et al., 2016). For both WAE-MMD and WAE-GAN variants introduced and tested in Section 5 herein, the basic idea is to minimize an objective function composed of a reconstruction penalty for handling criteria (i), and a Wassenstein distance measure between p(\mbox{\boldmathz}) and q_{\phi}(\mbox{\boldmathz}) (either MMD- or GAN-based) for addressing criteria (ii). Note that under the reported experimental design from (Tolstikhin et al., 2018), the WAE-GAN model more-or-less defaults to an adversarial autoencoder, although broader differentiating design choices are possible. The adversarially-regularized autoencoder proposed in (Zhao et al., 2018) can also be interpreted as a variant of the WAE-GAN customized to handle discrete data.
As with the approaches mentioned above, the two VAE stages we have proposed can also be motivated in one-to-one correspondence with criteria (i) and (ii). In brief, the first VAE stage addresses criteria (i) by pushing both the encoder variance, and the decoder variances selectively, towards zero such that accurate reconstruction is possible using a minimal number of active latent dimensions. However, our detailed analysis suggests that, although the resulting aggregate posterior q_{\phi}(\mbox{\boldmathz}) will occupy nonzero measure in -dimensional space (selectively filling out superfluous dimensions with random noise), it need not be close to \mathcal{N}(\mbox{\boldmathz}|0,I). This then implies that if we take samples from \mathcal{N}(\mbox{\boldmathz}|0,I) and pass them through the learned decoder, the result may not closely resemble real data.
Of course if we could somehow directly sample from q_{\phi}(\mbox{\boldmathz}), then we would not need to use \mathcal{N}(\mbox{\boldmathz}|0,I). And fortunately, because the first-stage VAE ensures that q_{\phi}(\mbox{\boldmathz}) will satisfy the conditions of Theorem 2, we know that a second VAE can in fact be learned to accurately sample from this distribution, which in turn addresses criteria (ii). Specifically, per the arguments from Section 4, sampling \mbox{\boldmathu}\sim\mathcal{N}(\mbox{\boldmathu}|{\bf 0},\mbox{\boldmathI}) and then \mbox{\boldmathz}\sim p_{\theta^{\prime}}(\mbox{\boldmathz}|\mbox{\boldmathu}) is akin to sampling \mbox{\boldmathz}\sim q_{\phi}(\mbox{\boldmathz}) even though the latter is not available in closed form. Such samples can then be passed through the first-stage VAE decoder to obtain samples of . Hence our framework provides a principled alternative to existing encoder-decoder structures designed to handle criteria (i) and (ii), leading to state-of-the-art results for this class of model in terms of FID scores under neutral testing conditions.
Finally, before closing this section, it is also informative to consider the vector-quantization VAE (VQ-VAE) from (van den Oord et al., 2017). This model of discrete distributions involves first training a deterministic autoencoder with a latent space parameterized by a vector quantization operator. Later, a PixelCNN (van den Oord et al., 2016) or related is trained to approximate the aggregated posterior of the corresponding discrete latent codes. In this sense the VQ-VAE is effectively satisfying criteria (ii), at least in the regime of discrete distributions. However, in (van den Oord et al., 2017) it is also mentioned that joint training of the VQ-VAE reconstruction module and the PixelCNN aggregated posterior estimate could improve performance. While this may be true for discrete distributions (verification is left to future work in (van den Oord et al., 2017)), at least for continuous data lying on a low-dimensional manifold as has been our focus, we reiterate that joint training of our two proposed continuous VAE stages is likely to degrade performance. This is for the detailed reasons provided in Section 4. In this regard, there can seemingly exist non-trivial differences in the properties of generative models of discrete versus continuous data.
2 Identifiability of Disentangled Representations
A number of VAE modifications have been putatively designed to encourage this type of disentanglement (likewise for other generative modeling paradigms). These efforts can be partitioned into those that involve some degree of supervision, to both define and isolate specific ground-truth factors of variation (which may be application-specific), and those that aspire to be completely unsupervised. For the latter, the VAE is typically refashioned to minimize, at least to the extent possible, some loose proxy for the degree of entanglement. An influential example is the total correlation,This naming convention is perhaps a misnomer, since total correlation measures statistical dependency beyond second-order correlations. defined for the present context as
where denotes the marginal distribution of . It follows that iff q_{\phi}(\mbox{\boldmathz})=\prod_{j}q_{\phi}(z_{j}), meaning that the aggregate posterior distribution, which generates input samples to the VAE decoder, involves independent factors of variation. It has been argued then that adjustments to the VAE objective that push the total correlation towards zero will favor desirable disentangled representations (Chen et al., 2018; Higgins et al., 2017). Note that if criteria (ii) from Section 6.1 is satisfied with a factorial prior p(\mbox{\boldmathz})=\prod_{j}p(z_{j}) such as \mathcal{N}(\mbox{\boldmathz}|0,I), then we have actually already achieved . In contrast, models with more complex priors as used in (Tomczak and Welling, 2018) would require further alterations to enforce this goal.
While we agree that methods penalizing TC (either directly or indirectly) will favor latent representations with independent dimensions, unfortunately this need not correspond with any semantically-meaningful form of disentanglement. Intuitively, this means that independent latent factors need not correspond with interpretable concepts like stroke width or digit slope. In fact, without at least some form of supervision or constraints on the space of possible representations, disentanglement defined with respect to any particular ground-truth semantic factors is not generally identifiable in the strict statistical sense.
As one easy way to see this, suppose that we have access to data generated as \mbox{\boldmathx}=f_{gt}(\mbox{\boldmathz}), where serves as an arbitrary ground-truth decoder and \mbox{\boldmathz}\sim\prod_{j}p_{gt}(z_{j}) represents ground-truth/disentangled latent factors of interest. However, we could just as well generate identical data via \mbox{\boldmathx}=f_{gt}\left(\mbox{\boldmathD}_{1}^{-1}\left[\mbox{\boldmathR}^{\top}\mbox{\boldmathD}_{2}^{-1}\left(\widetilde{\mbox{\boldmathz}}\right)\right]\right), where \widetilde{\mbox{\boldmathz}}\triangleq\mbox{\boldmathD}_{2}\left[\mbox{\boldmathR}\mbox{\boldmathD}_{1}(\mbox{\boldmathz})\right] with associated factorial distribution . Here the operator \mbox{\boldmathD}_{1} converts to a standardized Gaussian distribution, is an arbitrary rotation matrix the mixes the factors but retains a factorial representation with , and \mbox{\boldmathD}_{2} converts the resulting Gaussian to a new factorial distribution with arbitrary marginals.
By construction, each new latent factor will be composed as a mixture of the original factors of interest, and yet they will also have zero total correlation, i.e., they will be statistically independent. Therefore if we treat the composite operator \widetilde{f}(\cdot)\triangleq f_{gt}\left(\mbox{\boldmathD}_{1}^{-1}\left[\mbox{\boldmathR}^{\top}\mbox{\boldmathD}_{2}^{-1}\left(\cdot\right)\right]\right) as the effective decoder, we have a new generative process \mbox{\boldmathx}=\widetilde{f}(\widetilde{\mbox{\boldmathz}}) and \widetilde{\mbox{\boldmathz}}\sim\prod_{j}\widetilde{p}(\widetilde{z}_{j}). By construction, this process will produce identical observed data using a latent representation with zero total correlation, but with highly entangled factors with respect to the original ground truth delineation. Note also that the revised decoder need not necessarily be more complex than , except in special circumstances. For example, if we constrain our decoder architecture to be affine, then the model defaults to independent component analysis and the nonlinear operators \mbox{\boldmathD}_{1} and \mbox{\boldmathD}_{2} cannot be absorbed into an affine composite operator. In this case, the model is identifiable up to an arbitrary permutation and scaling (Hyvärinen and Oja, 2000).
To visualize this, let \mbox{\boldmathz}=[z_{1},z_{2}]^{\top} denote a 2D binary vector that serves as the latent representation. In Table 4 we present two candidate encodings, both of which display zero total correlation between factors and (i.e., knowledge of tells us nothing about the value of and vice versa). However, in terms of the semantically meaningful factors of gender and age, these two encodings are quite different. In the first, varying alters gender independently of age, while varies age independently of gender. In contrast, with the second encoding varying while keeping fixed changes both the age and gender, an entangled representation per these attributes. Without supervising labels to resolve such ambiguity, there is no possible way, even in principle, for any generative model (VAE or otherwise) to differentiate these two encodings such that the so-called disentangled representation is favored over the entangled one.
Although in some limited cases, it has been reported that disentangled representations can be at least partially identified, we speculate that this is a result of nuanced artifacts of the experimental design that need not translate to broader problems of interest. Indeed, in our own empirical tests we have not been able to consistently reproduce any stable disentangled representation across various testing conditions as expected. It was our original intent to include a thorough presentation of these results; however, in the process of preparing this manuscript we became aware of a contemporary work with extensive testing in this area (Locatello et al., 2018). In full disclosure, the results from (Locatello et al., 2018) are actually far more extensive than those that we have completed, so we simply defer the reader to this paper to examine the compelling battery of tests presented there.
A Comparison of Novel Samples Generated from our Model
Generation results for CelebA, MNIST, Fashion-MNIST and CIFAR-10 datasets of different methods are shown in Figures 710 respectively. When is fixed to be one, the generated samples are very blurry. If a learnable is used, the samples becomes sharper; however, there are many lingering artifacts as expected. In contrast, the proposed 2-Stage VAE can remove these artifacts and generate more realistic samples. For comparison purposes, we also show the results from WAE-MMD, WAE-GAN (Tolstikhin et al., 2018) and WGAN-GP (Gulrajani et al., 2017) for the CelebA dataset.
B Example Reconstructions of Training Data
Reconstruction results for MNIST, Fashion-MNIST, CIFAR-10 and CelebA datasets are shown in Figures 1114 respectively. On relatively simple datasets like MNIST and Fashion-MNIST, the VAE with learnable achieves almost exact reconstruction because of a better estimate of the underlying manifold consistent with theory. However, the VAE with fixed produces blurry reconstructions as expected. Note that the reconstruction of a 2-Stage VAE is the same as that of a VAE with learnable because the second-stage VAE has nothing to do with facilitating the reconstruction task.
C Additional Experimental Results Validating Theoretical Predictions
We first present more examples similar to Figure 2 from the main paper. Random noise is added to \mbox{\boldmath\mu}_{z} along different directions and the result is passed through the decoder network. Each row corresponds to a certain direction in the latent space and samples are shown for each direction. These dimensions/rows are ordered by the eigenvalues of \mbox{\boldmath\Sigma}_{z}. The larger is, the less impact a random perturbation along this direction will have as quantified by the reported image variance values. In the first two or three rows, the noise generates some images from different classes/objects/identities, indicating a significant visual difference. For a slightly larger , the corresponding dimensions encode relatively less significant attributes as predicted. For example, the fifth row of both MNIST and Fashion-MNIST contains images from the same class but with a slightly different style. The images in the fourth row of the CelebA dataset have very subtle differences. When , the corresponding dimensions become completely inactive and all the output images are exactly the same, as shown in the last rows for all the three datasets.
Additionally, as discussed in the main text and below in Section I, there are likely to be eigenvalues of \mbox{\boldmath\Sigma}_{z} converging to zero and eigenvalues converging to one. We plot the histogram of values for both MNIST and CelebA datasets in Figure 16. For both datasets, approximately converges to either to zero or one. However, since CelebA is a more complicated dataset than MNIST, the ground-truth manifold dimension of CelebA is likely to be much larger than that of MNIST. So more eigenvalues are expected to be near zero for the CelebA dataset. This is indeed the case, demonstrating that VAE has the ability to detect the manifold dimension and select the proper number of latent dimensions in practical environments.
D Network Structure and Experimental Settings
We first describe the network and training details used in producing Figure 6 from the main file, and for generating samples and reconstructions in the appendix. The first-stage VAE network is shown in Figure 17. Basically we use two Residual Blocks for each resolution scale, and we double the number of channels when downsampling and halve it when upsampling. The specific settings such as the number of channels and the number of scales are specified in the caption. The second VAE is much simpler. Both the encoder and decoder have three -dimensional hidden layers. Finally, the training details are presented below. Note that these settings were not tuned, we simply chose more epochs for more complex data sets and fewer for datasets with larger training samples. For each dataset just a single setting was tested as follows:
MNIST and Fashion-MNIST: The batch size is specified to be 100. We use the ADAM optimizer with the default hyperparameters in TensorFlow. The first VAE is trained for epochs. The initial learning rate is and we halve it every epochs. The second VAE is trained for epochs with the same initial learning rate, halved every epochs.
CIFAR-10: Since CIFAR-10 is more complicated than MNIST and Fashion-MNIST, we use more epochs for training. Specifically, we use and epochs for the two VAEs respectively and half the learning rate every and epochs for the two stages. The other settings are the same as that for MNIST.
CelebA: Because CelebA has many more examples, in the first stage we train epochs and half the learning rate every epochs. In the second stage, we train epochs and half the learning rate every epochs. The other settings are the same as that for MNIST, etc.
Finally, to fairly compare against various GAN models and VAE baselines using FID scores on a neutral architecture (i.e., the results from Table 1), we simply adopt the InfoGAN network structure consistent with the neutral setup from (Lucic et al., 2018) for the first-stage VAE. For the second-stage VAE we just use three -dimensional hidden layers, which contribute less than to the total number of parameters. Note that the small number of additional parameters contributing to the second stage do not improve the other VAE baselines when aggregated and trained jointly.
E Proof of Theorem 2
Additionally, let \mbox{\boldmath\xi}=G(\mbox{\boldmathz}) such that
and let \mbox{\boldmathx}^{\prime}=F^{-1}(\mbox{\boldmath\xi}) such that d\mbox{\boldmath\xi}=dF(\mbox{\boldmathx}^{\prime})=p_{gt}(\mbox{\boldmathx}^{\prime})d\mbox{\boldmathx}^{\prime}. Plugging this expression into the previous p_{\theta^{*}}(\mbox{\boldmathx}) we obtain
As , becomes infinitely small and \mathcal{N}\left(\mbox{\boldmathx}|\mbox{\boldmathx}^{\prime},\gamma^{*}_{t}\mbox{\boldmathI}\right) becomes a Dirac-delta function, resulting in
where is a Jacobian matrix. We omit the arguments and in , and hereafter to avoid unnecessary clutter. We first explain why is differentiable. Since is a composition of and according to (18), we only need to explain that both functions are differentiable. For , it is the inverse of a differentiable function . Moreover, the derivative of F(\mbox{\boldmathx}) is p_{gt}(\mbox{\boldmathx}), which is nonzero everywhere. So and therefore are both differentiable.
The true posterior p_{\theta^{*}_{t}}(\mbox{\boldmathz}|\mbox{\boldmathx}) and the approximate posterior are
respectively. We first transform to \mbox{\boldmathz}^{\prime} via
where \mbox{\boldmathz}^{*}=f_{\mu_{z}}(\mbox{\boldmathx}). The true and approximate posteriors p_{\theta^{*}_{t}}(\mbox{\boldmathz}|\mbox{\boldmathx}) and q_{\phi^{*}_{t}}(\mbox{\boldmathz}|\mbox{\boldmathx}) will then be transformed into another two distributions of \mbox{\boldmathz}^{\prime}, namely p^{\prime}_{\theta^{*}_{t}}(\mbox{\boldmathz}^{\prime}|\mbox{\boldmathx}) and q^{\prime}_{\phi^{*}_{t}}(\mbox{\boldmathz}|\mbox{\boldmathx}). The KL divergence between p_{\theta^{*}_{t}}(\mbox{\boldmathz}|\mbox{\boldmathx}) and q_{\phi^{*}_{t}}(\mbox{\boldmathz}|\mbox{\boldmathx}) will be the same as that between p^{\prime}_{\theta^{*}_{t}}(\mbox{\boldmathz}^{\prime}|\mbox{\boldmathx}) and q^{\prime}_{\phi^{*}_{t}}(\mbox{\boldmathz}|\mbox{\boldmathx}). We now prove that q^{\prime}_{\phi^{*}_{t}}(\mbox{\boldmathz}^{\prime}|\mbox{\boldmathx})/p^{\prime}_{\theta^{*}_{t}}(\mbox{\boldmathz}^{\prime}|\mbox{\boldmathx}) converges to a constant not related to \mbox{\boldmathz}^{\prime} as goes to . If this is true, the constant must be since both q^{\prime}_{\phi^{*}_{t}}(\mbox{\boldmathz}^{\prime}|\mbox{\boldmathx}) and p^{\prime}_{\theta^{*}_{t}}(\mbox{\boldmathz}^{\prime}|\mbox{\boldmathx}) are probability distributions. Then the KL divergence between them converges to as .
E.3 Generalization to the Case with κ>r𝜅𝑟\kappa>r
F Proof of Theorem 3
Now let the decoder mean function be given by
We next show that p_{\theta^{*}_{t}}(\mbox{\boldmathx}) diverges to infinite as for any . For a given , let \mbox{\boldmathu}^{*}=\varphi(\mbox{\boldmathx}) and B(\mbox{\boldmathu}^{*},\sqrt{\gamma^{*}_{t}}) be the closed ball centered at \mbox{\boldmathu}^{*} with radius . Then
According to the Lagrangian’s mean value theorem, there exists a \mbox{\boldmathu}^{\prime} between and \mbox{\boldmathu}^{*} such that
If we denote \Lambda(\mbox{\boldmathu}^{\prime})=\left(\frac{d\varphi^{-1}(\mbox{\boldmathu})}{d\mbox{\boldmathu}}|_{\mbox{\boldmathu}=\mbox{\boldmathu}^{\prime}}\right)^{\top}\left(\frac{d\varphi^{-1}(\mbox{\boldmathu})}{d\mbox{\boldmathu}}|_{\mbox{\boldmathu}=\mbox{\boldmathu}^{\prime}}\right), we then have that
where V\left(B(\mbox{\boldmathu}^{*},\sqrt{\gamma^{*}_{t}})\right) is the volume of the -dimensional ball B(\mbox{\boldmathu}^{*},\sqrt{\gamma^{*}}). The volume should be where is a constant related to the dimension . So
Since defines a diffeomorphism, D(\mbox{\boldmathu}^{*})<\infty. Moreover, \left(\min_{\mbox{\boldmathu}\in B(\mbox{\boldmathu}^{*},1)}p_{gt}^{u}(\mbox{\boldmathu})\right)>0 because p_{gt}^{u}(\mbox{\boldmathu}) is nonzero and continuous everywhere. We may then conclude that
for \mbox{\boldmathx}\in{\mathcal{X}}. This then implies that the stated average across with respect to will also be .
Similar to (24) and (25), let the encoder be
Following the proofs in Section E.2, we can prove the KL divergence between q_{\phi^{*}_{t}}(\mbox{\boldmathz}|\mbox{\boldmathx}) and p_{\theta^{*}_{t}}(\mbox{\boldmathz}|\mbox{\boldmathx}) converges to .
We begin with the probability distribution given by p_{\theta_{t}^{*}}(\mbox{\boldmathx}):
The second equation that interchanges the order of the integrations admitted by Fubini’s theorem. The third equation that interchanges the order of the integration and the limit is justified by the bounded convergence theorem. We now note that the term inside the first integration, \mathcal{N}(\mbox{\boldmathx}|\mbox{\boldmathx}^{\prime},\gamma_{t}^{*}\mbox{\boldmathI}), converges to a Dirac-delta function as . So the integration over depends on whether \mbox{\boldmathx}^{\prime} is inside or not, i.e.,
We separate the manifold into three parts: \mbox{\boldmath\chi}\cap\left(\mbox{\boldmathA}-\partial\mbox{\boldmathA}\right), \mbox{\boldmath\chi}\cap\left(\mbox{\boldmathA}^{c}-\partial\mbox{\boldmathA}\right) and \mbox{\boldmath\chi}\cap\partial\mbox{\boldmathA}. Then (55) can be separated into three parts accordingly. The first two parts can be derived as
For the third part, given the assumption that \mu_{gt}(\partial\mbox{\boldmathA})=0, we have
Note that this result involves a subtle requirement involving the boundary . This condition is only included to handle a minor, practically-inconsequential technicality. In brief, as a density p_{\theta_{t}^{*}}(\mbox{\boldmathx}) will apply zero mass exactly on any low-dimensional manifold, although it can apply all of its mass to any region in the neighborhood of . But suppose we choose some is a subset of , i.e, it is exclusively confined to the ground-truth manifold. Then the probability mass within assigned by will be nonzero while that given by p_{\theta_{t}^{*}}(\mbox{\boldmathx}) can still be zero. Of course this does not mean that p_{\theta_{t}^{*}}(\mbox{\boldmathx}) and do not match each other in any practical sense. This is because if we expand this specialized by an arbitrary small -dimensional volume, then p_{\theta_{t}^{*}}(\mbox{\boldmathx}) and will now supply essentially the same probability mass on this infinitesimally expanded set (which is arbitrary close to ).
G Proof of Theorem 4
From the main text, is the optimal solution with a fixed . The true posterior and the approximate posterior are
We first argue that the KL divergence between p_{\theta^{*}_{\gamma}}(\mbox{\boldmathz}|\mbox{\boldmathx}) and q_{\phi^{*}_{\gamma}}(\mbox{\boldmathz}|\mbox{\boldmathx}) is always strictly greater than zero. This can be proved by contradiction. Suppose the KL divergence exactly equals zero. Then p_{\theta^{*}_{\gamma}}(\mbox{\boldmathz}|\mbox{\boldmathx}) must also be a Gaussian distribution, meaning that the logarithm of p_{\theta^{*}_{\gamma}}(\mbox{\boldmathz}|\mbox{\boldmathx}) is a quadratic form in . In particular, we have
where we have absorbed all the terms not related to into a constant, and it must be that
for some matrix and vector . Then we have
The first inequality comes from the fact that is the optimal solution when is fixed at while is just one solution with . The second inequality holds because we chose to be a better solution than .
G.2 Case 2: r<d𝑟𝑑r<d
The first inequality holds discarding the KL term, which is non-negative. The second inequality holds because a quadratic term is removed. Furthermore, according to the proof in Section F, there exists a such that for any ,
Again, we select a such that and let . Then
H Proof of Theorem 5
Plugging these expressions into (2) we obtain
where we have omitted explicit inclusion of the parameters and in the functions and to avoid undue clutter. Now suppose
which contradicts the fact that converges to . So we must have that
Because the term inside the expectation, i.e., \left|\left|f_{\mu_{x}}\left[f_{\mu_{z}}(\mbox{\boldmathx})+f_{S_{z}}(\mbox{\boldmathx})\epsilon\right]-\mbox{\boldmathx}\right|\right|_{2}^{2}, is always non-negative, we can conclude that
And if we let , this equation then becomes
I Further Analysis of the VAE Cost as γ𝛾\gamma becomes small
In the main paper, we mentioned that the squared eigenvalues of f_{S_{z}}(\mbox{\boldmathx};\phi_{\gamma}^{*}) will become arbitrary small at a rate proportional to . To justify this, we borrow the simplified notation from the proof of Theorem 5 and expand f_{\mu_{x}}(\mbox{\boldmathz}) at \mbox{\boldmathz}=f_{\mu_{z}}(\mbox{\boldmathx}) using a Taylor series. Omitting the high order terms (in the present narrow context around the neighborhood of VAE global optima these will be small), this gives
Plug this expression and (71) into (73), we obtain
From these manipulations we may conclude that the optimal value of f_{S_{z}}(\mbox{\boldmathx})f_{S_{z}}(\mbox{\boldmathx})^{\top} must satisfy
Note that f_{\mu_{x}}^{\prime}\left[f_{\mu_{z}}(\mbox{\boldmathx})\right] is the tangent space of the manifold at f_{\mu_{x}}\left[f_{\mu_{z}}(\mbox{\boldmathx})\right], so the rank must be . f_{\mu_{x}}^{\prime}\left[f_{\mu_{z}}(\mbox{\boldmathx})\right]^{\top}f_{\mu_{x}}^{\prime}\left[f_{\mu_{z}}(\mbox{\boldmathx})\right] can be decomposed as \mbox{\boldmathU}^{\top}\mbox{\boldmathS}\mbox{\boldmathU}, where is a -dimensional orthogonal matrix and is a -dimensional diagonal matrix with nonzero elements. Denote \text{diag}[\mbox{\boldmathS}]=[S_{1},S_{2},...,S_{r},0,...,0]. Then
Case 1: . In this case, has no nonzero diagonal elements, and therefore
As , the eigenvalues of f_{S_{z}}(\mbox{\boldmathx})f_{S_{z}}(\mbox{\boldmathx})^{\top}, which are given by , converge to at a rate of .
Case 2: . In this case, the first eigenvalues also converge to at a rate of , but the remaining eigenvalues will be , meaning the redundant dimensions are simply filled with noise matching the prior p(\mbox{\boldmathz}) as desired.