Flow++: Improving Flow-Based Generative Models with Variational Dequantization and Architecture Design

Jonathan Ho, Xi Chen, Aravind Srinivas, Yan Duan, Pieter Abbeel

Introduction

Deep generative models – latent variable models in the form of variational autoencoders (Kingma & Welling, 2013), implicit generative models in the form of GANs (Goodfellow et al., 2014), and exact likelihood models like PixelRNN/CNN (van den Oord et al., 2016b, c), Image Transformer (Parmar et al., 2018), PixelSNAIL (Chen et al., 2017), NICE, RealNVP, and Glow (Dinh et al., 2014, 2016; Kingma & Dhariwal, 2018) – have recently begun to successfully model high dimensional raw observations from complex real-world datasets, from natural images and videos, to audio signals and natural language (Karras et al., 2017; Kalchbrenner et al., 2016b; van den Oord et al., 2016a; Kalchbrenner et al., 2016a; Vaswani et al., 2017).

Autoregressive models, a certain subclass of exact likelihood models, achieve state-of-the-art density estimation performance on many challenging real-world datasets, but generally suffer from slow sampling time due to their autoregressive structure (van den Oord et al., 2016b; Salimans et al., 2017; Chen et al., 2017; Parmar et al., 2018). Inverse autoregressive models can sample quickly and potentially have strong modeling capacity, but they cannot be trained efficiently by maximum likelihood (Kingma et al., 2016). Non-autoregressive flow-based models (which we will refer to as “flow models”), such as NICE, RealNVP, and Glow, are efficient for sampling, but have so far lagged behind autoregressive models in density estimation benchmarks (Dinh et al., 2014, 2016; Kingma & Dhariwal, 2018).

In the hope of creating an ideal likelihood-based generative model that simultaneously has fast sampling, fast inference, and strong density estimation performance, we seek to close the density estimation performance gap between flow models and autoregressive models. In subsequent sections, we present our new flow model, Flow++, which is powered by an improved training procedure for continuous likelihood models and a number of architectural extensions of the coupling layer defined by Dinh et al. (2014, 2016).

Flow Models

A flow model ff is constructed as an invertible transformation that maps observed data x\mathbf{x} to a standard Gaussian latent variable z=f(x)\mathbf{z}=f(\mathbf{x}), as in nonlinear independent component analysis (Bell & Sejnowski, 1995; Hyvärinen et al., 2004; Hyvärinen & Pajunen, 1999). The key idea in the design of a flow model is to form ff by stacking individual simple invertible transformations (Dinh et al., 2014, 2016; Kingma & Dhariwal, 2018; Rezende & Mohamed, 2015; Kingma et al., 2016; Louizos & Welling, 2017). Explicitly, ff is constructed by composing a series of invertible flows as f(x)=f1∘⋯∘fL(x)f(\mathbf{x})=f_{1}\circ\dotsb\circ f_{L}(\mathbf{x}), with each fif_{i} having a tractable inverse and a tractable Jacobian determinant. This way, sampling is efficient, as it can be performed by computing f−1(z)=fL−1∘⋯∘f1−1(z)f^{-1}(\mathbf{z})=f_{L}^{-1}\circ\dotsb\circ f_{1}^{-1}(\mathbf{z}) for z∼N(0,I)\mathbf{z}\sim\mathcal{N}(\mathbf{0},\mathbf{I}), and so is training by maximum likelihood, since the model density

is easy to compute and differentiate with respect to the parameters of the flows fif_{i}.

Flow++

In this section, we describe three modeling inefficiencies in prior work on flow models: (1) uniform noise is a suboptimal dequantization choice that hurts both training loss and generalization; (2) commonly used affine coupling flows are not expressive enough; (3) convolutional layers in the conditioning networks of coupling layers are not powerful enough. Our proposed model, Flow++, consists of a set of improved design choices: (1) variational flow-based dequantization instead of uniform dequantization; (2) logistic mixture CDF coupling flows; (3) self-attention in the conditioning networks of coupling layers.

Many real-world datasets, such as CIFAR10 and ImageNet, are recordings of continuous signals quantized into discrete representations. Fitting a continuous density model to discrete data, however, will produce a degenerate solution that places all probability mass on discrete datapoints (Uria et al., 2013). A common solution to this problem is to first convert the discrete data distribution into a continuous distribution via a process called “dequantization,” and then model the resulting continuous distribution using the continuous density model (Uria et al., 2013; Dinh et al., 2016; Salimans et al., 2017).

Consequently, maximizing the log-likelihood of the continuous model on uniformly dequantized data cannot lead to the continuous model degenerately collapsing onto the discrete data, because its objective is bounded above by the log-likelihood of a discrete model.

1.2 Variational dequantization

We will choose qq itself to be a conditional flow-based generative model of the form u=qx(ϵ)\mathbf{u}=q_{\mathbf{x}}(\bm{\epsilon}), where ϵ∼p(ϵ)=N(ϵ;0,I)\bm{\epsilon}\sim p(\bm{\epsilon})=\mathcal{N}(\bm{\epsilon};\mathbf{0},\mathbf{I}) is Gaussian noise. In this case, q(u∣x)=p(qx−1(u))⋅∣∂qx−1/∂u∣q(\mathbf{u}|\mathbf{x})=p(q_{\mathbf{x}}^{-1}(\mathbf{u}))\cdot\left|\partial q_{\mathbf{x}}^{-1}/\partial\mathbf{u}\right|, and thus we obtain the objective

2 Improved coupling layers

Recent progress in the design of flow models has involved carefully constructing flows to increase their expressiveness while preserving tractability of the inverse and Jacobian determinant computations. One example is the invertible 1×11\times 1 convolution flow, whose inverse and Jacobian determinant can be calculated and differentiated with standard automatic differentiation libraries (Kingma & Dhariwal, 2018). Another example, which we build upon in our work here, is the affine coupling layer (Dinh et al., 2016). It is a parameterized flow y=fθ(x)\mathbf{y}=f_{\theta}(\mathbf{x}) that first splits the components of x\mathbf{x} into two parts x1,x2\mathbf{x}_{1},\mathbf{x}_{2}, and then computes y=(y1,y2)\mathbf{y}=(\mathbf{y}_{1},\mathbf{y}_{2}), given by

Here, aθ\mathbf{a}_{\theta} and bθ\mathbf{b}_{\theta} are outputs of a neural network that acts on x1\mathbf{x}_{1} in a complex, expressive manner, but the resulting behavior on x2\mathbf{x}_{2} always remains an elementwise affine transformation – effectively, aθ\mathbf{a}_{\theta} and bθ\mathbf{b}_{\theta} together form a data-parameterized family of invertible affine transformations. This allows the affine coupling layer to express complex dependencies on the data while keeping inversion and log-likelihood computation tractable. Using ⋅\cdot and exp⁡\exp to respectively denote elementwise multiplication and exponentiation, the affine coupling layer is defined by:

The splitting operation x↦(x1,x2)\mathbf{x}\mapsto(\mathbf{x}_{1},\mathbf{x}_{2}) and merging operation (y1,y2)↦y(\mathbf{y}_{1},\mathbf{y}_{2})\mapsto\mathbf{y} are usually performed over channels or over space in a checkerboard-like pattern (Dinh et al., 2016).

We found in our experiments that density modeling performance of these coupling layers could be improved by augmenting the data-parameterized elementwise affine transformations by more general nonlinear elementwise transformations. For a given scalar component xx of x2\mathbf{x}_{2}, we apply the cumulative distribution function (CDF) for a mixture of KK logistics – parameterized by mixture probabilities, means, and log scales π,μ,s\bm{\pi},\bm{\mu},\mathbf{s} – followed by an inverse sigmoid and an affine transformation parameterized by aa and bb:

The transformation parameters π,μ,s,a,b\bm{\pi},\bm{\mu},\mathbf{s},a,b for each component of x2\mathbf{x}_{2} are produced by a neural network acting on x1\mathbf{x}_{1}. This neural network must produce these transformation parameters for each component of x2\mathbf{x}_{2}, hence it produces vectors aθ(x1)\mathbf{a}_{\theta}(\mathbf{x}_{1}) and bθ(x1)\mathbf{b}_{\theta}(\mathbf{x}_{1}) and tensors πθ(x1),μθ(x1),sθ(x1)\bm{\pi}_{\theta}(\mathbf{x}_{1}),\bm{\mu}_{\theta}(\mathbf{x}_{1}),\mathbf{s}_{\theta}(\mathbf{x}_{1}) (with last axis dimension KK). The coupling transformation is then given by:

where the formula for computing y2\mathbf{y}_{2} operates elementwise.

The inverse sigmoid ensures that the inverse of this coupling transformation always exists: the range of the logistic mixture CDF is (0,1)(0,1), so the domain of its inverse must stay within this interval. The CDF itself can be inverted efficiently with bisection, because it is a monotonically increasing function. Moreover, the Jacobian determinant of this transformation involves calculating the probability density function of the logistic mixtures, which poses no computational difficulty.

2.2 Expressive conditioning architectures with self-attention

In addition to improving the expressiveness of the elementwise transformations on x2\mathbf{x}_{2}, we found it crucial to improve the expressiveness of the conditioning on x1\mathbf{x}_{1} – that is, the expressiveness of the neural network responsible for producing the elementwise transformation parameters π,μ,s,a,b\bm{\pi},\bm{\mu},\mathbf{s},\mathbf{a},\mathbf{b}. Our best results were obtained by stacking convolutions and multi-head self attention into a gated residual network (Mishra et al., 2018; Chen et al., 2017), in a manner resembling the Transformer (Vaswani et al., 2017) with pointwise feedforward layers replaced by 3×33\times 3 convolutional layers. Our architecture is defined as a stack of blocks. Each block consists of the following two layers connected in a residual fashion, with layer normalization (Ba et al., 2016) after each residual connection:

With these blocks in hand, the network that outputs the elementwise transformation parameters is simply given by stacking blocks on top of each other, and finishing with a final convolution that increases the number of channels to the amount needed to specify the elementwise transformation parameters.

Experiments

Here, we show that Flow++ achieves state-of-the-art density modeling performance among non-autoregressive models on CIFAR10 and 32x32 and 64x64 ImageNet. We also present ablation experiments that quantify the improvements proposed in section 3, and we present example generative samples from Flow++ and compare them against samples from autoregressive models.

Our experiments employed weight normalization and data-dependent initialization (Salimans & Kingma, 2016). We used the checkerboard-splitting, channel-splitting, and downsampling flows of Dinh et al. (2016); we also used before every coupling flow an invertible 1x1 convolution flows of Kingma & Dhariwal (2018), as well as a variant of their “actnorm” flow that normalizes all activations independently (instead of normalizing per channel). Our CIFAR10 model used 4 coupling layers with checkerboard splits at 32x32 resolution, 2 coupling layers with channel splits at 16x16 resolution, and 3 coupling layers with checkerboard splits at 16x16 resolution; each coupling layer used 10 convolution-attention blocks, all with 96 filters. More details on architectures, as well as details for the other experiments, are in our source code release.

In table 1, we show that Flow++ achieves state-of-the-art density modeling results out of all non-autoregressive models, and it is competitive with autoregressive models: its performance is on par with the first generation of PixelCNN models (van den Oord et al., 2016b), and it outperforms Multiscale PixelCNN (Reed et al., 2017). Our results are reported using 16384 importance samples in our CIFAR experiment and 1 sample in our ImageNet experiments (Burda et al., 2015). With 1 sample only, our CIFAR model attains 3.12 bits/dim. Our listed ImageNet 32x32 and 64x64 results are evaluated on a NVIDIA DGX-1; they are worse by 0.01 bits/dim when evaluated on a NVIDIA Titan X GPU.

2 Ablations

We ran the following ablations of our model on unconditional CIFAR10 density estimation: variational dequantization vs. uniform dequantization; logistic mixture coupling vs. affine coupling; and stacked self-attention vs. convolutions only. As each ablation involves removing some component of the network, we increased the number of filters in all convolutional layers (and attention layers, if present) in order to match the total number of parameters with the full Flow++ model.

In fig. 1 and table 2, we compare the performance of these ablations relative to Flow++ at 400 epochs of training, which was not enough for these models to converge, but far enough to see their relative performance differences. Switching from our variational dequantization to the more standard uniform dequantization costs the most: approximately 0.1270.127 bits/dim. The remaining two ablations both cost approximately 0.030.03 bits/dim: switching from our logistic mixture coupling layers to affine coupling layers, and switching from our hybrid convolution-and-self-attention architecture to a pure convolutional residual architecture. Note that these performance differences are present despite all networks having approximately the same number of parameters: the improved performance of Flow++ comes from improved inductive biases, not simply from increased parameter count.

The most interesting result is probably the effect of the dequantization scheme on training and generalization loss. At 400 epochs of training, the full Flow++ model with variational dequantization has a train-test gap of approximately 0.02 bits/dim, but with uniform dequantization, the train-test gap is approximately 0.06 bits/dim. This confirms our claim in Section 3.1.2 that training with variational dequantization is a more natural task for the model than training with uniform dequantization.

3 Samples

We present the samples from our trained density models of Flow++ on CIFAR10, 32x32 ImageNet, 64x64 ImageNet, and 5-bit CelebA in figs. 2, 3, 5 and 4. The Flow++ samples match the perceptual quality of PixelCNN samples, showing that Flow++ captures both local and global dependencies as well as PixelCNN and is capable of generating diverse samples on large datasets. Moreover, sampling is fast: our CIFAR10 model takes approximately 0.32 seconds to generate a batch of 8 samples in parallel on one NVIDIA 1080 Ti GPU, making it more than an order of magnitude faster than PixelCNN++ with sampling speed optimizations (Ramachandran et al., 2017). More samples are available in the supplementary.

Related Work

Likelihood-based models constitute a large family of deep generative models. One subclass of such methods, based on variational inference, allows for efficient approximate inference and sampling, but does not admit exact log likelihood computation (Kingma & Welling, 2013; Rezende et al., 2014; Kingma et al., 2016). Another subclass, which we called exact likelihood models in this work, does admit exact log likelihood computation. These exact likelihood models are typically specified as invertible transformations that are parameterized by neural networks (Deco & Brauer, 1995; Larochelle & Murray, 2011; Uria et al., 2013; Dinh et al., 2014; Germain et al., 2015; van den Oord et al., 2016b; Salimans et al., 2017; Chen et al., 2017).

There is prior work that aims to improve the sampling speed of deep autoregressive models. The Multiscale PixelCNN (Reed et al., 2017) modifies the PixelCNN to be non-fully-expressive by introducing conditional independence assumptions among pixels in a way that permits sampling in a logarithmic number of steps, rather than linear. Such a change in the autoregressive structure allows for faster sampling but also makes some statistical patterns impossible to capture, and hence reduces the capacity of the model for density estimation. WaveRNN (Kalchbrenner et al., 2018) improves sampling speed for autoregressive models for audio via sparsity and other engineering considerations, some of which may apply to flow models as well.

There is also recent work that aims to improve the expressiveness of coupling layers in flow models. Kingma & Dhariwal (2018) demonstrate improved density estimation using an invertible 1x1 convolution flow, and demonstrate that very large flow models can be trained to produce photorealistic faces. Huang et al. (2018) show how to design elementwise transformations which themselves are neural networks. Müller et al. (2018) introduce piecewise polynomial couplings that are similar in spirit to our mixture of logistics couplings and found them to be more expressive than affine couplings, but reported little performance gains in density estimation. We leave a detailed comparison between our coupling layer and these other types of coupling layers for future work.

Conclusion

We presented Flow++, a new flow-based generative model that begins to close the performance gap between flow models and autoregressive models. Our work considers specific instantiations of design principles for flow models – dequantization, flow design, and conditioning architecture design – and we hope these principles will help guide future research in flow models and likelihood-based models in general.

Acknowledgements

We thank Evan Lohn for discovering that our ImageNet models attain slightly different performance on different GPU hardware. This work was funded in part by ONR PECASE N000141612723, Huawei, Amazon AWS, and Google Cloud.

References