Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
Emily Denton, Soumith Chintala, Arthur Szlam, Rob Fergus
Introduction
Building a good generative model of natural images has been a fundamental problem within computer vision. However, images are complex and high dimensional, making them hard to model well, despite extensive efforts. Given the difficulties of modeling entire scene at high-resolution, most existing approaches instead generate image patches. In contrast, in this work, we propose an approach that is able to generate plausible looking scenes at and . To do this, we exploit the multi-scale structure of natural images, building a series of generative models, each of which captures image structure at a particular scale of a Laplacian pyramid . This strategy breaks the original problem into a sequence of more manageable stages. At each scale we train a convolutional network-based generative model using the Generative Adversarial Networks (GAN) approach of Goodfellow et al. . Samples are drawn in a coarse-to-fine fashion, commencing with a low-frequency residual image. The second stage samples the band-pass structure at the next level, conditioned on the sampled residual. Subsequent levels continue this process, always conditioning on the output from the previous scale, until the final level is reached. Thus drawing samples is an efficient and straightforward procedure: taking random vectors as input and running forward through a cascade of deep convolutional networks (convnets) to produce an image.
Deep learning approaches have proven highly effective at discriminative tasks in vision, such as object classification . However, the same level of success has not been obtained for generative tasks, despite numerous efforts . Against this background, our proposed approach makes a significant advance in that it is straightforward to train and sample from, with the resulting samples showing a surprising level of visual fidelity, indicating a better density model than prior methods.
Generative image models are well studied, falling into two main approaches: non-parametric and parametric. The former copy patches from training images to perform, for example, texture synthesis or super-resolution . More ambitiously, entire portions of an image can be in-painted, given a sufficiently large training dataset . Early parametric models addressed the easier problem of texture synthesis , with Portilla & Simoncelli making use of a steerable pyramid wavelet representation , similar to our use of a Laplacian pyramid. For image processing tasks, models based on marginal distributions of image gradients are effective , but are only designed for image restoration rather than being true density models (so cannot sample an actual image). Very large Gaussian mixture models and sparse coding models of image patches can also be used but suffer the same problem.
A wide variety of deep learning approaches involve generative parametric models. Restricted Boltzmann machines , Deep Boltzmann machines , Denoising auto-encoders all have a generative decoder that reconstructs the image from the latent representation. Variational auto-encoders provide probabilistic interpretation which facilitates sampling. However, for all these methods convincing samples have only been shown on simple datasets such as MNIST and NORB, possibly due to training complexities which limit their applicability to larger and more realistic images.
Several recent papers have proposed novel generative models. Dosovitskiy et al. showed how a convnet can draw chairs with different shapes and viewpoints. While our model also makes use of convnets, it is able to sample general scenes and objects. The DRAW model of Gregor et al. used an attentional mechanism with an RNN to generate images via a trajectory of patches, showing samples of MNIST and CIFAR10 images. Sohl-Dickstein et al. use a diffusion-based process for deep unsupervised learning and the resulting model is able to produce reasonable CIFAR10 samples. Theis and Bethge employ LSTMs to capture spatial dependencies and show convincing inpainting results of natural textures.
Our work builds on the GAN approach of Goodfellow et al. which works well for smaller images (e.g. MNIST) but cannot directly handle large ones, unlike our method. Most relevant to our approach is the preliminary work of Mirza and Osindero and Gauthier who both propose conditional versions of the GAN model. The former shows MNIST samples, while the latter focuses solely on frontal face images. Our approach also uses several forms of conditional GAN model but is much more ambitious in its scope.
Approach
The basic building block of our approach is the generative adversarial network (GAN) of Goodfellow et al. . After reviewing this, we introduce our LAPGAN model which integrates a conditional form of GAN model into the framework of a Laplacian pyramid.
The conditional generative adversarial net (CGAN) is an extension of the GAN where both networks and receive an additional vector of information as input. This might contain, say, information about the class of the training example . The loss function thus becomes
where is, for example, the prior distribution over classes. This model allows the output of the generative model to be controlled by the conditioning variable . Mirza and Osindero and Gauthier both explore this model with experiments on MNIST and faces, using as a class indicator. In our approach, will be another image, generated from another CGAN model.
2 Laplacian Pyramid
The Laplacian pyramid is a linear invertible image representation consisting of a set of band-pass images, spaced an octave apart, plus a low-frequency residual. Formally, let be a downsampling operation which blurs and decimates a image , so that is a new image of size . Also, let be an upsampling operator which smooths and expands to be twice the size, so is a new image of size . We first build a Gaussian pyramid , where and is repeated applicationsi.e. . of to . is the number of levels in the pyramid, selected so that the final level has very small spatial extent ( pixels).
The coefficients at each level of the Laplacian pyramid are constructed by taking the difference between adjacent levels in the Gaussian pyramid, upsampling the smaller one with so that the sizes are compatible:
Intuitively, each level captures image structure present at a particular scale. The final level of the Laplacian pyramid is not a difference image, but a low-frequency residual equal to the final Gaussian pyramid level, i.e. . Reconstruction from a Laplacian pyramid coefficients is performed using the backward recurrence:
which is started with and the reconstructed image being . In other words, starting at the coarsest level, we repeatedly upsample and add the difference image at the next finer level until we get back to the full resolution image.
3 Laplacian Generative Adversarial Networks (LAPGAN)
Our proposed approach combines the conditional GAN model with a Laplacian pyramid representation. The model is best explained by first considering the sampling procedure. Following training (explained below), we have a set of generative convnet models , each of which captures the distribution of coefficients for natural images at a different level of the Laplacian pyramid. Sampling an image is akin to the reconstruction procedure in Eqn. 4, except that the generative models are used to produce the ’s:
The generative models are trained using the CGAN approach at each level of the pyramid. Specifically, we construct a Laplacian pyramid from each training image . At each level we make a stochastic choice (with equal probability) to either (i) construct the coefficients either using the standard procedure from Eqn. 3, or (ii) generate them using :
Breaking the generation into successive refinements is the key idea in this work. Note that we give up any “global” notion of fidelity; we never make any attempt to train a network to discriminate between the output of a cascade and a real image and instead focus on making each step plausible. Furthermore, the independent training of each pyramid level has the advantage that it is far more difficult for the model to memorize training examples – a hazard when high capacity deep networks are used.
As described, our model is trained in an unsupervised manner. However, we also explore variants that utilize class labels. This is done by add a 1-hot vector , indicating class identity, as another conditioning variable for and .
Model Architecture & Training
We apply our approach to three datasets: (i) CIFAR10 – 3232 pixel color images of 10 different classes, 100k training samples with tight crops of objects; (ii) STL – 9696 pixel color images of 10 different classes, 100k training samples (we use the unlabeled portion of data); and (iii) LSUN – 10M images of 10 different natural scene types, downsampled to 6464 pixels.
For each dataset, we explored a variety of architectures for . We now detail the best performing models, selected using a combination of log-likelihood and visual appearance of the samples. Complete Torch specification files for all models are provided in supplementary material . For all models, the noise vector is drawn from a uniform distribution.
Initial scale: This operates at resolution, using densely connected nets for both & with 2 hidden layers and ReLU non-linearities. uses Dropout and has 600 units/layer vs 1200 for . is a 100-d vector.
Subsequent scales: For CIFAR10, we boost the training set size by taking four crops from the original images. Thus the two subsequent levels of the pyramid are and . For STL, we have 4 levels going from . For both datasets, & are convnets with 3 and 2 layers, respectively (see ). The noise input to is presented as a 4th “color plane” to low-pass , hence its dimensionality varies with the pyramid level. For CIFAR10, we also explore a class conditional version of the model, where a vector encodes the label. This is integrated into & by passing it through a linear layer whose output is reshaped into a single plane feature map which is then concatenated with the 1st layer maps. The loss in Eqn. 2 is trained using SGD with an initial learning rate of 0.02, decreased by a factor of at each epoch. Momentum starts at 0.5, increasing by 0.0008 at epoch up to a maximum of 0.8. During training, we monitor log-likelihood using a Parzen-window estimator and retain the best performing model. Training time depends on the models size and pyramid level, with smaller models taking hours to train and larger models taking several days.
2 LSUN
The larger size of this dataset allows us to train a separate LAPGAN model for each the 10 different scene classes. During evaluation, so that we may understand the variation captured by our models, we commence the sampling process with validation set imagesThese were not used in any way during training., downsampled to resolution.
The four subsequent scales use a common architecture for & at each level. is a 5-layer convnet with feature maps and a linear output layer. filters, ReLUs, batch normalization and Dropout are used at each hidden layer. has 3 hidden layers with maps plus a sigmoid output. See for full details. Note that and are substantially larger than those used for CIFAR10 and STL, as afforded by the larger training set.
Experiments
We evaluate our approach using 3 different methods: (i) computation of log-likelihood on a held out image set; (ii) drawing sample images from the model and (iii) a human subject experiment that compares (a) our samples, (b) those of baseline methods and (c) real images.
A traditional method for evaluating generative models is to measure their log-likelihood on a held out set of images. But, like the original GAN method , our approach does not have a direct way of computing the probability of an image. Goodfellow et al. propose using a Gaussian Parzen window estimate to compute log-likelihoods. Despite showing poor performance in high dimensional spaces, this approach is the best one available for estimating likelihoods of models lacking an explicitly represented density function.
Our LAPGAN model allows for an alternative method of estimating log-likelihood that exploits the multi-scale structure of the model. This new approach uses a Gaussian Parzen window estimate to compute a probability at each scale of the Laplacian pyramid. We use this procedure, described in detail in Appendix A, to compute the log-likelihoods for CIFAR10 and STL images (both at resolution). The parameter (controlling the Parzen window size) was chosen using the validation set. We also compute the Parzen window based log-likelihood estimates of the standard GAN model, using 50k samples for both the CIFAR10 and STL estimates. Table 1 shows our model achieving a significantly higher log-likelihood on both datasets. Comparisons to further approaches, notably , are problematic due to different normalizations used on the data.
2 Model Samples
We show samples from models trained on CIFAR10, STL and LSUN datasets. Additional samples can be found in the supplementary material .
Fig. 3 shows samples from our models trained on CIFAR10. Samples from the class conditional LAPGAN are organized by class. Our reimplementation of the standard GAN model produces slightly sharper images than those shown in the original paper. We attribute this improvement to the introduction of data augmentation. The LAPGAN samples improve upon the standard GAN samples. They appear more object-like and have more clearly defined edges. Conditioning on a class label improves the generations as evidenced by the clear object structure in the conditional LAPGAN samples. The quality of these samples compares favorably with those from the DRAW model of Gregor et al. and also Sohl-Dickstein et al. . The rightmost column of each image shows the nearest training example to the neighboring sample (in L2 pixel-space). This demonstrates that our model is not simply copying the input examples.
Fig. 4 shows samples from our LAPGAN model trained on STL. Here, we lose clear object shape but the samples remain sharp. Fig. 4 shows the generation chain for random STL samples.
Fig. 5 shows samples from LAPGAN models trained on three LSUN categories (tower, bedroom, church front). The validation image used to start the generation process is shown in the first column, along with 10 different samples, which illustrate the inherent variation captured by the model. Collectively, these show the models capturing long-range structure within the scenes, being able to recompose scene elements into credible looking images. To the best of our knowledge, no other generative model has been able to produce samples of this complexity. The substantial gain in quality over the CIFAR10 and STL samples is likely due to the much larger training LSUN training set which allowed us to train bigger and deeper models.
3 Human Evaluation of Samples
To obtain a quantitative measure of quality of our samples, we asked 15 volunteers to participate in an experiment to see if they could distinguish our samples from real images. The subjects were presented with the user interface shown in Fig. 6(right) and shown at random four different types of image: samples drawn from three different GAN models trained on CIFAR10 ((i) LAPGAN, (ii) class conditional LAPGAN and (iii) standard GAN ) and also real CIFAR10 images. After being presented with the image, the subject clicked the appropriate button to indicate if they believed the image was real or generated. Since accuracy is a function of viewing time, we also randomly pick the presentation time from one of 11 durations ranging from 50ms to 2000ms, after which a gray mask image is displayed. Before the experiment commenced, they were shown examples of real images from CIFAR10. After collecting 10k samples from the volunteers, we plot in Fig. 6 the fraction of images believed to be real for the four different data sources, as a function of presentation time. The curves show our models produce samples that are far more realistic than those from standard GAN .
Discussion
By modifying the approach in to better respect the structure of images, we have proposed a conceptually simple generative model that is able to produce high-quality sample images that are both qualitatively and quantitatively better than other deep generative modeling approaches. A key point in our work is giving up any “global” notion of fidelity, and instead breaking the generation into plausible successive refinements. We note that many other signal modalities have a multiscale structure that may benefit from a similar approach.
Appendix A
in a moment we will carefully define the functions . For now, suppose that , , and for each fixed , . Then we can check that has unit integral:
Now we define the with Parzen window approximations to the densities of each of the scales. For , we take a set of training samples , and construct the density function . We fix to define .For pyramids with more levels, we continue in the same way for each of the finer scales. Note we always use the true low pass at each scale, and measure the true high pass against the high pass samples generated from the model. Thus for a pyramid with levels, the final log likelihood will be: .