Guided Image Generation with Conditional Invertible Neural Networks

Lynton Ardizzone, Carsten Lüth, Jakob Kruse, Carsten Rother, Ullrich Köthe

Introduction

Generative adversarial networks (GANs) produce ever larger and more realistic samples . Hence they have become the primary choice for a majority of image generation tasks. As such, their conditional variants (cGANs) would appear to be the natural tool for conditional image generation as well, and they have successfully been applied in many scenarios . Despite numerous improvements, significant expertise and computational resources are required to find a training configuration for large GANs that is stable, and produces diverse images. A lack in diversity is especially common when the condition itself is an image, and special precautions have to be taken to avoid mode collapse.

Conditional variational autoencoders (cVAEs) do not suffer from the same problems. Training is generally stable, and since every data point is assigned a region in latent space, sampling yields the full variety of data seen during training. However cVAEs come with drawbacks of their own: The assumption of a Gaussian posterior on the decoder side implies an L2 reconstruction loss, which is known to cause blurriness. In addition, the partition of the latent space into diagonal Gaussians leads to either mode-mixing issues or regions of poor sample quality . There has also been some success in combining aspects of both approaches for certain tasks, such as .

We propose a third approach, by extending Invertible Neural Networks (INNs, ) for the task of conditional image generation, by adding conditioning inputs to their core building blocks. INNs are neural networks which are by construction bijective, efficiently invertible, and have a tractable Jacobian determinant. They represent transport maps between the input distribution p(x)p(\mathbf{x}) and a prescribed, easy-to-sample-from latent distribution p(z)p(\mathbf{z}). During training, the likelihood of training samples from p(x)p(\mathbf{x}) is maximized in latent space, while at inference time, z\mathbf{z}-samples can trivially be transformed back to the data domain. Previously, INNs have been used successfully for unconditional image generation, e.g. by and .

Unconditional INN training is related to that of VAEs, but it compensates for some key disadvantages: Firstly, since reconstructions are perfect by design, no reconstruction loss is needed, and generated images do not become blurry. Secondly, each x\mathbf{x} maps to exactly one z\mathbf{z} in latent space, and there is no need for posteriors p(z ∣ x)p(\mathbf{z}\,|\,\mathbf{x}). This avoids the VAE problem of disjoint or overlapping regions in latent space. In terms of training stability and sample diversity, INNs show the same strengths as autoencoder architectures, but with superior image quality. We find that these positive aspects apply to conditional INNs (cINNs) as well.

One limitation of INNs is that their design restricts the use of some standard components of neural networks, such as pooling and batch normalization layers. Our conditional architecture alleviates this problem, as the conditional inputs can be preprocessed by a conditioning network with a standard feed-forward architecture, which can be learned jointly with the cINN to greatly improve its generative capabilities. We demonstrate the qualities of cINNs for conditional image generation, and uncover emergent properties of the latent space, for the tasks of conditional MNIST generation and diverse colorization of ImageNet.

Our work makes the following contributions:

We propose a new architecture called conditional invertible neural network (cINN), which combines an INN with an unconstrained feed-forward network for conditioning. It generates diverse images with high realism and thus overcomes limitations of existing approaches.

We demonstrate a stable, maximum likelihood-based training procedure for jointly optimizing the parameters of the INN and the conditioning network.

We take advantage of our bidirectional cINN architecture to explore and manipulate emergent properties of the latent space. We illustrate this for MNIST digit generation and image colorization.

Related work

Conditional Generative Modeling. Modern generative models learn to transform noise (usually sampled from multivariate Gaussians) into desired target distributions. Methods differ by the model-family these transformations are picked from and by the losses determining optimal solutions.

Conditional generative adversarial networks (cGANs) train a pair of neural networks: a generator transforms a pair of conditioning and noise vectors to images, and a discriminator penalizes unrealistic looking images. The conditioning information is either concatenated to the noise , or fed into the network via conditional batch-norm layers . Ensuring diversity of the generated images (for fixed conditioning) appears to be challenging in this approach. Recent BigGANs successfully address this problem by using very large networks and batch sizes, but require parallel training on up to 512 TPUs. PacGANs employ augmented discriminators, which evaluate entire batches of real or generated images together rather than one image at a time. CausalGANs train two additional discriminator networks, called “labeler” and “anti-labeler”, with the latter explicitly penalizing the lack of diversity. Pix2pix addresses the important special case when the target is conditioned on an image in a different modality, e.g. to generate satellite images from maps. In addition to the discriminator loss, it minimizes the L1 distance between generated and ground-truth targets using a paired training set, which contains corresponding images from both modalities. This leads to impressive image quality, but lack of diversity seems to be an especially hard problem in this case. In contrast, our method does not need explicit precautions to promote diversity.

Bidirectional architectures augment generator networks with complementary encoder networks that learn the generator’s inverse and enable reconstruction losses, which exploit cycle consistency requirements. Conditional variational autoencoders (cVAEs) replace all distributions in a standard VAE by the appropriate conditional distributions, and are trained to minimize the evidence lower bound (ELBO loss). Since variational distributions are typically Gaussian, the reconstruction penalty is equivalent to squared loss, resulting in rather blurry generated images. This is avoided by AGE networks and CycleGANs , which combine standard cGAN discriminators with L1 reconstruction loss in the data domain, and bidirectional conditional GANs , which extend the GAN discriminator to act on the distributions in data and latent space jointly. SPADE , building upon pix2pix and pix2pixHD , augments cGANs with additional VAE encoders to shape the latent space such that diversity is ensured.

Instead of enforcing bijectivity through cycle losses, invertible neural networks are bidirectional by design, since encoder and generator are realized by forward and backward processing within a single bijective model. We focus on architectures whose forward and backward pass require the same computational effort. The coupling layer designs pioneered by NICE and RealNVP emerged as very powerful and flexible model families under this restriction. Using additive coupling layers, i-RevNets demonstrated that the lack of information reduction from data space to latent space does not cause overfitting. The Glow architecture combines affine coupling layers with invertible 1x1 convolutions and achieves impressive attribute manipulations (e.g. age, hair color) in generated faces images. This approach was recently generalized to video .

Thanks to tractable Jacobian determinants, the coupling layer architecture enables maximum likelihood training , but experimental comparisons with other training methods are inconclusive so far. For instance, found minimization of an adversarial loss to be superior to maximum likelihood training in RealNVPs, trained i-RevNets in the same manner as adversarial auto-encoders, i.e. with a discriminator acting in latent rather than data space, and Flow-GANs performed best using bidirectional training, a combination of maximum likelihood and adversarial loss. On the other hand, maximum likelihood training worked well within Glow , and i-ResNets could even be trained with approximated Jacobian determinants. In this work we reinforce the view that high-quality generative models can be trained by maximum likelihood loss alone. To the best of our knowledge, we are the first to apply the coupling layer design for conditional generative models, with the exception of , who use it to compute posteriors for (relatively small) inverse problems, but do not consider image generation.

Colorization. State-of-the-art regression models for colorization produce visually near-perfect images , but do not account for the ambiguity inherent in this inverse problem. To address this, models would ideally define a conditional distribution of plausible color images for a given grayscale input, instead of just returning a single “best” solution.

Popular existing approaches for diverse colorization predict per-pixel color histograms from a U-Net or from hypercolumns of an adapted VGG network . However, sampling from these local histograms independently can not lead to a spatially consistent colorization, requiring additional heuristic post-processing steps to avoid artefacts.

In terms of generative models, both VAEs and cGANs have been proposed for the task. However, their solutions do not reach the quality of the regression-based models, and cGANs in particular often lack diversity. To compensate, modifications and extensions to generative approaches have been developed, such as auto-regressive models and CRFs . However, these methods are computationally very expensive and often unable to scale to realistic image sizes.

Conceptually closest to our proposed method is the work of , where an encoder network maps color information to a latent space and a generator network learns the inverse transform, both conditioned on the grayscale image. Their experiments however are limited to a data set with only cars, and just three latent dimensions, leading to global, but no local diversity.

In contrast to the above, our flow-based cINN generates diverse colorizations in one standard feed-forward pass. It models the distribution of all pixels jointly, and allows for meaningful latent space manipulations.

Method

Our method is an extension of the affine coupling block architecture established in . There, each network block splits its input u\mathbf{u} into two parts [u1,u2][\mathbf{u}_{1},\mathbf{u}_{2}] and applies affine transformations between them that have strictly upper or lower triangular Jacobians:

The outputs [v1,v2][\mathbf{v}_{1},\mathbf{v}_{2}] are concatenated again and passed to the next coupling block. The internal functions sjs_{j} and tjt_{j} can be represented by arbitrary neural networks, and are only ever evaluated in the forward direction, even when the coupling block is inverted:

As shown in , the logarithm of the Jacobian determinant for such a coupling block is simply the sum of s1s_{1} and s2s_{2} over image dimensions.

We adapt the design of Eqs. 1 and 2 to produce a conditional version of the coupling block. Because the subnetworks sjs_{j} and tjt_{j} are never inverted, we can concatenate conditioning data c\mathbf{c} to their inputs without losing the invertibility, replacing s1(u2)s_{1}(\mathbf{u}_{2}) with s1(u2,c)s_{1}(\mathbf{u}_{2},\mathbf{c}) etc. Our conditional coupling block design is illustrated in Fig. 2.

In general, we will refer to a cINN with network parameters θ\theta as f(x;c,θ)f(\mathbf{x};\mathbf{c},\theta), and the inverse as g(z;c,θ)g(\mathbf{z};\mathbf{c},\theta). For any fixed condition c\mathbf{c}, the invertibility is given as

2 Maximum likelihood training of cINNs

By prescribing a probability distribution pZ(z)p_{Z}(\mathbf{z}) on latent space ZZ, the model ff assigns any input x\mathbf{x} a probability, dependent on both the network parameters θ\theta and the conditioning c\mathbf{c}, through the change-of-variables formula:

Here, we use the Jacobian matrix ∂f/∂x{\partial f}/{\partial\mathbf{x}}. We will denote the determinant of the Jacobian, evaluated at some training sample xi\mathbf{x}_{i}, as J_{i}\equiv\text{det}\big{(}{\partial f}/{\partial\mathbf{x}}|_{\mathbf{x}_{i}}\big{)}. Bayes’ theorem gives us the posterior over model parameters as p(θ;x,c)∝pX(x;c,θ)⋅pθ(θ)p(\theta;\mathbf{x},\mathbf{c})\propto p_{X}(\mathbf{x};\mathbf{c},\theta)\cdot p_{\theta}(\theta). Our goal is to find network parameters that maximize its logarithm, i.e. we minimize the loss

which is the same as in classical Bayesian model fitting.

Inserting Eq. 4 with a standard normal distribution for pZ(z)p_{Z}(\mathbf{z}), as well as a Gaussian prior on the weights θ\theta with 1/2σθ2≡τ1/2\sigma_{\theta}^{2}\equiv\tau, we obtain

The latter term represents L2 weight regularization, while the former is the maximum likelihood loss.

Training a network with this loss yields an estimate of the maximum likelihood network parameters θ^ML\hat{\theta}_{\text{ML}}. From there, we can perform conditional generation for a fixed c\mathbf{c} by sampling z\mathbf{z} and using the inverted network gg: xgen=g(z;c,θ^ML)\mathbf{x}_{\text{gen}}=g(\mathbf{z};\mathbf{c},\hat{\theta}_{\text{ML}}), with z∼pZ(z)\mathbf{z}\sim p_{Z}(\mathbf{z}).

Training with the maximum likelihood method makes it virtually impossible for mode collapse to occur: If any mode in the training set has low probability under the current guess pX(x;c,θ)p_{X}(\mathbf{x};\mathbf{c},\theta), the corresponding latent vectors will lie far outside the normal distribution pZp_{Z} and receive big loss from the first L2-term in Eq. 6. In contrast, the discriminator of a GAN only supplies a weak signal, proportional to the mode’s relative frequency in the training data, so that the generator is not penalized much for ignoring a mode completely.

3 Conditioning network

In complex settings, we expect that higher-level features of c\mathbf{c} need to be extracted for the conditioning to be effective, e.g. global semantic information from an image as in Section 4.2. In such cases, feeding the condition c\mathbf{c} directly into the cINN would place an unreasonable burden on the ss and tt networks, as higher-level features would have to be re-learned in each coupling block.

4 Important details

For cINNs to match the performance of well-established architectures for conditional generation, we introduce a number of minor modifications and adjustments to the architecture and training procedure. With these adaptions, our training setup is very stable and converges every time. Ablation results are presented in Sec. 4.4.

Noise as data augmentation. We add a small amount of noise to the inputs x\mathbf{x} as part of the standard data augmentation. This helps to smooth out quantization artifacts in the input, and prevents sparse gradients when large parts of the image are completely flat (as e.g. in MNIST).

Soft clamping of scale coefficients. We apply an additional nonlinear function to the scale coefficients ss, of the form

which yields sclamp≈ss_{\text{clamp}}\approx s for ∣s∣≪α|s|\ll\alpha and sclamp≈±αs_{\text{clamp}}\approx\pm\alpha for ∣s∣≫α|s|\gg\alpha. This prevents any instabilities stemming from exploding magnitude of the exponential exp⁡(sclamp)\exp(s_{\text{clamp}}). We find α=1.9\alpha=1.9 to be a good value for most architectures.

Initialization. Heuristically, we find that Xavier initialization leads to stable training from the start. We experienced training instability when initial parameter values were too high. Similar to , we also initialize the last convolution in all ss and tt subnetworks to zero, so training starts from an identity transform.

Soft channel permutations. We use random orthogonal matrices to mix the information between the channels. This allows for more interaction between the two information streams u1,u2\mathbf{u}_{1},\mathbf{u}_{2} in the coupling blocks. A similar technique was used in , but our matrices stay fixed throughout training and are guaranteed to be cheaply invertible.

Haar wavelet downsampling. All prior INN architectures use checkerboard patterns for reshaping to lower spatial resolutions. We find it helpful to instead perform downsampling with Haar wavelets , which essentially decompose images into an average pooling channel as well as vertical, horizontal and diagonal derivatives, see Fig. 3. The three derivative channels contain high resolution information which we can split off early, transforming only the remaining information further in later stages of the cINN. This also contributes to mixing the variables between layers, complementing the soft permutations. Similarly, uses a discrete cosine transform as a final transformation in their INN, to replace global average pooling.

Experiments

We present results and explore the latent space of our models for two conditional image generation tasks: MNIST digit generation and image colorization.

As a first experiment, we perform simple class-conditional generation of MNIST digits. We construct a cINN of 24 coupling blocks using fully connected subnetworks ss and tt, which receive the conditioning directly as a one-hot vector (Fig. 5). No conditioning network hh is used. For data augmentation we only add a small amount of noise to the images (σ=0.02\sigma=0.02), as described in Section 3.4.

Samples generated by the model are shown in Fig. 6. We find that the cINN learns latent representations that are shared across conditions c\mathbf{c}. Keeping the latent vector z\mathbf{z} fixed while varying c\mathbf{c} produces different digits in the same style. This property, in conjunction with our network’s invertibility, can directly be used for style transfer, as demonstrated in Fig. 7. This outcome is not obvious – the trained cINN could also decompose into 10 essentially separate subnetworks, one for each condition. In this case, the latent space of each class would be structured differently, and inter-class transfer of latent vectors would be meaningless. The structure of the latent space is further illustrated in Fig. 4, where we identify three latent axes with interpretable meanings. Note that while the latent space is learned without supervision, we found the axes in a semi-automatic fashion: We perform PCA on the latent vectors of the test set, without the noise augmentation, and manually identify meaningful directions in the subspace of the first four principal components.

2 Diverse ImageNet colorization

For a more challenging task, we turn to colorization of natural images. The common approach for this task is to represent images in LabLab color space and generate color channels a,b\mathbf{a},\mathbf{b} by a model conditioned on the luminance channel L\mathbf{L}.

We train on the ImageNet dataset , again adding low noise to the a,b\mathbf{a},\mathbf{b} channels (σ=0.05\sigma=0.05). As the color channels do not require as much resolution as the luminance channel, we condition on 256×256256\times 256 pixel grayscale images, but generate 64×6464\times 64 pixel color information. This is in accordance with the majority of existing colorization methods.

As with most generative INN architectures, we do not keep the resolution and channels fixed throughout the network, for the sake of computational cost. Instead, we use 4 resolution stages, as illustrated in Fig. 8. At each stage, the data is reshaped to a lower resolution and more channels, after which a fraction of the channels are split off as one part of the latent code. As the high resolution stages have a smaller receptive field and less expressive power, the corresponding parts of the latent vector encode local structures and noise. More global information is passed on to the lower resolution sections of the cINN.

We initially train the cINN and the hkh_{k}, keeping the parameters of the conditioning network hh fixed, for 30 00030\,000 iterations. After this, we train both jointly until convergence, for 3 days on 3 Nvidia GTX1080 GPUs. The Adam optimizer is essential for fast convergence, and we lower the learning rate when the maximum likelihood loss levels off.

At inference time, we use joint bilateral upsampling to match the resolution of the generated color channels a^,b^\hat{\mathbf{a}},\hat{\mathbf{b}} to that of the luminance channel L\mathbf{L}. This produces visually slightly more pleasing edges than bicubic upsampling, but has little to no impact on the results. It was not used in the quantitative results table, to ensure an unbiased comparison.

The cINN compares favourably to existing methods, as shown in Table 1, and has the best diversity and best-of-8 accuracy of the compared methods. The cGAN apparently ignores the latent code, and relies only on the condition. As a result, we do not measure any significant diversity, in line with results from .

In terms of FID score, the cGAN performs best, although its results do not appear more realistic to the human eye, cf. Fig. 13. This may be due to the fact that FID is sensitive to outliers, which are unavoidable for a truly diverse method (see Fig. 12), or because the discriminator loss implicitly optimizes for the similarity of deep CNN activations. The VGG classification accuracy of generative methods is decreased compared to CNN, because occasional outliers may lead to misclassification. Latent space interpolations and color transfer are shown in Figs. 14 and 15.

3 Diverse bedrooms colorization

To provide a simpler model for more in-depth experiments and ablations, we additionally train a cINN for colorization on the LSUN bedrooms dataset . We use a smaller model than for ImageNet, and train the conditioning network jointly from scratch, without pretraining. Both the conditioning input, as well as the generated color channels have a resolution of 64×6464\times 64 pixels. The entire model trains in under 4 hours on a single GTX 1080Ti GPU.

To our knowledge, the only diversity-enforcing cGAN architecture previously used for colorization is the colorGAN , which is also trained exclusively on the bedrooms dataset. Training the colorGAN for comparison, we find it requires over 24 hours to converge stably, after multiple restarts. The results are generally worse than those of the cINN, as shown in Fig. 9. While the resulting pixel-wise color variance is slightly higher for the colorGAN, it is not clear whether this captures the true variance, or whether it is due to unrealistically colorful outputs, such as in the second row in Fig. 9.

4 Ablation of training improvements

To demonstrate the improved stability and training speed through the improvements from Sec. 3.4, we perform ablations, see Fig. 10. The ablations for colorization were performed for the LSUN bedrooms task, due to training speed.

We find that for stable training at Adam learning rates of 10−310^{-3}, the clamping and Haar wavelet downsampling are strictly necessary. Without these, the network has to be trained with much lower learning rates and more careful and specialized initialization, as used e.g. in . Beyond this, the noise augmentation and permutations lead to the largest improvement in final result. The effect of the noise is more pronounced for MNIST, as large parts of the image are completely black otherwise. For natural images, dequantization of the data is likely to be the main advantage of the added noise. The initialization only improves the final result by a small margin, but also converges noticeably faster.

Conclusion and Outlook

We have proposed a conditional invertible neural network architecture which enables guided generation of diverse images with high realism. For image colorization, we believe that even better results can be achieved when employing latest tricks from large-scale GAN frameworks. Especially the non-invertible nature of the conditioning network make cINNs a suitable method for other computer vison tasks such as diverse semantic segmentation.

Acknowledgments

This work is supported by Deutsche Forschungsgemeinschaft (DFG) under Germany’s Excellence Strategy EXC-2181/1 - 390900948 (the Heidelberg STRUCTURES Excellence Cluster). LA received funding by the Federal Ministry of Education and Research of Germany, project ‘High Performance Deep Learning Framework’ (No 01IH17002). JK, CR and UK received financial support from the European Research Council (ERC) under the European Unions Horizon 2020 research and innovation program (grant agreement No 647769). JK received funding by Informatics for Life funded by the Klaus Tschira Foundation. Computations were performed on an HPC Cluster at the Center for Information Services and High Performance Computing (ZIH) at TU Dresden.

References