Adversarial Latent Autoencoders

Stanislav Pidhorskyi, Donald Adjeroh, Gianfranco Doretto

Introduction

Generative Adversarial Networks (GAN) have emerged as one of the dominant unsupervised approaches for computer vision and beyond. Their strength relates to their remarkable ability to represent complex probability distributions, like the face manifold , or the bedroom images manifold , which they do by learning a generator map from a known distribution onto the data space. Just as important are the approaches that aim at learning an encoder map from the data to a latent space. They allow learning suitable representations of the data for the task at hand, either in a supervised , or unsupervised manner.

Autoencoder (AE) networks are unsupervised approaches aiming at combining the “generative” as well as the “representational” properties by learning simultaneously an encoder-generator map. General issues subject of investigation in AE structures are whether they can: (a) have the same generative power as GANs; and, (b) learn disentangled representations . Several works have addressed (a) . An important testbed for success has been the ability for an AE to generate face images as rich and sharp as those produced by a GAN . Progress has been made but victory has not been declared. A sizable amount of work has addressed also (b) , but not jointly with (a).

We introduce an AE architecture that is general, and has generative power comparable to GANs while learning a less entangled representation. We observed that every AE approach makes the same assumption: the latent space should have a probability distribution that is fixed a priori and the autoencoder should match it. On the other hand, it has been shown in , the state-of-the-art for synthetic image generation with GANs, that an intermediate latent space, far enough from the imposed input space, tends to have improved disentanglement properties.

We designed two ALAEs, one with a multilayer perceptron (MLP) as encoder with a symmetric generator, and another with the generator derived from a StyleGAN , which we call StyleALAE. For this one, we designed a companion encoder and a progressively growing architecture. We verified qualitatively and quantitatively that both architectures learn a latent space that is more disentangled than the imposed one. In addition, we show qualitative and quantitative results about face and bedroom image generation that are comparable with StyleGAN at the highest resolution of 1024×10241024\times 1024. Since StyleALAE learns also an encoder network, we are able to show at the highest resolution, face reconstructions as well as several image manipulations based on real images rather then generated.

Related Work

Our approach builds directly on the vanilla GAN architecture . Since then, a lot of progress has been made in the area of synthetic image generation. LAPGAN and StackGAN train a stack of GANs organized in a multi-resolution pyramid to generate high-resolution images. HDGAN improves by incorporating hierarchically-nested adversarial objectives inside the network hierarchy. In they use a multi-scale generator and discriminator architecture to synthesize high-resolution images with a GAN conditioned on semantic label maps, while in BigGAN they improve the synthesis by applying better regularization techniques. In PGGAN it is shown how high-resolution images can be synthesized by progressively growing the generator and the discriminator of a GAN. The same principle was used in StyleGAN , the current state-of-the-art for face image generation, which we adapt it here for our StyleALAE architecture. Other recent work on GANs has focussed on improving the stability and robustness of the training . New loss functions have been introduced , along with gradient regularization methods , weight normalization techniques , and learning rate equalization . Our framework is amenable to these improvements, as we explain in later sections.

Variational AE architectures have not only been appreciated for their theoretical foundation, but also for their stability during training, and the ability to provide insightful representations. Indeed, they stimulated research in the area of disentanglement , allowing learning representations with controlled degree of disentanglement between factors of variation in , and subsequent improvements in , leading to more elaborate metrics for disentanglement quantification , which we also use to analyze the properties of our approach. VAEs have also been extended to learn a latent prior different than a normal distribution, thus achieving significantly better models .

A lot of progress has been made towards combining the benefits of GANs and VAEs. AAE has been the precursor of those approaches, followed by VAE/GAN with a more direct approach. BiGAN and ALI provide an elegant framework fully adversarial, whereas VEEGAN and AGE pioneered the use of the latent space for autoencoding and advocated the reduction of the architecture complexity. PIONEER and IntroVAE followed this line, with the latter providing the best generation results in this category. Section 4.1 describes how the proposed approach compares with those listed here.

Finally, we quickly mention other approaches that have shown promising results with representing image data distributions. Those include autoregressive and flow-based methods . The former forego the use of a latent representation, but the latter does not.

Preliminaries

Following the more general formulation introduced in , the GAN learning problem entails finding the minimax with respect to the pair (G,D)(\mathtt{G},\mathtt{D}) (i.e., the Nash equilibrium), of the value function defined as

Adversarial Latent Autoencoders

We introduce a novel autoencoder architecture by modifying the original GAN paradigm. We begin by decomposing the generator G\mathtt{G} and the discriminator D\mathtt{D} in two networks: FF, GG, and EE, DD, respectively. This means that

see Figure 1. In addition, we assume that the latent spaces at the interface between FF and GG, and between EE and DD are the same, and we indicate them as W\mathcal{W}. In the most general case we assume that FF is a deterministic map, whereas we allow EE and GG to be stochastic. In particular, we assume that GG might optionally depend on an independent noisy input η\eta, with a known fixed distribution pη(η)p_{\eta}(\eta). We indicate with G(w,η)G(w,\eta) this more general stochastic generator.

Under the above conditions we now consider the distributions at the output of every network. The network FF simply maps p(z)p(z) onto qF(w)q_{F}(w). At the output of GG the distribution can be written as

where qG(x∣w,η)q_{G}(x|w,\eta) is the conditional distribution representing GG. Similarly, for the output of EE the distribution becomes

where qE(w∣x)q_{E}(w|x) is the conditional distribution representing EE. In (4) if we replace q(x)q(x) with pD(x)p_{\mathcal{D}}(x) we obtain the distribution qE,D(w)q_{E,\mathcal{D}}(w), which describes the output of EE when the real data distribution is its input.

Since optimizing (1) leads toward the synthetic distribution matching the real one, i.e., q(x)=pD(x)q(x)=p_{\mathcal{D}}(x), it is obvious from (4) that doing so also leads toward having qE(w)=qE,D(w)q_{E}(w)=q_{E,\mathcal{D}}(w). In addition to that, we propose to ensure that the distribution of the output of EE be the same as the distribution at the input of GG. This means that we set up an additional goal, which requires that

In this way we could interpret the pair of networks (G,E)(G,E) as a generator-encoder network that autoencodes the latent space W\mathcal{W}.

If we indicate with Δ(p∥q)\Delta(p\|q) a measure of discrepancy between two distributions pp and qq, we propose to achieve the goal (5) via regularizing the GAN loss (1) by alternating the following two optimizations

where the left and right arguments of Δ\Delta indicate the distributions generated by the networks mapping p(z)p(z), which correspond to qF(w)q_{F}(w) and qE(w)q_{E}(w), respectively. We refer to a network optimized according to (6) (7) as an Adversarial Latent Autoencoder (ALAE). The building blocks of an ALAE architecture are depicted in Figure 1.

Data distribution. In architectures composed by an encoder network and a generator network, the task of the encoder is to map input data onto a space characterized by a latent distribution, whereas the generator is tasked to map latent codes onto a space described by a data distribution. Different strategies are used to learn the data distribution. For instance, some approaches impose a similarity criterion on the output of the generator , or even learn a similarity metric . Other techniques instead, set up an adversarial game to ensure the generator output matches the training data distribution . This latter approach is what we use for ALAE.

Latent distribution. For the latent space instead, the common practice is to set a desired target latent distribution, and then the encoder is trained to match it either by minimizing a divergence type of similarity , or by setting up an adversarial game . Here is where ALAE takes a fundamentally different approach. Indeed, we do not impose the latent distribution, i.e., qE(w)q_{E}(w), to match a target distribution. The only condition we set, is given by (5). In other words, we do not want FF to be the identity map, and are very much interested in letting the learning process decide what FF should be.

Reciprocity. Another aspect of autoecoders is whether and how they achieve reciprocity. This property relates to the ability of the architecture to reconstruct a data sample xx from its code ww, and viceversa. Clearly, this requires that x=G(E(x))x=G(E(x)), or equivalently that w=E(G(w))w=E(G(w)). In the first case, the network must contain a reconstruction term that operates in the data space. In the latter one, the term operates in the latent space. While most approaches follow the first strategy , there are some that implement the second , including ALAE. Indeed, this can be achieved by choosing the divergence in (7) to be the expected coding reconstruction error, as follows

StyleALAE

We use ALAE to build an autoencoder that uses a StyleGAN based generator. For this we make our latent space W\mathcal{W} play the same role as the intermediate latent space in . Therefore, our GG network becomes the part of StyleGAN depicted on the right side of Figure 2. The left side is a novel architecture that we designed to be the encoder EE.

Since at every layer, GG is driven by a style input, we design EE symmetrically, so that from a corresponding layer we extract style information. We do so by inserting Instance Normalization (IN) layers , which provide instance averages and standard deviations for every channel. Specifically, if yiEy^{E}_{i} is the output of the ii-th layer of EE, the IN module extracts the statistics μ(yiE)\mu(y^{E}_{i}) and σ(yiE)\sigma(y^{E}_{i}) representing the style at that level. The IN module also provides as output the normalized version of the input, which continues down the pipeline with no more style information from that level. Given the information flow between EE and GG, the architecture is effectively mimicking a multiscale style transfer from EE to GG, with the difference that there is not an extra input image that provides the content .

The set of styles that are inputs to the Adaptive Instance Normalization (AdaIN) layers in GG are related linearly to the latent variable ww. Therefore, we propose to combine the styles output by the encoder, and to map them onto the latent space, via the following multilinear map

where the CiC_{i}’s are learnable parameters, and NN is the number of layers.

Similarly to we use progressive growing. We start from low-resolution images (4×44\times 4 pixels) and progressively increase the resolution by smoothly blending in new blocks to EE and GG. For the FF and DD networks we implement them using MLPs. The Z\mathcal{Z} and W\mathcal{W} spaces, and all layers of FF and DD have the same dimensionality in all our experiments. Moreover, for StyleALAE we follow , and chose FF to have 8 layers, and we set DD to have 3 layers.

Implementation

Adversarial losses and regularization. We use a non-saturating loss , which in (1) we introduce by setting f(⋅)f(\cdot) to be a SoftPlus function . This is a smooth version of the rectifier activation function, defined as f(t)=softplus⁡(t)=log⁡(1+exp⁡(t))f(t)=\operatorname{softplus}(t)=\log(1+\exp(t)). In addition, we use gradient regularization techniques . We utilize R1R_{1} , a zero-centered gradient penalty term which acts only on real data, and is defined as γ2E⁡pD(x)[∥∇D∘E(x)∥2]\frac{\gamma}{2}\operatorname{E}_{p_{\mathcal{D}}(x)}\left[\|\nabla D\circ E(x)\|^{2}\right], where the gradient is taken with respect to the parameters θE\theta_{E} and θD\theta_{D} of the networks EE and DD, respectively.

Training. In order to optimizate (6) (7) we use alternating updates. One iteration is composed of three updating steps: two for (6) and one for (7). Step I updates the discriminator (i.e., networks EE and DD). Step II updates the generator (i.e., networks FF and GG). Step III updates the latent space autoencoder (i.e., networks GG and EE). The procedural details are summarized in Algorithm 1. For updating the weights we use the Adam optimizer with β1=0.0\beta_{1}=0.0 and β2=0.99\beta_{2}=0.99, coupled with the learning rate equalization technique described below. For non-growing architectures (i.e., MLPs) we use a learning rate of 0.0020.002, and batch size of 128. For growing architectures (i.e., StyleALAE) learning rate and batch size depend on the resolution.

Experiments

Code and uncompressed images are available at https://github.com/podgorskiy/ALAE.

We train ALAE with MNIST , and then use the feature representation for classification, reconstruction, and analyzing disentanglement. We use the permutation-invariant setting, where each 28×2828\times 28 MNIST image is treated as a 784784D vector without spatial structure, which requires to use a MLP instead of a CNN. We follow and use a three layer MLP with a latent space size of 5050D. Both networks, EE and GG have two hidden layers with 1024 units each. In the features used are the activations of the layer before the last of the encoder, which are 10241024D vectors. We refer to those as long features. We also use, as features, the 5050D vectors taken from the latent space, W\mathcal{W}. We refer to those as short features.

MNIST has an official split into training and testing sets of sizes 60000 and 10000 respectively. We refer to it as different writers (DW) setting since the human writers of the digits for the training set are different from those who wrote the testing digits. We consider also a same writers (SW) setting, which uses only the official training split by further splitting it in two parts: a train split of size 50000 and a test split of size 10000, while the official testing split is ignored. In SW the pools of writers in the train and test splits overlap, whereas in DW they do not. This makes SW an easier setting than DW.

The most significant result of Table 2 is drawn by comparing the 1NN with the corresponding linear SVM columns. Since 1NN does not presume disentanglement in order to be effective, but linear SVM does, larger performance drops signal stronger entanglement. ALAE is the approach that remains more stable when switching from 1NN to linear SVM, suggesting a greater disentanglement of the space. This is true especially for short features, whereas for long features this effect fades away because linear separability grows.

Another observation is about SW vs. DW. 1NN generalizes less effectively for DW, as expected, but linear SVM provides a small improvement. This is unclear, but we speculate that DW might have fewer writers in the test set, and potentially slightly less challenging.

Figure 3 shows qualitative reconstruction results. It can be seen that BiGAN reconstructions are subject to semantic label flipping much more often than ALAE. Finally, Figure 4 shows two traversals: one obtained by interpolating in the Z\mathcal{Z} space, and the other by interpolating in the W\mathcal{W} space. The second shows a smoother image space transition, suggesting a lesser degree of entanglement.

2 Learning style representations

FFHQ. We evaluate StyleALAE with the FFHQ dataset. It is very recent and consists of 70000 images of people faces aligned and cropped at resolution of 1024×10241024\times 1024. In contrast to , we split FFHQ into a training set of 60000 images and a testing set of 10000 images. We do so in order to measure the reconstruction quality for which we need images that were not used during training.

We implemented our approach with PyTorch. Most of the experiments were conducted on a machine with 4×\times GPU Titan X, but for training the models at resolution 1024×10241024\times 1024 we used a server with 8×\times GPU Titan RTX. We trained StyleALAE for 147 epochs, 18 of which were spent at resolution 1024×10241024\times 1024. Starting from resolution 4×44\times 4 we grew StyleALAE up to 1024×10241024\times 1024. When growing to a new resolution level we used 500500k training samples during the transition, and another 500500k samples for training stabilization. Once reached the maximum resolution of 1024×10241024\times 1024, we continued training for 11M images. Thus, the total training time measured in images was 1010M. In contrast, the total training time for StyleGAN was 2525M images, and 1515M of them were used at resolution 1024×10241024\times 1024. At the same resolution we trained StyleALAE with only 11M images, so, 15 times less.

Table 3 reports the FID score for generations and reconstructions. Source images for reconstructions are from the test set and were not used during training. The scores of StyleALAE are higher, and we regard the large training time difference between StyleALAE and StyleGAN (11M vs 1515M) as the likely cause of the discrepancy.

Table 4 reports the perceptual path length (PPL) of SyleALAE. This is a measurement of the degree of disentanglement of representations. We compute the values for representations in the W\mathcal{W} and the Z\mathcal{Z} space, where StyleALAE is trained with style mixing in both cases. The StyleGAN score measured in Z\mathcal{Z} corresponds to a traditional network, and in W\mathcal{W} for a style-based one. We see that the PPL drops from Z\mathcal{Z} to W\mathcal{W}, indicating that W\mathcal{W} is perceptually more linear than Z\mathcal{Z}, thus less entangled. Also, note that for our models the PPL is lower, despite the higher FID scores.

Figure 6 shows a random collection of generations obtained from StyleALAE. Figure 5 instead shows a collection of reconstructions. In Figure 9 instead, we repeat the style mixing experiment in , but with real images as sources and destinations for style combinations. We note that the original images are faces of celebrities that we downloaded from the internet. Therefore, they are not part of FFHQ, and come from a different distribution. Indeed, FFHQ is made of face images obtained from Flickr.com depicting non-celebrity people. Often the faces do not wear any makeup, neither have the images been altered (e.g., with Photoshop). Moreover, the imaging conditions of the FFHQ acquisitions are very different from typical photoshoot stages, where professional equipment is used. Despite this change of image statistics, we observe that StyleALAE works effectively on both reconstruction and mixing.

LSUN. We evaluated StyleALAE with LSUN Bedroom . Figure 7 shows generations and reconstructions from unseen images during training. Table 3 reports the FID scores on the generations and the reconstructions.

CelebA-HQ. CelebA-HQ is an improved subset of CelebA consisting of 30000 images at resolution 1024×10241024\times 1024. We follow and use CelebA-HQ downscaled to 256×256256\times 256 with training/testing split of 27000/3000. Table 5 reports the FID and PPL scores, and Figure 8 compares StyleALE reconstructions of unseen faces with two other approaches.

Conclusions

We introduced ALAE, a novel autoencoder architecture that is simple, flexible and general, as we have shown to be efective with two very different backbone generator-encoder networks. Differently from previous work it allows learning the probability distribution of the latent space, when the data distribution is learned in adversarial settings. Our experiments confirm that this enables learning representations that are likely less entangled. This allows us to extend StyleGAN to StyleALAE, the first autoencoder capable of generating and manipulating images in ways not possible with SyleGAN alone, while maintaining the same level of visual detail.

This material is based upon work supported by the National Science Foundation under Grants No. OIA-1920920, and OAC-1761792.

References