Adversarial Latent Autoencoders
Stanislav Pidhorskyi, Donald Adjeroh, Gianfranco Doretto
Introduction
Generative Adversarial Networks (GAN) have emerged as one of the dominant unsupervised approaches for computer vision and beyond. Their strength relates to their remarkable ability to represent complex probability distributions, like the face manifold , or the bedroom images manifold , which they do by learning a generator map from a known distribution onto the data space. Just as important are the approaches that aim at learning an encoder map from the data to a latent space. They allow learning suitable representations of the data for the task at hand, either in a supervised , or unsupervised manner.
Autoencoder (AE) networks are unsupervised approaches aiming at combining the “generative” as well as the “representational” properties by learning simultaneously an encoder-generator map. General issues subject of investigation in AE structures are whether they can: (a) have the same generative power as GANs; and, (b) learn disentangled representations . Several works have addressed (a) . An important testbed for success has been the ability for an AE to generate face images as rich and sharp as those produced by a GAN . Progress has been made but victory has not been declared. A sizable amount of work has addressed also (b) , but not jointly with (a).
We introduce an AE architecture that is general, and has generative power comparable to GANs while learning a less entangled representation. We observed that every AE approach makes the same assumption: the latent space should have a probability distribution that is fixed a priori and the autoencoder should match it. On the other hand, it has been shown in , the state-of-the-art for synthetic image generation with GANs, that an intermediate latent space, far enough from the imposed input space, tends to have improved disentanglement properties.
We designed two ALAEs, one with a multilayer perceptron (MLP) as encoder with a symmetric generator, and another with the generator derived from a StyleGAN , which we call StyleALAE. For this one, we designed a companion encoder and a progressively growing architecture. We verified qualitatively and quantitatively that both architectures learn a latent space that is more disentangled than the imposed one. In addition, we show qualitative and quantitative results about face and bedroom image generation that are comparable with StyleGAN at the highest resolution of . Since StyleALAE learns also an encoder network, we are able to show at the highest resolution, face reconstructions as well as several image manipulations based on real images rather then generated.
Related Work
Our approach builds directly on the vanilla GAN architecture . Since then, a lot of progress has been made in the area of synthetic image generation. LAPGAN and StackGAN train a stack of GANs organized in a multi-resolution pyramid to generate high-resolution images. HDGAN improves by incorporating hierarchically-nested adversarial objectives inside the network hierarchy. In they use a multi-scale generator and discriminator architecture to synthesize high-resolution images with a GAN conditioned on semantic label maps, while in BigGAN they improve the synthesis by applying better regularization techniques. In PGGAN it is shown how high-resolution images can be synthesized by progressively growing the generator and the discriminator of a GAN. The same principle was used in StyleGAN , the current state-of-the-art for face image generation, which we adapt it here for our StyleALAE architecture. Other recent work on GANs has focussed on improving the stability and robustness of the training . New loss functions have been introduced , along with gradient regularization methods , weight normalization techniques , and learning rate equalization . Our framework is amenable to these improvements, as we explain in later sections.
Variational AE architectures have not only been appreciated for their theoretical foundation, but also for their stability during training, and the ability to provide insightful representations. Indeed, they stimulated research in the area of disentanglement , allowing learning representations with controlled degree of disentanglement between factors of variation in , and subsequent improvements in , leading to more elaborate metrics for disentanglement quantification , which we also use to analyze the properties of our approach. VAEs have also been extended to learn a latent prior different than a normal distribution, thus achieving significantly better models .
A lot of progress has been made towards combining the benefits of GANs and VAEs. AAE has been the precursor of those approaches, followed by VAE/GAN with a more direct approach. BiGAN and ALI provide an elegant framework fully adversarial, whereas VEEGAN and AGE pioneered the use of the latent space for autoencoding and advocated the reduction of the architecture complexity. PIONEER and IntroVAE followed this line, with the latter providing the best generation results in this category. Section 4.1 describes how the proposed approach compares with those listed here.
Finally, we quickly mention other approaches that have shown promising results with representing image data distributions. Those include autoregressive and flow-based methods . The former forego the use of a latent representation, but the latter does not.
Preliminaries
Following the more general formulation introduced in , the GAN learning problem entails finding the minimax with respect to the pair (i.e., the Nash equilibrium), of the value function defined as
Adversarial Latent Autoencoders
We introduce a novel autoencoder architecture by modifying the original GAN paradigm. We begin by decomposing the generator and the discriminator in two networks: , , and , , respectively. This means that
see Figure 1. In addition, we assume that the latent spaces at the interface between and , and between and are the same, and we indicate them as . In the most general case we assume that is a deterministic map, whereas we allow and to be stochastic. In particular, we assume that might optionally depend on an independent noisy input , with a known fixed distribution . We indicate with this more general stochastic generator.
Under the above conditions we now consider the distributions at the output of every network. The network simply maps onto . At the output of the distribution can be written as
where is the conditional distribution representing . Similarly, for the output of the distribution becomes
where is the conditional distribution representing . In (4) if we replace with we obtain the distribution , which describes the output of when the real data distribution is its input.
Since optimizing (1) leads toward the synthetic distribution matching the real one, i.e., , it is obvious from (4) that doing so also leads toward having . In addition to that, we propose to ensure that the distribution of the output of be the same as the distribution at the input of . This means that we set up an additional goal, which requires that
In this way we could interpret the pair of networks as a generator-encoder network that autoencodes the latent space .
If we indicate with a measure of discrepancy between two distributions and , we propose to achieve the goal (5) via regularizing the GAN loss (1) by alternating the following two optimizations
where the left and right arguments of indicate the distributions generated by the networks mapping , which correspond to and , respectively. We refer to a network optimized according to (6) (7) as an Adversarial Latent Autoencoder (ALAE). The building blocks of an ALAE architecture are depicted in Figure 1.
Data distribution. In architectures composed by an encoder network and a generator network, the task of the encoder is to map input data onto a space characterized by a latent distribution, whereas the generator is tasked to map latent codes onto a space described by a data distribution. Different strategies are used to learn the data distribution. For instance, some approaches impose a similarity criterion on the output of the generator , or even learn a similarity metric . Other techniques instead, set up an adversarial game to ensure the generator output matches the training data distribution . This latter approach is what we use for ALAE.
Latent distribution. For the latent space instead, the common practice is to set a desired target latent distribution, and then the encoder is trained to match it either by minimizing a divergence type of similarity , or by setting up an adversarial game . Here is where ALAE takes a fundamentally different approach. Indeed, we do not impose the latent distribution, i.e., , to match a target distribution. The only condition we set, is given by (5). In other words, we do not want to be the identity map, and are very much interested in letting the learning process decide what should be.
Reciprocity. Another aspect of autoecoders is whether and how they achieve reciprocity. This property relates to the ability of the architecture to reconstruct a data sample from its code , and viceversa. Clearly, this requires that , or equivalently that . In the first case, the network must contain a reconstruction term that operates in the data space. In the latter one, the term operates in the latent space. While most approaches follow the first strategy , there are some that implement the second , including ALAE. Indeed, this can be achieved by choosing the divergence in (7) to be the expected coding reconstruction error, as follows
StyleALAE
We use ALAE to build an autoencoder that uses a StyleGAN based generator. For this we make our latent space play the same role as the intermediate latent space in . Therefore, our network becomes the part of StyleGAN depicted on the right side of Figure 2. The left side is a novel architecture that we designed to be the encoder .
Since at every layer, is driven by a style input, we design symmetrically, so that from a corresponding layer we extract style information. We do so by inserting Instance Normalization (IN) layers , which provide instance averages and standard deviations for every channel. Specifically, if is the output of the -th layer of , the IN module extracts the statistics and representing the style at that level. The IN module also provides as output the normalized version of the input, which continues down the pipeline with no more style information from that level. Given the information flow between and , the architecture is effectively mimicking a multiscale style transfer from to , with the difference that there is not an extra input image that provides the content .
The set of styles that are inputs to the Adaptive Instance Normalization (AdaIN) layers in are related linearly to the latent variable . Therefore, we propose to combine the styles output by the encoder, and to map them onto the latent space, via the following multilinear map
where the ’s are learnable parameters, and is the number of layers.
Similarly to we use progressive growing. We start from low-resolution images ( pixels) and progressively increase the resolution by smoothly blending in new blocks to and . For the and networks we implement them using MLPs. The and spaces, and all layers of and have the same dimensionality in all our experiments. Moreover, for StyleALAE we follow , and chose to have 8 layers, and we set to have 3 layers.
Implementation
Adversarial losses and regularization. We use a non-saturating loss , which in (1) we introduce by setting to be a SoftPlus function . This is a smooth version of the rectifier activation function, defined as . In addition, we use gradient regularization techniques . We utilize , a zero-centered gradient penalty term which acts only on real data, and is defined as , where the gradient is taken with respect to the parameters and of the networks and , respectively.
Training. In order to optimizate (6) (7) we use alternating updates. One iteration is composed of three updating steps: two for (6) and one for (7). Step I updates the discriminator (i.e., networks and ). Step II updates the generator (i.e., networks and ). Step III updates the latent space autoencoder (i.e., networks and ). The procedural details are summarized in Algorithm 1. For updating the weights we use the Adam optimizer with and , coupled with the learning rate equalization technique described below. For non-growing architectures (i.e., MLPs) we use a learning rate of , and batch size of 128. For growing architectures (i.e., StyleALAE) learning rate and batch size depend on the resolution.
Experiments
Code and uncompressed images are available at https://github.com/podgorskiy/ALAE.
We train ALAE with MNIST , and then use the feature representation for classification, reconstruction, and analyzing disentanglement. We use the permutation-invariant setting, where each MNIST image is treated as a D vector without spatial structure, which requires to use a MLP instead of a CNN. We follow and use a three layer MLP with a latent space size of D. Both networks, and have two hidden layers with 1024 units each. In the features used are the activations of the layer before the last of the encoder, which are D vectors. We refer to those as long features. We also use, as features, the D vectors taken from the latent space, . We refer to those as short features.
MNIST has an official split into training and testing sets of sizes 60000 and 10000 respectively. We refer to it as different writers (DW) setting since the human writers of the digits for the training set are different from those who wrote the testing digits. We consider also a same writers (SW) setting, which uses only the official training split by further splitting it in two parts: a train split of size 50000 and a test split of size 10000, while the official testing split is ignored. In SW the pools of writers in the train and test splits overlap, whereas in DW they do not. This makes SW an easier setting than DW.
The most significant result of Table 2 is drawn by comparing the 1NN with the corresponding linear SVM columns. Since 1NN does not presume disentanglement in order to be effective, but linear SVM does, larger performance drops signal stronger entanglement. ALAE is the approach that remains more stable when switching from 1NN to linear SVM, suggesting a greater disentanglement of the space. This is true especially for short features, whereas for long features this effect fades away because linear separability grows.
Another observation is about SW vs. DW. 1NN generalizes less effectively for DW, as expected, but linear SVM provides a small improvement. This is unclear, but we speculate that DW might have fewer writers in the test set, and potentially slightly less challenging.
Figure 3 shows qualitative reconstruction results. It can be seen that BiGAN reconstructions are subject to semantic label flipping much more often than ALAE. Finally, Figure 4 shows two traversals: one obtained by interpolating in the space, and the other by interpolating in the space. The second shows a smoother image space transition, suggesting a lesser degree of entanglement.
2 Learning style representations
FFHQ. We evaluate StyleALAE with the FFHQ dataset. It is very recent and consists of 70000 images of people faces aligned and cropped at resolution of . In contrast to , we split FFHQ into a training set of 60000 images and a testing set of 10000 images. We do so in order to measure the reconstruction quality for which we need images that were not used during training.
We implemented our approach with PyTorch. Most of the experiments were conducted on a machine with 4 GPU Titan X, but for training the models at resolution we used a server with 8 GPU Titan RTX. We trained StyleALAE for 147 epochs, 18 of which were spent at resolution . Starting from resolution we grew StyleALAE up to . When growing to a new resolution level we used k training samples during the transition, and another k samples for training stabilization. Once reached the maximum resolution of , we continued training for M images. Thus, the total training time measured in images was M. In contrast, the total training time for StyleGAN was M images, and M of them were used at resolution . At the same resolution we trained StyleALAE with only M images, so, 15 times less.
Table 3 reports the FID score for generations and reconstructions. Source images for reconstructions are from the test set and were not used during training. The scores of StyleALAE are higher, and we regard the large training time difference between StyleALAE and StyleGAN (M vs M) as the likely cause of the discrepancy.
Table 4 reports the perceptual path length (PPL) of SyleALAE. This is a measurement of the degree of disentanglement of representations. We compute the values for representations in the and the space, where StyleALAE is trained with style mixing in both cases. The StyleGAN score measured in corresponds to a traditional network, and in for a style-based one. We see that the PPL drops from to , indicating that is perceptually more linear than , thus less entangled. Also, note that for our models the PPL is lower, despite the higher FID scores.
Figure 6 shows a random collection of generations obtained from StyleALAE. Figure 5 instead shows a collection of reconstructions. In Figure 9 instead, we repeat the style mixing experiment in , but with real images as sources and destinations for style combinations. We note that the original images are faces of celebrities that we downloaded from the internet. Therefore, they are not part of FFHQ, and come from a different distribution. Indeed, FFHQ is made of face images obtained from Flickr.com depicting non-celebrity people. Often the faces do not wear any makeup, neither have the images been altered (e.g., with Photoshop). Moreover, the imaging conditions of the FFHQ acquisitions are very different from typical photoshoot stages, where professional equipment is used. Despite this change of image statistics, we observe that StyleALAE works effectively on both reconstruction and mixing.
LSUN. We evaluated StyleALAE with LSUN Bedroom . Figure 7 shows generations and reconstructions from unseen images during training. Table 3 reports the FID scores on the generations and the reconstructions.
CelebA-HQ. CelebA-HQ is an improved subset of CelebA consisting of 30000 images at resolution . We follow and use CelebA-HQ downscaled to with training/testing split of 27000/3000. Table 5 reports the FID and PPL scores, and Figure 8 compares StyleALE reconstructions of unseen faces with two other approaches.
Conclusions
We introduced ALAE, a novel autoencoder architecture that is simple, flexible and general, as we have shown to be efective with two very different backbone generator-encoder networks. Differently from previous work it allows learning the probability distribution of the latent space, when the data distribution is learned in adversarial settings. Our experiments confirm that this enables learning representations that are likely less entangled. This allows us to extend StyleGAN to StyleALAE, the first autoencoder capable of generating and manipulating images in ways not possible with SyleGAN alone, while maintaining the same level of visual detail.
This material is based upon work supported by the National Science Foundation under Grants No. OIA-1920920, and OAC-1761792.