Diffusion Models already have a Semantic Latent Space

Mingi Kwon, Jaeseok Jeong, Youngjung Uh

Introduction

In image synthesis, diffusion models have advanced to achieve state-of-the-art performance regarding quality and mode coverage since the introduction of denoising diffusion probabilistic models (Ho et al., 2020). They disrupt images by adding noise through multiple steps of forward process and generate samples by progressive denoising through multiple steps of reverse (i.e., generative) process. Since their deterministic version provides nearly perfect reconstruction of original images (Song et al., 2020a), they are suitable for image editing, which renders target attributes on the real images. However, simply editing the latent variables (i.e., intermediate noisy images) causes degraded results (Kim & Ye, 2021). Instead, they require complicated procedures: providing guidance in the reverse process or finetuning models for an attribute.

Figure 1(a-c) briefly illustrates the existing approaches. Image guidance mixes the latent variables of the guiding image with unconditional latent variables (Choi et al., 2021; Lugmayr et al., 2022; Meng et al., 2021). Though it provides some control, it is ambiguous to specify which attribute to reflect among the ones in the guide and the unconditional result, and it lacks intuitive control for the magnitude of change. Classifier guidance manipulates images by imposing gradients of a classifier on the latent variables in the reverse process to match the target class (Dhariwal & Nichol, 2021; Avrahami et al., 2022; Liu et al., 2021). It requires training an extra classifier for the latent variables, i.e., noisy images. Furthermore, computing gradients through the classifier during sampling is costly. Finetuning the whole model can steer the resulting images to the target attribute without the above problems (Kim & Ye, 2021). Still, it requires multiple models to reflect multiple descriptions.

On the other hand, generative adversarial networks (Goodfellow et al., 2020) inherently provide straightforward image editing in their latent space. Given a latent vector for an original image, we can find the direction in the latent space that maximizes the similarity of the resulting image with a target description in CLIP embedding (Patashnik et al., 2021). The latent direction found on one image leads to the same manipulation of other images. However, given a real image, finding its exact latent vector is often challenging and produces unexpected appearance changes.

It would allow admirable image editing if the diffusion models with the nearly perfect inversion property have such a semantic latent space. Preechakul et al. (2022) introduces an additional input to the reverse diffusion process: a latent vector from an original image embedded by an extra encoder. This latent vector contains the semantics to condition the process. However, it requires training from scratch and does not match with pretrained diffusion models.

In this paper, we propose an asymmetric reverse process (Asyrp) which discovers the semantic latent space of a frozen diffusion model such that modifications in the space edits attributes of the original images. Our semantic latent space, named h-space, has the properties necessary for editing applications as follows. The same shift in this space results in the same attribute change in all images. Linear changes in this space lead to linear changes in attributes. The changes do not degrade the quality of the resulting images. The changes throughout the timesteps are almost identical to each other for a desired attribute change. Figure 1(d) illustrates some of these properties and §\S 5.3 provides detailed analyses. To the best of our knowledge, it is the first attempt to discover the semantic latent space in the frozen pretrained diffusion models. Spoiler alert: our semantic latent space is different from the intermediate latent variables in the diffusion process. Moreover, we introduce a principled design of the generative process for versatile editing and quality boosting by quantifiable measures: editing strength of an interval and quality deficiency at a timestep. Extensive experiments demonstrate that our method is generally applicable to various architectures (DDPM++, iDDPM, and ADM) and datasets (CelebA-HQ, AFHQ-dog, LSUN-church, LSUN-bedroom, and MetFaces).

Background

We briefly describe essential backgrounds. The rest of the related work is deferred to Appendix A.

DDPM is a latent variable model that learns a data distribution by denoising noisy images (Ho et al., 2020). The forward process diffuses the data samples through Gaussian transitions parameterized with a Markov process:

where {βt}t=1T\{\beta_{t}\}^{T}_{t=1} is the variance schedule and αt=∏s=1t(1−βs)\alpha_{t}=\prod^{t}_{s=1}(1-\beta_{s}). Then the reverse process becomes pθ(x0:T):=p(xT)∏t=1Tpθ(xt−1∣xt)p_{\theta}\left({\bm{x}}_{0:T}\right):=p\left({\bm{x}}_{T}\right)\prod_{t=1}^{T}p_{\theta}\left({\bm{x}}_{t-1}\mid{\bm{x}}_{t}\right), starting from xT∼N(0,I){\bm{x}}_{T}\sim\mathcal{N}(0,\mathbf{I}) with noise predictor ϵtθ\bm{\epsilon}^{\theta}_{t}:

where zt∼N(0,I){\bm{z}}_{t}\sim\mathcal{N}(0,\mathbf{I}) and σt2\sigma_{t}^{2} is a variance of the reverse process which is set to σt2=βt\sigma_{t}^{2}=\beta_{t} by DDPM.

2 Denoising Diffusion Implicit Model (DDIM)

DDIM redefines Eq. (1) as qσ(xt−1∣xt,x0)=N(αt−1x0+1−αt−1−σt2⋅xt−αtx01−αt,σt2I)q_{\sigma}({\bm{x}}_{t-1}|{\bm{x}}_{t},{\bm{x}}_{0})=\mathcal{N}(\sqrt{\alpha_{t-1}}{\bm{x}}_{0}+\sqrt{1-\alpha_{t-1}-\sigma_{t}^{2}}\cdot\frac{{\bm{x}}_{t}-\sqrt{\alpha_{t}}{\bm{x}}_{0}}{\sqrt{1-\alpha_{t}}},\sigma_{t}^{2}\bm{I}) which is a non-Markovian process (Song et al., 2020a). Accordingly, the reverse process becomes

where σt=η(1−αt−1)/(1−αt)1−αt/αt−1\sigma_{t}=\eta\sqrt{\left(1-\alpha_{t-1}\right)/\left(1-\alpha_{t}\right)}\sqrt{1-\alpha_{t}/\alpha_{t-1}}. When η=1\eta=1 for all tt, it becomes DDPM. As η=0\eta=0, the process becomes deterministic and guarantees nearly perfect inversion.

3 Image manipulation with CLIP

CLIP learns multimodal embeddings with an image encoder EIE_{I} and a text encoder ETE_{T} whose similarity indicates semantic similarity between images and texts (Radford et al., 2021). Compared to directly minimizing the cosine distance between the edited image and the target description (Patashnik et al., 2021), directional loss with cosine distance achieves homogeneous editing without mode collapse (Gal et al., 2021):

Discovering semantic latent space in diffusion models

This section explains why naive approaches do not work and proposes a new controllable reverse process. Then we describe the techniques for controlling the generative process. Throughout this paper, we use an abbreviated version of Eq. (3):

where Pt(ϵtθ(xt))\mathbf{P}_{t}(\bm{\epsilon}^{\theta}_{t}({\bm{x}}_{t})) denotes the predicted x0{\bm{x}}_{0} and Dt(ϵtθ(xt))\mathbf{D}_{t}(\bm{\epsilon}^{\theta}_{t}({\bm{x}}_{t})) denotes the direction pointing to xt{\bm{x}}_{t}. We omit σtzt\sigma_{t}{\bm{z}}_{t} for brevity, except when η≠0\eta\neq 0. We further abbreviate Pt(ϵtθ(xt))\mathbf{P}_{t}(\bm{\epsilon}^{\theta}_{t}({\bm{x}}_{t})) as Pt\mathbf{P}_{t} and Dt(ϵtθ(xt))\mathbf{D}_{t}(\bm{\epsilon}^{\theta}_{t}({\bm{x}}_{t})) as Dt\mathbf{D}_{t} when the context clearly specifies the arguments.

We aim to allow semantic latent manipulation of images x0{\bm{x}}_{0} generated from xT{\bm{x}}_{T} given a pretrained and frozen diffusion model. The easiest idea to manipulate x0{\bm{x}}_{0} is simply updating xT{\bm{x}}_{T} to optimize the directional CLIP loss given text prompts with Eq. (4). However, it leads to distorted images or incorrect manipulation (Kim & Ye, 2021).

An alternative approach is to shift the noise ϵtθ\bm{\epsilon}^{\theta}_{t} predicted by the network at each sampling step. However, it does not achieve manipulating x0{\bm{x}}_{0} because the intermediate changes in Pt\mathbf{P}_{t} and Dt\mathbf{D}_{t} cancel out each other resulting in the same pθ(x0:T)p_{\theta}({\bm{x}}_{0:T}), similarly to destructive interference.

2 Asymmetric reverse process

In order to break the interference, we propose a new controllable reverse process with asymmetry:

3 h-space

Note that ϵtθ\bm{\epsilon}^{\theta}_{t} is implemented as U-Net in all state-of-the-art diffusion models. We choose its bottleneck, the deepest feature maps ht{\bm{h}}_{t}, to control ϵtθ\bm{\epsilon}^{\theta}_{t}. By design, ht{\bm{h}}_{t} has smaller spatial resolutions and high-level semantics than ϵtθ\bm{\epsilon}^{\theta}_{t}. Accordingly, the sampling equation becomes

We observe that h-space in Asyrp has the following properties that others do not have.

The same Δh\Delta{\bm{h}} leads to the same effect on different samples.

Linearly scaling Δh\Delta{\bm{h}} controls the magnitude of attribute change, even with negative scales.

Adding multiple Δh\Delta{\bm{h}} manipulates the corresponding multiple attributes simultaneously.

Δh\Delta{\bm{h}} preserves the quality of the resulting images without degradation.

Δht\Delta{\bm{h}}_{t} is roughly consistent across different timesteps tt.

The above properties are demonstrated thoroughly in §\S 5.3. Appendix D.3 provides details of h-space and suboptimal results from alternative choices.

4 Implicit neural directions

Generative process design

This section describes the entire editing process, which consists of three phases: editing with Asyrp, traditional denoising, and quality boosting. We design formulas to determine the length of each phase with quantifiable measures.

2 Quality boosting with stochastic noise injection

3 Overall process of image editing

The visual overview and comprehensive algorithms of the entire process are in Appendix I.

Experiments

In this section, we show the effectiveness of semantic latent editing in h-space with Asyrp on various attributes, datasets and architectures in §\S 5.1. Moreover, we provide quantitative results including user study in §\S 5.2. Lastly, we provide detailed analyses for the properties of the semantic latent space on h-space and alternatives in §\S 5.3.

1 Versatility of h-space with Asyrp

Figure 5 shows the effectiveness of our method on various datasets and any existing U-Net based architectures. Our method can synthesize the attributes that are not even included in the training dataset, such as church →\xrightarrow{} {department, factory, and temple}. Even for dogs, our method synthesizes smiling Poodle and Yorkshire, the species that barely smile in the dataset. Figure 5 provides results for changing human faces to different identities, painting styles, and ancient primates. More result can be found in Appendix N. Versatility of our method is surprising because we do not alter the models but only shift the bottleneck feature maps in h-space with Asyrp during inference.

2 Quantitative comparison

3 Analysis on h-space

We provide detailed analyses to validate the properties of semantic latent space for diffusion models: homogeneity, linearity, robustness, and consistency across timesteps.

In Figure 7, we observe that linearly scaling a Δh\Delta{\bm{h}} reflects the amount of change in the visual attributes. Surprisingly, it generalizes to negative scales that are not seen during training. Moreover, Figure 10 shows that combinations of different Δh\Delta{\bm{h}}’s yield their combined semantic changes in the resulting images. Appendix N.2 provides mixed interpolation between multiple attributes.

Figure 10 compares the effect of adding random noise in h-space and ϵ\bm{\epsilon}-space. The random noises are chosen to be the vectors with random directions and magnitude of the example Δht\Delta{\bm{h}}_{t} and Δϵt\Delta\bm{\epsilon}_{t} in Figure 7 on each space. Perturbation in h-space leads to realistic images with a minimal difference or some semantic changes. On the contrary, perturbation in ϵ\bm{\epsilon}-space severely distorts the resulting images. See Appendix D.2 for more analyses.

Conclusion

We proposed a new generative process, Asyrp, which facilitates image editing in a semantic latent space h-space for pretrained diffusion models. h-space has nice properties as in the latent space of GANs: homogeneity, linearity, robustness, and consistency across timesteps. The full editing process is designed to achieve versatile editing and high quality by measuring editing strength and quality deficiency at timesteps. We hope that our approach and detailed analyses help cultivate a new paradigm of image editing in the semantic latent space of diffusion models. Combining previous finetuning or guidance techniques would be an interesting research direction.

This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (Ministry of Science and ICT) (No. 2021-0-00155)

References

Appendix

After Sohl-Dickstein et al. (2015), denoising diffusion probabilistic models (DDPMs) provide a universal approach for generative modeling (Ho et al., 2020). On the other hand, Song et al. (2020b) suggests score-based model and unifies SDEs incorporating diffusion models with score-based models. Subsequent works renovate diffusion models by focusing on architectures, scheduling, weighting, and fast sampling (Nichol & Dhariwal (2021), Karras et al. (2022), Choi et al. (2022), Song et al. (2020a), Watson et al. (2022)). They mainly consider random generation rather than controlled generation.

In the meantime, Dhariwal & Nichol (2021) introduces classifier guidance not only improving the quality of images but also retrieving specific class of images. Since it can apply any guidance, its variants have emerged (Sehwag et al. (2022), Avrahami et al. (2022), Liu et al. (2021), Nichol et al. (2021)). However it requires a noise-dependent classifier (or any off-the-shelf models) and additional cost to compute gradients for the guidance during its sampling process. The other works try to control the generative process using image-space guidance (Choi et al. (2021), Meng et al. (2021), Lugmayr et al. (2022), Avrahami et al. (2022)). They manipulate resulting images by matching noisy images with target images during the reverse process. Still, it is hard to expect delicate control of the reverse process from the image guidance. Furthermore, Preechakul et al. (2022) introduces an extra encoder which encodes the semantic features of a real image in order to condition the generative process. Although the semantics allow one to control diffusion models, it requires additional training from scratch with the encoder and inherently can not use the other pretrained diffusion models.

For controllability, Rombach et al. (2022) and Vahdat et al. (2021) apply another approach which adapts VAE (Kingma & Welling, 2013) and autoencoder (Rumelhart et al., 1985) to diffusion models. In spite of their great success in editing, their diffusion models learn the distribution of the learned embeddings in VAE or autoencoder, not the images. Kim & Ye (2021) proposes another strategy: fine-tuning a whole diffusion model for image editing. It shows valid performance but it requires each fine-tuned model corresponding each attribute.

In comparison, Asyrp enables outstanding manipulation without high computation, specifically designed architectures, or fine-tuning whole models.

Meanwhile, generative adversarial network Goodfellow et al. (2020) address their latent space for image editing (Ling et al. (2021), Härkönen et al. (2020), Chefer et al. (2021), Shen et al. (2020), Yüksel et al. (2021), Patashnik et al. (2021), Gal et al. (2021), Dai et al. (2019), Xu et al. (2022)). However they have to conduct ‘inversion’ to their latent space for real image editing and ‘GAN inversion’ is often challenging and produces unexpected appearance changes.

On the contrary, Asyrp enables to use latent space of real images by nearly perfect easy inversion of DDIM.

Appendix B More discussion

In this section, we discuss the pros and cons of diffusion-model-based and GAN-based methods. And we provide guidelines for further improvements.

GAN-based latent manipulation methods Patashnik et al. (2021); Gal et al. (2021)require careful inversion from real images to latent codes for real image editing. On the contrary, our proposed method based on diffusion models has a powerful advantage; the sophisticated inversion method is not necessary. This means that we can obtain the latent code of an arbitrary real image even if the image is not in the trained domain. On the other hand, several inversion methods have been proposed for GANs to obtain the latent of the real image, and the corresponding latent manipulation method should be considered for each inversion method. For example, it is difficult to apply the method of editing in ww space to the method of inversion using w+w^{+} space.

However, GANs have the advantage of fast sampling. In addition, diffusion models have a relatively slow sampling time. Additionally, we have to be aware of the time steps of diffusion models, which is still less well known.

The advantage of being free from Inversion provides the following milestones. The manipulation in the latent of the diffusion models is the same as the editing in real images. It can be expanded to segmentation, clustering, classification, etc. in h-space for real-world images.

It would be an interesting research direction to employ previous techniques. Our method can be used in conjunction with gradient guidance methods. Although we do not focus on random sampling, ours works effectively for sampling with stochastic. (See §\S M.) It may bring more diverse methods to steer diffusion models.

h-space in the latent diffusion models such as stable diffusion, is another interesting research direction. The main contribution of our paper is only modifying Pt\mathbf{P}_{t} while preserving Dt\mathbf{D}_{t}, and can be adapted with latent diffusion models. However, since the latent meaning may be different due to structural differences, research on this is needed.

Furthermore, all of the properties of h-space according to the time step has not been fully discussed so far. Research on them can be expected to expand further.

Editing with Asyrp seldom yields changes in overall style or peripheral objects but edits attributes of the main object. Style transfer using frozen diffusion models is our future work.

Techniques for high-quality image manipulation such as Asyrp should be accompanied by social and/or technical solutions to prevent abuse. We acknowledge the potential ethical implications that may arise from the use of our image manipulation technique, Asyrp. We advocate for the development and implementation of social and technical solutions to prevent potential abuses such as spreading disinformation or propaganda. We are committed to ensuring fairness and non-discrimination, legal compliance, and research integrity in our work.

Appendix C Proof of Theorem 1

Appendix D Additional supports for h-space with Asyrp

In §\S 3.1, we argue that if both Pt\mathbf{P}_{t} and Dt\mathbf{D}_{t} are shifted, we can not manipulate x0{\bm{x}}_{0}. In Figure 13, we do not observe the noticeable difference between (a) and (b) which are the result of the original reverse process of DDIM and the one with shifting both terms, respectively.

D.2 Robustness and semantics in h-space and ϵitalic-ϵ\epsilon-space with Asyrp

In §\S 3.2, we also argue that h-space is more robust than ϵ\bm{\epsilon}-space with Asyrp. In Figure 13, we observe that small random noise z∼N(0,I)z\sim\mathcal{N}(0,\mathbf{I}) in ϵ\bm{\epsilon}-space degrades the resulting image without semantic changes (c) and much larger random noise in h-space yields random semantic changes without severe artifacts (d).

D.3 Choice of h-space in U-Net

As shown in Figure 14, there are many other candidates for h-space in the architecture. Among the layers, we choose the 8th layer, the bridge of the U-Net based architecture. The layer is not influenced by any skip connection, has the smallest spatial dimension with compressed information, and is located just before the upsampling blocks. Thus, we assume that it could possibly be considered as the most suitable latent embedding. To confirm the assumption, we train ft{\bm{f}}_{t} on the other layers. The results are shown in Figure 16. We carefully tuned the training hyperparameters (λCLIP\lambda_{CLIP} and λrecon\lambda_{recon}) for fair comparison. The 1st to the 6th layers hardly bring visible changes. The 7th and 9th layers bring not only the desired changes but also difficulty in finding optimal hyperparameters. After the 9th layer, the results bear severe artifacts.

Appendix E Implicit neural directions

Figure 16 illustrates the neural implicit function ft{\bm{f}}_{t}. It has only two 1x1 convolution layers with 512 channels. Note that we haven’t explored the network architecture much.

Figure 17 shows the quality improvements by non-accelerated sampling with scaled Δht\Delta{\bm{h}}_{t} described in §\S 3.4. Even with different number of inference steps, we observe similar changes of an attribute if we preserve the sum of Δht\Delta{\bm{h}}_{t}. This scaling technique allows non-accelerated sampling with 1000 steps for the models trained by accelerated training with 40 steps. Using non-accelerated sampling with scaled Δht\Delta{\bm{h}}_{t} leads to the same magnitude of manipulation and higher-quality images. In our experiments, it takes about 1.5 seconds to sample for 40 steps and 40 seconds for 1000 steps.

Appendix G Editing strength and editing flexibility

Appendix H Quality boosting

We validate the effectiveness of our quality boosting (§\S 4.2) in the original DDIM process and in Asyrp.

Figure 21 shows quality improvements by our quality boosting in Asyrp.

Appendix I Algorithm

Figure 24 illustrates generative process. Algorithm 1 and 2 describe training algorithm and inference algorithm of Asyrp, respectively.

Appendix J Training details

J.2 Training with random sampling instead of the training datasets

Apparently, for training, inverting real-images can be replaced by random sampling. It refers to using xT∼N(0,I){\bm{x}}_{T}\sim\mathcal{N}(0,\mathbf{I}) instead of xT=q(x0)⋅∏t=1Tq(xt∣xt−1){\bm{x}}_{T}=q({\bm{x}}_{0})\cdot\prod_{t=1}^{T}q\left({\bm{x}}_{t}\mid{\bm{x}}_{t-1}\right) where x0∼pdata(x){\bm{x}}_{0}\sim p_{data}(x). It allows us to train Asyrp only with the pretrained network and without extra dataset. Using random samples has tradeoff between preservation of contents and possible amount in editing. It take advantages when a target attribute requires large amount changes. We assume that the inversion of real-image is in the long tail of a Gaussian distribution because of realistic background or detailed clothes. On the other hand, random noise is considered to be closer to the mean of the normal, so it is easier to find directions. It can easily bring larger changes but also easily alter the contents. On the contrary, training with inversion shows the opposite property.

We train ft{\bm{f}}_{t} with random sampling for attributes whose identity preservation is not important to take advantage of these properties. The rightmost column in Table 3 shows the choices.

Appendix K Evaluation

We conduct user study to compare the performance of Asyrp and DiffusionCLIP (Kim & Ye, 2021) on Celeba-HQ (Liu et al., 2015) and LSUN-church (Yu et al., 2015). We use official checkpoints provided by DiffusionCLIP except for some facial attributes whose checkpoint do not exist. We tried our best to tune their hyperparameters following the manual for fair comparison. Example images are shown in Figure 26-27.

We use smiling and sad for in-domain CelebA-HQ attributes, Pixar and Neanderthal for unseen-domain CelebA-HQ attributes, and department store, ancient, and wooden for LSUN-church.

In unseen domain and Lsun-church, we use official checkpoints provided by DiffusionCLIP. We also randomly select 8 images for each CelebA-HQ attribute and 12 images for each LSUN-church attribute.

We observe that DiffusionCLIP works better in changing the holistic style of images. At the same time, it is short of the ability to bring semantic changes and suffers noisy results and a lack of diversity. The problems would be caused by fine-tuning the whole diffusion model.

We use the following questions for the survey. 1) Quality: Which image quality do you think is better? (clear and less noisy) 2) Attribute: Which image do you think is “Attribute(e.g., Smiling) naturally”? 3) Overall: Which image do you think is better considering the above evaluation criteria?

As for LSUN-church, we provide a set of four images at once and add a question: 3) Diversity: Which group do you think has a more diverse style? 4) Overall: Which image do you think is better considering the above evaluation criteria?

K.2 Segmentation consistency and directional CLIP similarity

Appendix L Directions

Figure 29 and Figure 30 show that the effects of mean direction and global direction are quite similar with Δht\Delta h_{t} by ft{\bm{f}}_{t} in various attributes. We compute mean direction and global direction from 20 different images.

L.2 Compare three methods

In this section, we compare three methods: implicit neural direction ftf_{t}, optimized Δht\Delta h_{t}, and optimized Δhglobal\Delta h^{global}.

ft≈Δhglobal<Δhtf_{t}\approx\Delta h^{global}<\Delta h_{t}

We have to optimize each Δht\Delta h_{t} for each time step tt. Additionally, it needs specific hyperparameters for each Δht\Delta h_{t}, e.g., higher learning rates for larger tt. On the contrary, time-consuming for ftf_{t} is similar to optimizing Δhglobal\Delta h^{global}.

ft≈Δht>Δhglobalf_{t}\approx\Delta h_{t}>\Delta h^{global}

ftf_{t} and Δht\Delta h_{t}, where directions can be obtained for each timestep, have the best quality. As can be shown in Figure 10, Δhglobal\Delta h^{global} is sometimes accompanied by slight differences in hair, etc.

Δht\Delta h_{t} can be obtained from ftf_{t}, and Δhglobal\Delta h^{global} can be obtained by aggregating Δht\Delta h_{t}.

We opt to use ftf_{t} for above three advantages.

Appendix M Random sampling

We conduct extra experiments: generating images with target attributes using Asyrp not from inversion but from random Gaussian noises. As a consequence, the generative process can be used for conditional random sampling. We provide the results in Figure 31. However, it is beyond the scope of this paper.

Appendix N More samples

We conduct extra experiments: editing images with target class using Asyrp with ImageNet pretrained model. We verified that models trained on large datasets, such as ImageNet, can be edited using Asyrp. However, we also observed that in this case the latent space is not partitioned by classes. For an orange, we have different latents for a single orange, for many oranges, for a cross-section of cut orange, and for a single piece of orange. Therefore, we learned the implicit function by collecting similar images to find the direction.

N.2 Multi-interpolation

Figure 33 provides mixed interpolation between multiple attributes. We observe that any interpolation with any attribute is possible.

N.3 More results on all datasets

We provide more results on CelebA-HQ (Figure 34), LSUN-church (Figure 35), AFHQ, LSUN-bedroom, MetFaces (Figure 36).