Image2StyleGAN++: How to Edit the Embedded Images?

Rameen Abdal, Yipeng Qin, Peter Wonka

Introduction

Recent GANs demonstrated that synthetic images can be generated with very high quality. This motivates research into embedding algorithms that embed a given photograph into a GAN latent space. Such embedding algorithms can be used to analyze the limitations of GANs , do image inpainting , local image editing , global image transformations such as image morphing and expression transfer , and few-shot video generation .

In this paper, we propose to extend a very recent embedding algorithm, Image2StyleGAN . In particular, we would like to improve this previous algorithm in three aspects. First, we noticed that the embedding quality can be further improved by including Noise space optimization into the embedding framework. The key insight here is that stable Noise space optimization can only be conducted if the optimization is done sequentially with W+W^{+} space and not jointly. Second, we would like to improve the capabilities of the embedding algorithm to increase the local control over the embedding. One way to improve local control is to include masks in the embedding algorithm with undefined content. The goal of the embedding algorithm should be to find a plausible embedding for everything outside the mask, while filling in reasonable semantic content in the masked pixels. Similarly, we would like to provide the option of approximate embeddings, where the specified pixel colors are only a guide for the embedding. In this way, we aim to achieve high quality embeddings that can be controlled by user scribbles. In the third technical part of the paper, we investigate the combination of embedding algorithm and direct manipulations of the activation maps (called activation tensors in our paper).

We propose Noise space optimization to restore the high frequency features in an image that cannot be reproduced by other latent space optimization of GANs. The resulting images are very faithful reconstructions of up to 45 dB compared to about 20 dB (PSNR) for the previously best results.

We propose an extended embedding algorithm into the W+W^{+} space of StyleGAN that allows for local modifications such as missing regions and locally approximate embeddings.

We investigate the combination of embedding and activation tensor manipulation to perform high quality local edits along with global semantic edits on images.

We apply our novel framework to multiple image editing and manipulation applications. The results show that the method can be successfully used to develop a state-of-the-art image editing software.

Related Work

Generative Adversarial Networks (GANs) are one of the most popular generative models that have been successfully applied to many computer vision applications, e.g. object detection , texture synthesis , image-to-image translation and video generation . Backing these applications are the massive improvements on GANs in terms of architecture , loss function design , and regularization . On the bright side, such improvements significantly boost the quality of the synthesized images. To date, the two highest quality GANs are StyleGAN and BigGAN . Between them, StyleGAN produces excellent results for unconditional image synthesis tasks, especially on face images; BigGAN produces the best results for conditional image synthesis tasks (e.g. ImageNet ). While on the dark side, these improvements make the training of GANs more and more expensive that nowadays it is almost a privilege of wealthy institutions to compete for the best performance. As a result, methods built on pre-trained generators start to attract attention very recently. In the following, we would like to discuss previous work of two such approaches: embedding images into a GAN latent space and the manipulation of GAN activation tensors.

The embedding of an image into the latent space is a longstanding topic in both machine learning and computer vision. In general, the embedding can be implemented in two ways: i) passing the input image through an encoder neural network (e.g. the Variational Auto-Encoder ); ii) optimizing a random initial latent code to match the input image . Between them, the first approach dominated for a long time. Although it has an inherent problem to generalize beyond the training dataset, it produces higher quality results than the naive latent code optimization methods . While recently, Abdal et al. obtained excellent embedding results by optimizing the latent codes in an enhanced W+W^{+} latent space instead of the initial ZZ latent space. Their method suggests a new direction for various image editing applications and makes the second approach interesting again.

Activation Tensor Manipulation.

With fixed neural network weights, the expression power of a generator can be fully utilized by manipulating its activation tensors. Based on this observation, Bau et al. investigated what a GAN can and cannot generate by locating and manipulating relevant neurons in the activation tensors . Built on the understanding of how an object is “drawn” by the generator, they further designed a semantic image editing system that can add, remove or change the appearance of an object in an input image . Concurrently, Frühstück et al. investigated the potential of activation tensor manipulation in image blending. Observing that boundary artifacts can be eliminated by by cropping and combining activation tensors at early layers of a generator, they proposed an algorithm to create large-scale texture maps of hundreds of megapixels by combining outputs of GANs trained on a lower resolution.

Overview

Our paper is structured as follows. First, we describe an extended version of the Image2StyleGAN embedding algorithm (See Sec. 4). We propose two novel modifications: 1) to enable local edits, we integrate various spatial masks into the optimization framework. Spatial masks enable embeddings of incomplete images with missing values and embeddings of images with approximate color values such as user scribbles. In addition to spatial masks, we explore layer masks that restrict the embedding into a set of selected layers. The early layers of StyleGAN encode content and the later layers control the style of the image. By restricting embeddings into a subset of layers we can better control what attributes of a given image are extracted. 2) to further improve the embedding quality, we optimize for an additional group of variables nn that control additive noise maps. These noise maps encode high frequency details and enable embedding with very high reconstruction quality.

Second, we explore multiple operations to directly manipulate activation tensors (See Sec. 5). We mainly explore spatial copying, channel-wise copying, and averaging,

Interesting applications can be built by combining multiple embedding steps and direct manipulation steps. As a stepping stone towards building interesting application, we describe in Sec. 6 common building blocks that consist of specific settings of the extended optimization algorithm.

Finally, in Sec. 7 we outline multiple applications enabled by Image2StyleGAN++: improved image reconstruction, image crossover, image inpainting, local edits using scribbles, local style transfer, and attribute level feature transfer.

An Extended Embedding Algorithm

We implement our embedding algorithm as a gradient-based optimization that iteratively updates an image starting from some initial latent code. The embedding is performed into two spaces using two groups of variables; the semantically meaningful W+W^{+} space and a Noise space NsN_{s} encoding high frequency details. The corresponding groups of variables we optimize for are w∈W+w\in W^{+} and n∈Nsn\in N_{s}. The inputs to the embedding algorithm are target RGB images xx and yy (they can also be the same image), and up to three spatial masks (MsM_{s}, MmM_{m}, and MpM_{p})

Algorithm 1 is the generic embedding algorithm used in the paper.

Our objective function consists of three different types of loss terms, i.e. the pixel-wise MSE loss, the perceptual loss , and the style loss .

Where MsM_{s}, MmM_{m} , MpM_{p} denote the spatial masks, ⊙\odot denotes the Hadamard product, GG is the StyleGAN generator, nn are the Noise space variables, ww are the W+W^{+} space variables, LstyleL_{style} denotes style loss from ‘conv3_3’‘conv3\_3’ layer of an ImageNet pretrained VGG-16 network , LperceptL_{percept} is the perceptual loss defined in Image2StyleGAN . Here, we use layers ‘conv1_1’‘conv1\_1’, ‘conv1_2’‘conv1\_2’, ‘conv2_2’‘conv2\_2’ and ‘conv3_3’‘conv3\_3’ of VGG-16 for the perceptual loss. Note that the perceptual loss is computed for four layers of the VGG network. Therefore, MpM_{p} needs to be downsampled to match the resolutions of the corresponding VGG-16 layers in the computation of the loss function.

2 Optimization Strategies

Optimization of the variables w∈W+w\in W^{+} and n∈Nsn\in N_{s} is not a trivial task. Since only w∈W+w\in W^{+} encodes semantically meaningful information, we need to ensure that as much information as possible is encoded in ww and only high frequency details in the Noise space.

The first possible approach is the joint optimization of both groups of variables ww and nn. Fig.2 (b) shows the result using the perceptual and the pixel-wise MSE loss. We can observe that many details are lost and were replaced with high frequency image artifacts. This is due to the fact that the perceptual loss is incompatible with optimizing noise maps. Therefore, a second approach is to use pixel-wise MSE loss only (see Fig. 2 (c)). Although the reconstruction is almost perfect, the representation (w,n)(w,n) is not suitable for image editing tasks. In Fig. 2 (d), we show that too much of the image information is stored in the noise layer, by resampling the noise variables nn. We would expect to obtain another very good, but slightly noisy embedding. Instead, we obtain a very low quality embedding. Also, we show the result of jointly optimizing the variables and using perceptual and pixel-wise MSE loss for ww variables and pixel-wise MSE loss for the noise variable. Fig. 2 (e) shows the reconstructed image is not of high perceptual quality. The PSNR score decreases to 33.3 dB. We also tested these optimizations on other images. Based on our results, we do not recommend using joint optimization.

The second strategy is an alternating optimization of the variables ww and nn. In Fig. 3, we show the result of optimizing ww while keeping nn fixed and subsequently optimizing nn while keeping ww fixed. In this way, most of the information is encoded in ww which leads to a semantically meaningful embedding. Performing another iteration of optimizing ww (Fig. 3 (d)) reveals a smoothing effect on the image and the PSNR reduces from 39.5 dB to 20 dB. Subsequent Noise space optimization does not improve PSNR of the images. Hence, repetitive alternating optimization does not improve the quality of the image further. In summary, we recommend to use alternating optimization, but each set of variables is only optimized once. First we optimize ww, then nn.

Activation Tensor Manipulations

Due to the progressive architecture of StyleGAN, one can perform meaningful tensor operations at different layers of the network . We consider the following editing operations: spatial copying, averaging, and channel-wise copying. We define activation tensor AlIA_{l}^{I} as the output of the ll-th layer in the network initialized with variables (w,n)(w,n) of the embedded image II. They are stored as tensors AlI∈RWl×Hl×ClA_{l}^{I}\in R^{W_{l}\times H_{l}\times C_{l}}. Given two such tensors AlIA_{l}^{I} and BlIB_{l}^{I}, copying replaces high-dimensional pixels ∈R1×1×Cl\in R^{1\times 1\times C_{l}} in AlIA_{l}^{I} by copying from BlIB_{l}^{I}. Averaging forms a linear combination λAlI+(1−λ)BlI\lambda A_{l}^{I}+(1-\lambda)B_{l}^{I}. Channel-wise copying creates a new tensor by copying selected channels from AlIA_{l}^{I} and the remaining channels from BlIB_{l}^{I}. In our tests we found that spatial copying works a bit better than averaging and channel-wise copying.

Frequently Used Building Blocks

We identify four fundamental building blocks that are used in multiple applications described in Sec. 7. While terms of the loss function can be controlled by spatial masks (Ms,Mm,MpM_{s},M_{m},M_{p}), we also use binary masks wmw_{m} and nmn_{m} to indicate what subset of variables should be optimized during an optimization process. For example, we might set wmw_{m} to only update the ww variables corresponding to the first kk layers. In general, wmw_{m} and nmn_{m} contain 11s for variables that should be updated and s for variables that should remain constant. In addition to the listed parameters, all building blocks need initial variable values winiw_{ini} and ninin_{ini}. For all experiments, we use a 32GB Nvidia V100 GPU.

Masked W+W^{+} optimization (WlW_{l}): This function optimizes w∈W+w\in W^{+}, leaving nn constant. We use the following parameters in the loss function (L) Eq. 1: λs=0\lambda_{s}=0, λmse1=10−5\lambda_{mse_{1}}=10^{-5}, λmse2=0\lambda_{mse_{2}}=0, λp=10−5\lambda_{p}=10^{-5}. We denote the function as:

where wmw_{m} is a mask for W+W^{+} space. We either use Adam with learning rate 0.01 or gradient descent with learning rate 0.8, depending on the application. Some common settings for Adam are: β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999, and ϵ=1e−8\epsilon=1e^{-8}. In Sec. 7, we use Adam unless specified.

For this optimization, we use Adam with learning rate 5, β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999, and ϵ=1e−8\epsilon=1e^{-8}. Note that the learning rate is very high.

Masked Style Transfer(MstM_{st}): This function optimizes ww to achieve a given target style defined by style image yy. We set following parameters in the loss function (L) Eq. 1: λs=5×10−7\lambda_{s}=5\times 10^{-7}, λmse1=0\lambda_{mse_{1}}=0, λmse2=0\lambda_{mse_{2}}=0, λp=0\lambda_{p}=0. We denote the function as:

where ww is the whole W+W^{+} space. For this optimization, we use Adam with learning rate 0.01, β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999, and ϵ=1e−8\epsilon=1e^{-8}.

Masked activation tensor operation (IattI_{att}): This function describes an activation tensor operation. Here, we represent the generator G(w,n,t)G(w,n,t) as a function of W+W^{+} space variable ww, Noise space variable nn, and input tensor tt. The operation is represented by:

where AlI1A_{l}^{I_{1}} and BlI2B_{l}^{I_{2}} are the activations corresponding to images I1I_{1} and I2I_{2} at layer ll, and M1M_{1} and M2M_{2} are the masks downsampled using nearest neighbour interpolation to match the Hl×WlH_{l}\times W_{l} resolution of the activation tensors.

Applications

In the following we describe various applications enabled by our framework.

As shown in Fig. 4, any image can be embedded by optimizing for variables w∈W+w\in W^{+} and n∈Nsn\in N_{s}. Here we describe the details of this embedding (See Alg. 2). First, we initialize: winiw_{ini} is a mean face latent code or random code sampled from UU depending on whether the embedding image is a face or a non-face, and ninin_{ini} is sampled from a standard normal distribution N(0,I)N(0,I) . Second, we apply masked W+W^{+} optimization (WlW_{l}) without using spatial masks or masking variables. That means all masks are set to 11. ImI_{m} is the target image we try to reconstruct. Third, we perform masked noise optimization (Mkn{Mk}_{n}), again without making use of masks. The images reconstructed are of high fidelity. The PNSR score range of 39 to 45 dB provides an insight of how expressive the Noise space in StyleGAN is. Unlike the W+W^{+} space, the Noise space is used for spatial reconstruction of high frequency features. We use 5000 iterations of WlW_{l} and 3000 iterations of Mkn{Mk}_{n} to get PSNR scores of 44 to 45 dB. Additional iterations did not improve the results in our tests.

2 Image Crossover

We define the image crossover operation as copying parts from a source image yy into a target image xx and blending the boundaries. As initialization, we embed the target image xx to obtain the W+W^{+} code w∗w^{*}. We then perform masked W+W^{+} optimization (WlW_{l}) with blurred masks MblurM_{blur} to embed the regions in xx and yy that contribute to the final image. Blurred masks are obtained by convolution of the binary mask with a Gaussian filter of suitable size. Then, we perform noise optimization. Details are provided in Alg. 3.

Other notations are the same as described in Sec 7.1. Fig. 5 and Fig. 1 show example results. We deduce that the reconstruction quality of the images is quite high. For the experiments, we use 1000 iterations in the function masked W+W^{+} optimization and 1000 iterations in Mkn{Mk}_{n}.

3 Image Inpainting

In order to perform a semantically meaningful inpainting, we embed into the early layers of the W+W^{+} space to predict the missing content and in the later layers to maintain color consistency. We define the image xx as a defective image (IdefI_{def}). Also, we use the mask wmw_{m} where the value is 1 corresponding to the first 9 (1 to 9), 17th17^{th} and 18th18^{th} layer of W+W^{+}. As an initialization, we set winiw_{ini} to the mean face latent code . We consider MM as the mask describing the defective region. Using these parameters, we perform the masked W+W^{+} optimization WlW_{l}. Then we perform the masked noise optimization Mkn{Mk}_{n} using Mblur+M_{blur+} which is the slightly larger blurred mask used for blending. Here λmse2\lambda_{mse_{2}} is taken to be 10−410^{-4}. Other notations are the same as described in Sec 7.1. Alg. 4 shows the details of the algorithm. We perform 200 steps of gradient descent optimizer for masked W+W^{+} optimization WlW_{l} and 1000 iterations of masked noise optimization Mkn{Mk}_{n}. Fig.6 shows example inpainting results. The results are comparable with the current state of the art, partial convolution . The partial convolution method frequently suffers from regular artifacts (see Fig.6 (third column)). These artifacts are not present in our method. In Fig.7 we show different inpainting solutions for the same image achieved by using different initializations of winiw_{ini} , which is an offset to mean face latent code sampled independently from a uniform distribution U[−0.4,0.4]U[-0.4,0.4]. The initialization mainly affects layers 10 to 16 that are not altered during optimization. Multiple inpainting solutions cannot be computed with existing state-of-the-art methods.

4 Local Edits using Scribbles

Another application is performing semantic local edits guided by user scribbles. We show that simple scribbles can be converted to photo-realistic edits by embedding into the first 4 to 6 layers of W+W^{+} (See Fig.8). This enables us to do local edits without training a network. We define an image xx as a scribble image (IscrI_{scr}). Here, we also use the mask wmw_{m} where the value is 1 corresponding to the first 4,5 or 6 layers of the W+W^{+} space. As initialization, we set the winiw_{ini} to w∗w^{*} which is the W+W^{+} code of the image without scribble. We perform masked W+W^{+} optimization using these parameters. Then we perform masked noise optimization Mkn{Mk}_{n} using MblurM_{blur}. Other notations are the same as described in Sec 7.1. Alg. 5 shows the details of the algorithm. We perform 1000 iterations using Adam with a learning rate of 0.1 of masked W+W^{+} optimization WlW_{l} and then 1000 steps of masked noise optimization Mkn{Mk}_{n} to output the final image.

5 Local Style Transfer

Local style transfer modifies a region in the input image xx to transform it to the style defined by a style reference image. First, we embed the image in W+W^{+} space to obtain the code w∗w^{*}. Then we apply the masked W+W^{+} optimization WlW_{l} along with masked style transfer MstM_{st} using blurred mask MblurM_{blur}. Finally, we perform the masked noise optimization Mkn{Mk}_{n} to output the final image. Alg. 6 shows the details of the algorithm. Results for the application are shown in Fig.9. We perform 1000 steps to obtain of WlW_{l} along with MstM_{st} and then perform 1000 iterations of Mkn{Mk}_{n}.

6 Attribute level feature transfer

We extend our work to another application using tensor operations on the images embedded in W+W^{+} space. In this application we perform the tensor manipulation corresponding to the tensors at the output of the 4th4^{th} layer of StyleGAN. We feed the generator with the latent codes (ww, nn) of two images I1I_{1} and I2I_{2} and store the output of the fourth layer as intermediate activation tensors AlI1A_{l}^{I_{1}} and BlI2B_{l}^{I_{2}}. A mask MsM_{s} specifies which values to copy from AlI1A_{l}^{I_{1}} and which to copy from BlI2B_{l}^{I_{2}}. The operation can be denoted by Iatt(Ms,Ms,w,nini,4)I_{att}(M_{s},M_{s},w,n_{ini},4). In Fig.10, we show results of the operation. A design parameter of this application is what style code to use for the remaining layers. In the shown example, the first image is chosen to provide the style. Notice, in column 2 of Fig.10, in-spite of the different alignment of the two faces and objects, the images are blended well. We also show results of blending for the LSUN-car and LSUN-bedroom datasets. Hence, unlike global edits like image morphing, style transfer, and expression transfer , here different parts of the image can be edited independently and the edits are localized. Moreover, along with other edits, we show a video in the supplementary material that further shows that other semantic edits e.g. masked image morphing can be performed on such images by linear interpolation of W+W^{+} code of one image at a time.

Conclusion

We proposed Image2StyleGAN++, a powerful image editing framework built on the recent Image2StyleGAN. Our framework is motivated by three key insights: first, high frequency image features are captured by the additive noise maps used in StyleGAN, which helps to improve the quality of reconstructed images; second, local edits are enabled by including masks in the embedding algorithm, which greatly increases the capability of the proposed framework; third, a variety of applications can be created by combining embedding with activation tensor manipulation. From the high quality results presented in this paper, it can be concluded that our Image2StyleGAN++ is a promising framework for general image editing. For future work, in addition to static images, we aim to extend our framework to process and edit videos.

Acknowledgement This work was supported by the KAUST Office of Sponsored Research (OSR) under Award No. OSR-CRG2018-3730.

References

Additional Results

To evaluate the results quantitatively, we use three standard metrics, SSIM, MSE loss and PSNR score to compare our method with the state-of-the-art Partial Convolution and Gated Convolution methods.

As different methods produce outputs at different resolutions, we bi-linearly interpolate the output images to test the methods at three resolutions 1024×10241024\times 1024, 512×512512\times 512 and 256×256256\times 256 respectively. We use 77 masks (Fig. 11) and 1010 ground truth images (Fig. 12) to create 1010 defective images (i.e. images with missing regions) for the evaluation. These masks and images are chosen to make the inpainting a challenging task: i) the masks are selected to contain very large missing regions, up to half of an image; ii) the ground truth images are selected to be of high variety that cover different genders, ages, races, etc.

Table 1 shows the quantitative comparison results. It can be observed that our method outperforms both Partial Convolution and Gated Convolution across all the metrics. More importantly, the advantages of our method can be easily verified by visual inspection. As Fig. 13 and Fig. 14 show, although previous methods (e.g. Partial convolution) perform well when the missing region is small, both of them struggle when the missing region covers a significant area (e.g. half) of the image. Specifically, Partial Convolution fails when the mask covers half of the input image (Fig. 13); due to the relatively small resolution (256×256256\times 256) model, Gated Convolution can fill in the details of large missing regions, but of much lower quality compared to the proposed method (Fig. 14).

In addition, our method is flexible and can generate different inpainting results (Fig. 15), which cannot be fulfilled by any of the above-mentioned methods. All our inpainting results are of high perceptual quality.

Although better than the two state-of-the-art methods, our inpainting results still leave room for improvement. For example in Fig. 13, the lighting condition (first row), age (second row) and skin color (third and last row) are not learnt that well. We propose to address them in the future work.

2 Image Crossover

To further evaluate the expressibility of the Noise space, we show additional results on image crossover in Fig. 16. We show that the space is able to crossover parts of images from different races (see second and third column).

3 Local Edits using Scribbles

In order to evaluate the quality of the local edits using scribbles, we evaluate the face attribute scores on edited images. We perform some common edits of adding baldness, adding a beard, smoothing wrinkles and adding a moustache on the face images to evaluate how photo-realistic the edited images are. Table 2 shows the average change in the confidence of the classifier after a particular edit is performed. We also show additional results of the Local edits in Fig. 17. For our method, one remaining challenge is that sometimes the edited region is overly smooth (e.g. first row).

4 Attribute Level Feature Transfer

We show a video in which attribute interpolation can be performed on the base image by copying the content from an attribute image. Here different attributes can be taken from different images embedded in the W+W^{+} space and applied to the base image. These attributes can be independently interpolated and the results show that the blending quality of the framework is quite high. We also show additional results on LSUN Cars and LSUN Bedrooms in the video (also see Fig. 18). Notice that in the LSUN bedrooms, for instance, the style and the position of the beds can be customized without changing the room layout.

In order to evaluate the perceptual quality of attribute level feature transfer, we compute perceptual length between the images produced by independently interpolated attributes (called masked interpolation). StyleGAN showed that the metric evaluates how perceptually smooth the transitions are. Here, perceptual length measures the changes produced by feature transfer which may be affected especially by the boundary of the blending. The boundary may tend to produce additional artifacts or introduce additional features which is clearly undesirable.

We compute the perceptual length across 1000 samples using two masks shown in Fig. 11 (First and Seventh column). In Table 3 we show the results of the computation of the perceptual length (both for masked and non-masked interpolation) on FFHQ, LSUN Cars and LSUN Bedrooms pretrained StyleGAN. We compare these scores as the non-masked interpolation gives us the upper bound of the perceptual length for a model (in this case there is no constraint on what features of the face should change). As a particular area of the image is interpolated rather than the whole image, note that our results on FFHQ pretrained StyleGAN produce lower score than the non-masked interpolation. The low perceptual length score suggests that there is a less drastic change. Hence, we conclude that the output images have comparable perceptual quality with non-masked interpolation.

LSUN Cars and LSUN Bedrooms produce relatively higher perceptual length score. We attribute this result to the fact that the images in these datasets can translate and the position of the features is not fixed. Hence, the two images produced at random might have different orientation in which case the blending does not work as good.

5 Channel wise feature average

We perform another operation denoted by Iatt(1,0,wx,,nini,6)I_{att}(1,0,w_{x},,n_{ini},6), where wxw_{x} can be the W+W^{+} code for images I1I_{1} or I2I_{2}. In Fig. 19, we show the result of this operation which is initialized with two different W+W^{+} codes. The resulting faces contain the characteristics of both faces and the styles are modulated by the input W+W^{+} codes.