GAN Inversion for Out-of-Range Images with Geometric Transformations

Kyoungkook Kang, Seongtae Kim, Sunghyun Cho

Introduction

Generative adversarial networks (GANs) are generative models that can synthesize realistic-looking images . Typically, GANs learn a mapping function from a random noise vector sampled from a pre-defined distribution to a realistic-looking image through the adversarial training of a generator and a discriminator. For the past several years, a significant progress has been made to improve the quality and diversity of synthesized images . As a result, recent GAN models such as StyleGAN , StyleGAN2 , and BigGAN can produce extremely high-quality images of high resolution.

Recently, it has been shown that rich semantic information is encoded in the intermediate features and the latent space of GANs, and furthermore, that images can be effectively edited in a semantically meaningful way by modifying features or latent code . To enable such semantic editing for real images, GAN inversion has attracted much attention lately . GAN inversion maps a real image into the latent space of a pre-trained GAN model. Once an inverted latent code is obtained, the image can be semantically edited by modifying its latent code or intermediate features generated from the code.

For successful semantic editing of real images, it is critical to find an in-domain latent code that aligns with the domain of a pre-trained GAN model . As shown in , there may exist more than one latent codes that can reconstruct a given input image, and some of them may be out of the domain. The semantic knowledge encoded in the latent space does not apply for such out-of-domain codes, thus semantic editing of such codes fails to produce proper results.

Unfortunately, such in-domain latent codes can be found only for a small fraction of real images that align with the training images of a pre-trained GAN model. For example, most GAN models use geometrically aligned face images as their training data for ease of training. As a result, images with a small amount of translation or other geometric transformations are out of their ranges, and the previous GAN inversion methods cannot find in-domain latent codes for such out-of-range images. This severely limits the applicability of semantic editing of real images using GAN inversion. Fig. 1 shows real-world examples. The input images in (a) and (d) are random images downloaded from internet. As they are out-of-range with different rotation, scaling and translation with respect to the training dataset (FFHQ ), directly applying a previous GAN inversion method produces unacceptable results as shown in (b) and (e).

One solution would be to align a target image before GAN inversion, but accurate alignment of an image to the training data can be difficult or even impossible especially in the case of arbitrary natural images. For example, for the image in Fig. 1(d), a face alignment method completely fails due to severe cropping.

In this paper, we propose a novel GAN inversion approach to semantic editing of out-of-range images, which is dubbed Base-Detail Invert (BDInvert). BDInvert inverts a geometrically unaligned image with the training images for StyleGAN and StyleGAN2 . Specifically, BDInvert is designed to cover geometric transformations such as translation, rotation, and scaling, and supports various types of editing for out-of-range images that are not supported by previous approaches.

Our key idea is as follows. It is impossible to invert an out-of-range image to an in-domain latent code in the original latent space of a pre-trained GAN model. Instead, we propose to invert an image into another space that we refer to as the F/W+\mathcal{F}/\mathcal{W}^{+}, which consists of two subspaces F\mathcal{F} and W+\mathcal{W}^{+}. The base code space F\mathcal{F} encodes geometric transformations and also supports diverse local variations that enable more faithful reconstruction of an input image. On the other hand, the detail code space W+\mathcal{W}^{+} is independent of geometric transformations and supports semantic manipulations.

To find a latent code in the F/W+\mathcal{F}/\mathcal{W}^{+} space that faithfully reconstructs an input image, we adopt an optimization-based approach. However, naïve optimization of a reconstruction loss does not guarantee a latent code that supports semantic editing. To enable semantic editing, we also propose a regularization approach based on an encoder network. Fig. 1(c) and (f) show our reconstruction and editing results of real-world images. Thanks to our F/W+\mathcal{F}/\mathcal{W}^{+} space and inversion approach, we can successfully reconstruct and edit the out-of-range real-world input images.

Our main contributions can be summarized as follows.

We propose BDInvert, a novel GAN inversion approach to semantic editing of real images with geometric transformations that are not aligned with the training images of a pre-trained GAN model.

BDInvert projects an image into an alternative latent space F/W+\mathcal{F}/\mathcal{W}^{+} that supports more faithful reconstruction and semantic editing of out-of-range images with geometric transformations and diverse local variations.

We propose a novel regularization method to find a proper solution in the F/W+\mathcal{F}/\mathcal{W}^{+} space that supports semantic image editing.

Related Work

In order to embed real images into the latent space of GANs, various approaches have been proposed in two directions. One direction is to train an encoder using a data-driven approach . The other direction is to initialize a latent vector randomly or to the output of a pre-trained encoder, and then to optimize it to reconstruct a target image . However, inverting a real image remains a difficult problem because of the limited expressiveness of the latent space of GANs.

Recently, in order to enhance the inversion quality, several attempts to widen the latent space have been made . Gu et al. improved the reconstruction quality by mixing features from several latent codes. Pan et al. fine-tune a generator on-the-fly for more faithful reconstruction. Huh et al. find geometric transformation parameters to transform an image region to be more suitable for BigGAN inversion. Meanwhile, Abdal et al. showed high-quality embedding results for StyleGAN using an extended latent space W+\mathcal{W}^{+}. Afterwards, many studies focusing on StyleGAN have been proposed . Abdal et al. and Karras et al. optimize the noise channel for more accurate embedding. For successful image editing, embedding an image into GAN’s domain is essential. To this end, Zhu et al. train an encoder that projects an image into StyleGAN’s domain, and optimize a latent code with the guidance of the encoder. Tewari et al. introduced a hierarchical optimization that first embeds an image into the W\mathcal{W} space and then embeds it into the W+\mathcal{W}^{+} space for better editing. Zhu et al. proposed the P−norm+\mathcal{P}-norm^{+} space for in-domain inversion. However, most existing works cannot handle out-of-range images.

A widely used approach to semantic image editing using GAN is to modify a latent code along semantically meaningful directions. Härkönen et al. identify semantic directions by applying the principal component analysis (PCA) on sampled latent codes. Shen et al. use attribute classifiers to discover semantic directions. Shen and Zhou proposed an unsupervised method that factorizes the weights of latent code transformation layers to find semantic directions that cause large changes to the output.

\mathcal{F}/\mathcal{W}^{+} In this section, we first review state-of-the-art GAN inversion approaches and discuss their limitations on out-of-range images. Then, we introduce an alternative latent space F/W+\mathcal{F}/\mathcal{W}^{+} to overcome the limitations.

Our approach is based on StyleGAN and StyleGAN2 , which produce high-quality synthesis results. Both GAN frameworks use a mapping network f:Z→Wf:\mathcal{Z}\rightarrow\mathcal{W} based on a multi-layer perceptron (MLP) that maps a latent code z∈Z\mathbf{z}\in\mathcal{Z} to an intermediate latent code w∈W\mathbf{w}\in\mathcal{W} as shown in Fig. 2(a). Compared to the latent space Z\mathcal{Z}, the intermediate latent space W\mathcal{W} provides less entangled representations of different attributes so that different attributes can be more easily adjusted in the image generation process. Another noticeable feature of StyleGAN and StyleGAN2 is their multi-scale image synthesis approaches, which enable scale-wise disentanglement of different attributes. To control the generation process in a multi-scale manner, both StyleGAN and StyleGAN2 feed the intermediate latent code w\mathbf{w} to multiple layers of different scales of the generator. In addition, to enhance the diversity of synthesized images, both StyleGAN and StyleGAN2 utilize noise randomly sampled from a Gaussian distribution for each image generation.

Later, Zhu et al. showed that, for semantic image manipulation, it is essential to find an in-domain latent code instead of a latent code that precisely reconstructs an input image. They also showed that real images can be effectively inverted to an in-domain latent code in W+\mathcal{W}^{+} with a domain-guided encoder and domain-regularized optimization.

Nonetheless, GAN inversion to the extended latent space W+\mathcal{W}^{+} still fails to find an in-domain latent code for out-of-range images as discussed in Sec. 1. To overcome this limitation, we propose another latent space F/W+\mathcal{F}/\mathcal{W}^{+}. Each element w∗\mathbf{w}^{*} in F/W+\mathcal{F}/\mathcal{W}^{+} is defined as w∗=(f,wM+)\mathbf{w}^{*}=(\mathbf{f},\mathbf{w}_{M+}) where f\mathbf{f} is a base code and wM+\mathbf{w}_{M+} is a detail code. wM+\mathbf{w}_{M+} is a set of latent codes for the fine scales of the generator, which is defined as wM+={wM,⋯ ,wN}\mathbf{w}_{M+}=\{\mathbf{w}_{M},\cdots,\mathbf{w}_{N}\}. f\mathbf{f} is a coarse-scale feature map of the generator before the layer that takes wM\mathbf{w}_{M}. Specifically, for StyleGAN , we define f\mathbf{f} as the feature map right before the first adaptive instance normalization (AdaIN) layer at a certain scale. For StyleGAN2 , we define f\mathbf{f} as the feature map after a pair of upsampling and convolution layers at a certain scale. Fig. 2(b) and (c) depict the latent space F/W+\mathcal{F}/\mathcal{W}^{+} of StyleGAN and StyleGAN2, respectively. In our experiments, we test two different scales, 8×88\times 8 and 16×1616\times 16, for f\mathbf{f}.

In the case of StyleGAN2 , the generator needs a feature map corresponding to an RGB image upsampled from the previous scale (Fig. 2(c)). While we may include a small-scale feature map as a part of our latent space, we observed that the feature maps at the coarse scales have values close to zero and have little impact on image generation results. Thus, we simply set them to zero in our experiments as depicted by the green box in Fig. 2(c).

The F/W+\mathcal{F}/\mathcal{W}^{+} space provides a couple of nice properties that enable semantic editing of out-of-range images. First, compared to {w1,⋯ ,wM−1}\{\mathbf{w}_{1},\cdots,\mathbf{w}_{M-1}\}, the base code f\mathbf{f} can represent a wider range of images including images with geometric transformations. For example, as f\mathbf{f} is a feature map of a convolutional neural network (CNN), we can simply shift f\mathbf{f} along the xx- or yy-axis to represent the feature map of a shifted image. Second, the detail code wM+\mathbf{w}_{M+} is invariant to translations of images. Specifically, in the case of StyleGAN , wM+\mathbf{w}_{M+} controls the parameters of the AdaIN layers of the generator. Similarly, in the case of StyleGAN2 , wM+\mathbf{w}_{M+} controls the parameters of the demodulation layers. Both AdaIN and demodulation operations are global operations that are applied to CNN features in a translation-invariant manner.

Thanks to the aforementioned properties, we can describe the relationship between an image II and its transformed image T(I)T(I) where TT is a geometric transformation operator as follows. Suppose that II is generated from w∗\mathbf{w}^{*}, i.e., I=G(w∗)=G(f,wM+)I=G(\mathbf{w}^{*})=G(\mathbf{f},\mathbf{w}_{M+}) where GG is the generator of a pre-trained GAN model. Then, T(I)T(I) can be expressed as:

where T′T^{\prime} is a geometric transformation operator corresponding to TT whose scale is adjusted according to the relative scale of f\mathbf{f} to II. This relationship can be also used for semantic image manipulation of T(I)T(I). As T′(f)T^{\prime}(\mathbf{f}) is a CNN feature map and wM+\mathbf{w}_{M+} is a set of parameters for global operations, for editing T(I)T(I), we can manipulate wM+\mathbf{w}_{M+} in the same way for II and achieve similar editing results.

Fig. 3 shows an example that illustrates the relationship in Eq. (1). In this example, we sample an in-domain latent code (f,wM+)(\mathbf{f},\mathbf{w}_{M+}) and generate an in-range image in Fig. 3(a) using StyleGAN2 . Shifting f\mathbf{f}, we can generate a shifted image of Fig. 3(a) as shown in Fig. 3(b). While they are not exactly the same due to the zero padding and noise component in StyleGAN2, they look almost identical proving the relationship in Eq. (1). Fig. 3(c) and (d) show the semantic editing results of (a) and (b) using the same manipulated latent code wM+′\mathbf{w}^{\prime}_{M+}. The results show that we can effectively perform semantic editing for geometrically transformed images in the same way as for in-range images.

The discussion above shows that, as long as (f,wM+)(\mathbf{f},\mathbf{w}_{M+}) is in-domain, (T′(f),wM+)(T^{\prime}(\mathbf{f}),\mathbf{w}_{M+}) for an arbitrary T′T^{\prime} also supports semantic image editing. Based on this, we define an extended domain of w∗\mathbf{w}^{*} as a set of geometrically transformed latent codes (T′(f),wM+)(T^{\prime}(\mathbf{f}),\mathbf{w}_{M+}) of in-domain latent codes (f,wM+)(\mathbf{f},\mathbf{w}_{M+}) for arbitrary transformations T′T^{\prime}.

While the discussion above discusses only geometric transformations, we note that our latent space F/W+\mathcal{F}/\mathcal{W}^{+} supports not only geometric transformations but also diverse local variations as the base code f\mathbf{f} supports locally different information. This leads to more faithful reconstruction even for images without geometric transformations as will be shown in Sec. 5. We also note that the latent space F/W+\mathcal{F}/\mathcal{W}^{+} does not support semantic editing that require coarse-scale wi\mathbf{w}_{i}’s such that i<Mi<M. However, our experiments show that it still supports various types of semantic editing as we define f\mathbf{f} as a very coarse-scale feature map.

\mathcal{F}/\mathcal{W}^{+} For inversion of an image, we adopt the optimization-based approach since it generally achieves higher reconstruction quality compared to the encoder-based approach . In this section, we introduce our optimization approach both for StyleGAN and StyleGAN2 .

Given an input image II, to find a latent code w∗\textbf{w}^{*} that reconstructs II, we optimize an objective function with a reconstruction loss LreconL_{recon}, which is defined as:

where LMSEL_{MSE} and LperL_{per} are mean-squared-error (MSE) and perceptual losses, respectively. ωper\omega_{per} is a weight for LperL_{per}. LMSEL_{MSE} is defined as LMSE(w∗)=∥I−G(w∗)∥2L_{MSE}(\textbf{w}^{*})=\|I-G(\textbf{w}^{*})\|^{2} where GG is the generator of a pre-trained StyleGAN model. LperL_{per} is defined as Lper(w∗)=∥F(I)−F(G(w∗))∥2L_{per}(\textbf{w}^{*})=\|F(I)-F(G(\textbf{w}^{*}))\|^{2}, where FF is a LPIPS network to compute the perceptual distance .

By optimizing Eq. (2), e.g., using the Adam optimizer , we can obtain latent codes that produce high-quality reconstruction results even for out-of-range images thanks to the high expressive power of the F/W+\mathcal{F}/\mathcal{W}^{+} space. However, such latent codes do not support semantic image editing as they are out-of-domain. Fig. 4 shows an example using StyleGAN2 . Fig. 4(a) shows a target image, which is out-of-range due to translation. Fig. 4(e) is a reference image for style mixing, which is a semantic image editing operation . Optimizing the reconstruction loss in Eq. (2), we can obtain a latent code that accurately reconstructs the target image as shown in Fig. 4(b). However, the estimated latent code is out-of-domain, so it fails to produce an appropriate style mixing result as shown in Fig. 4(f).

To enable semantic editing of out-of-range images, both f\mathbf{f} and wM+\mathbf{w}_{M+} must be in proper domains. To guide our optimization process to a solution in a proper domain, we adopt regularization both on f\mathbf{f} and wM+\mathbf{w}_{M+}. The following subsections discuss our regularization schemes one by one.

\mathbf{w}_{M+} To promote in-domain wM+\mathbf{w}_{M+}, we adopt the P−norm+\mathcal{P}-norm^{+} space-based regularization scheme proposed by Zhu et al. . Specifically, at each iteration of the iterative optimization of our objective function, we transform the current estimate of wM+\mathbf{w}_{M+} into the P−norm+\mathcal{P}-norm^{+} space. Then, we clip the values that are out of a certain range. In our experiments, we used the range [−5σ,5σ][-5\sigma,5\sigma] as suggested in where σ\sigma is the standard deviation of in-domain latent codes. We then transform the clipped values back into the W+\mathcal{W}^{+} space. We refer the readers to for more details.

3 Regularization on Base Code 𝐟𝐟\mathbf{f}

While optimizing Eq. (2) with the regularization on wM+\mathbf{w}_{M+} results in an in-domain solution for wM+\mathbf{w}_{M+}, it still produces an improper solution for f\mathbf{f} that results in the failure of semantic image editing. Fig. 4(c) shows an inversion result using the reconstruction loss with the regularization on wM+\mathbf{w}_{M+}. Thanks to the hard clipping in the P−norm+\mathcal{P}-norm^{+} space, the estimated wM+\mathbf{w}_{M+} is always in a desired range. However, the estimated f\mathbf{f} is still out-of-domain, and produces an incorrect style mixing result in Fig. 4(g).

To overcome this, we introduce a regularization method that encourages f\mathbf{f} to be in the extended domain of f\mathbf{f} defined in Sec. 3. Our method is a two-step approach. For an input image II, we first find an initial base code fo\mathbf{f}^{o} that lies in the extended domain of f\mathbf{f} using an encoder EE. Then, while optimizing Eq. (2), we find a base code f\mathbf{f} that is close to fo\mathbf{f}^{o}. To achieve this, we define a regularization loss for f\mathbf{f} as:

Our final objective function is then defined as:

where ωf\omega_{\mathbf{f}} is a weight for the regularization loss LfL_{\mathbf{f}}. Our final approach optimizes Eq. (4) with the regularization on wM+\mathbf{w}_{M+}. Fig. 4(d) and (h) show that our final approach can successfully invert an out-of-range image and support semantic image editing, respectively.

4 Encoder for Base Code 𝐟𝐟\mathbf{f}

Our encoder estimates an initial base code fo\mathbf{f}^{o} of an input image. As fo\mathbf{f}^{o} has a small spatial resolution, e.g., 16×1616\times 16, the encoder does not require an input image of the original resolution or a heavy network architecture. Thus, the encoder is designed to take a downsampled image of the resolution 8×8\times larger than f\mathbf{f}, e.g., 128×128128\times 128. The encoder has a VGG-like architecture consisting of 11 convolution blocks and three pooling layers without fully connected layers. More details can be found in the supplementary material.

For the training of the encoder, we randomly sample a batch of latent codes from the latent space Z\mathcal{Z} at each iteration. From each sampled latent code z\mathbf{z}, we obtain its corresponding latent code (fgt,wM+gt)(\mathbf{f}^{gt},\mathbf{w}_{M+}^{gt}) and its image II. Using the sampled latent codes and their images, we train our encoder with a loss function defined as:

where I↓I_{\downarrow} is a downsampled version of II. The first and second terms on the right-hand side are a MSE loss and a perceptual loss. The loss minimizes the difference between the training image II and its reconstructed image using the latent code obtained by the encoder.

As we have fgt\mathbf{f}^{gt}, we may use a loss term based on the distance between E(I↓)E(I_{\downarrow}) and fgt\mathbf{f}^{gt}, e.g, ∥E(I↓)−fgt∥2\|E(I_{\downarrow})-\mathbf{f}^{gt}\|^{2}. However, we found that using it instead of the loss terms in Eq. (4.4) leads to less accurate reconstruction of an input image.

Our training procedure does not use geometrically transformed images. Nevertheless, our encoder still performs effectively for geometrically transformed images thanks to the spatially-invariant property of CNNs. For example, for a shifted image, our encoder estimates a shifted feature map fo\mathbf{f}^{o} that lies in the extended domain of f\mathbf{f}.

Although Eq. (4.4) does not have any terms to encourage to predict a latent code in the extended domain, our encoder can effectively find a latent code that supports semantic image editing. As the encoder is trained using a large amount of images with a large batch size, we found that it is not necessary to include any other constraints such as the loss term based on the latent code distance.

Experiments

In our implementation, we downsample images to 256×256256\times 256 to compute the perceptual losses in LL and LencL_{enc} following previous works . In our experiments, we set ωper=10\omega_{per}=10, ωf=10\omega_{\mathbf{f}}=10 and λper=10\lambda_{per}=10. For training the encoder, we set the batch size to 1616 and the number of iterations to 10,000. We initially set the learning rate to 0.001 and reduced it by a factor of 0.1 every 2,000 iterations. For the inversion, we use 1,200 iterations with learning rate of 0.01. We use the Adam optimizer both for the training of the encoder and GAN inversion. We conducted our experiments using pre-trained models of StyleGANhttps://github.com/genforce/idinvert_pytorch and StyleGAN2https://github.com/genforce/genforce.

In our experiments, we implement semantic editing operations by adding a semantic editing vector to a latent code, i.e., wedit=w+αv\mathbf{w}^{edit}=\mathbf{w}+\alpha\mathbf{v} where α\alpha is a user parameter to control the editing strength and v\mathbf{v} is an editing vector following . Specifically, we use editing vectors provided by IDinvert and SeFa for StyleGAN and StyleGAN2 , respectively. For a latent code in F/W+\mathcal{F}/\mathcal{W}^{+}, we add an editing vector only to a detail code wM+\mathbf{w}_{M+}.

Reconstruction comparison

We first compare the reconstruction quality of our method with those of previous state-of-the-art inversion methods on the CelebA-HQ dataset using a StyleGAN2 model pre-trained on the FFHQ dataset . For the comparison, we constructed a test set composed of 50 images randomly extracted from the CelebA-HQ dataset . In order to investigate the inversion performance on out-of-range images with geometric transformations, we applied different transformations to the test set. Specifically, we applied translation of 50, 100, and 150 pixels in random directions, rotation by 10, 20, and 30 degrees randomly in a counterclockwise and clockwise direction, and scaling by 7/8, 3/4, 9/8, and 5/4.

We compare our method with state-of-the-art methods: Im2StyleGAN , StyleGAN2 inversion , P-norm+ , and PSP . PSP is an encoder-based method while the others are optimization-based ones. We used the authors’ code for StyleGAN2 and PSP. We implemented Im2StyleGAN and P-norm+ as their code is not available. We also compare two versions of our method, which use a base code f of size 8×88\times 8 and 16×1616\times 16, respectively.

Fig. 5 shows a qualitative comparison. As shown in the figure, all the methods except for Im2StyleGAN and ours fail to reconstruct the input images. Table 1 reports a quantitative comparison in PSNR and FID . We refer the readers to our supplementary material for additional comparison in SSIM and RMSE. The table shows that our 16×1616\times 16 version achieves the highest reconstruction quality both in PSNR and FID for all geometric transformations. Both in the figure and table, Im2StyleGAN shows high-quality reconstruction results. However, due to the lack of in-domain constraints, Im2StyleGAN tends to produce out-of-domain latent codes that are not semantically editable as will be seen later in this section. The table also shows that the performances of the previous methods degrade quickly for larger translations and rotations. For example, the performance of P-norm+ drops by 3.86 dB for the rotation by 30 degrees. Our 8×88\times 8 version performs worse than the 16×1616\times 16 version as it uses a more constrained latent space. We also note that our 16×1616\times 16 version outperforms all the other methods even for images without geometric transformations (Translation = 0 in Table 1) thanks to the base code f\mathbf{f} supporting local variations.

Inversion of natural images

Due to the large diversity of natural images, it is difficult to accurately reconstruct and edit a natural image using previous GAN inversion approaches. On the other hand, thanks to the high degree-of-freedom of the F/W+\mathcal{F}/\mathcal{W}^{+} space, our approach is especially effective in handling natural images. To verify this, we compare the reconstruction and editing quality of previous methods and ours on natural images. For evaluation, we use StyleGAN and StyleGAN2 models pre-trained on the LSUN bedroom, tower and cat datasets . We also collected 25 bedroom, tower and cat images each from the internet and used them as our test sets so that the images in the test sets are of the same classes as the training images of the pre-trained models, but not aligned with the training images. For these datasets, we use 3,000 iterations.

We compare our method against Im2StyleGAN , which shows high-quality reconstruction results in the previous experiment, and IDinvert , which finds an in-domain latent code for semantic editing. We use the authors’ code for IDinvert. Fig. 6 shows a qualitative comparison of the reconstruction and editing qualities. Both Im2StyleGAN and IDinvert produce less accurate reconstruction results than ours. Their editing results also show artifacts due to the out-of-range input images. Especially, the editing results of Im2StyleGAN have severe artifacts as its out-of-domain latent codes. In contrast, our method shows high-quality reconstruction and editing results for all three cases. Table 2 shows a quantitative comparison of the reconstruction qualities. The table also shows that our method achieves high reconstruction quality on natural images compared to the other methods. More results can be found in the supplementary material.

Ablation study

Fig. 7 shows a qualitative comparison of variants of our method using StyleGAN2 to verify the effectiveness of our regularization scheme. While all the variants show excellent reconstruction results thanks to the high degree of freedom of the F/W+\mathcal{F}/\mathcal{W}^{+} space, the editing results of the variants that use only the reconstruction loss or regularization on the detail code wM+\textbf{w}_{M+} are severely degraded. On the other hand, the editing result of our final model in (d) looks the most natural thanks to our regularized inversion scheme. More examples and a quantitative evaluation are in the supplementary material.

Editing operations v.s. scale of base code 𝐟𝐟\mathbf{f}

Finally, we analyze the effect of the scale of the base code f\mathbf{f} on image editing. Using a feature map at a finer-scale for the base code f\mathbf{f} leads to higher reconstruction quality as shown in Fig. 5 and Table 1. On the other hand, it also reduces the diversity of semantic editing operations. Especially, it makes it difficult to perform semantic operations that rely on coarse-scale latent codes wi\mathbf{w}_{i} in the W\mathcal{W} space. Fig. 8 shows an example. While our method with f\mathbf{f} of size 8×88\times 8 supports both pose changing and aging, ours with f\mathbf{f} of size 16×1616\times 16 does not support pose changing since the pose changing operation requires to edit small-scale latent codes.

Conclusion

In this paper, we proposed BDInvert, a novel GAN inversion approach for semantic editing of out-of-range images with geometric transformations. Based on the StyleGAN and StyleGAN2 frameworks , we presented an alternative latent space F/W+\mathcal{F}/\mathcal{W}^{+} that supports geometric transformations of an image as well as its semantic manipulation. To find a proper solution in the F/W+\mathcal{F}/\mathcal{W}^{+} space that is semantically editable, we introduced a novel regularized optimization approach. We verified the effectiveness of our approach both qualitatively and quantitatively.

As discussed in Secs. 3 and 5, the F/W+\mathcal{F}/\mathcal{W}^{+} space reduces the diversity of semantic editing operations. Also as our approach is based on optimization, it requires a relatively long computation time. With an Nvidia RTX 3090 GPU, it takes about 3 minutes for a 1024×10241024\times 1024-sized image. Our approach cannot handle images with severe geometric transformations. However, this can be easily resolved by rough alignment of an input image as our method does not require accurate alignment. Finally, our method cannot handle images that are too different from the training dataset. See the supplementary material for examples.

Acknowledgements

This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government(MSIT) (No.2019-0-01906, Artificial Intelligence Graduate School Program(POSTECH)) and National Research Foundation of Korea (NRF) grant funded by the Korea government(MSIT) (NRF-2018R1A5A1060031, No. 2020R1C1C1014863).

References