Pivotal Tuning for Latent-based Editing of Real Images
Daniel Roich, Ron Mokady, Amit H. Bermano, Daniel Cohen-Or
Introduction
In recent years, unconditional image synthesis has made huge progress with the emergence of Generative Adversarial Networks (GANs) . In essence, GANs learn the domain (or manifold) of the desired image set and produce new samples from the same distribution. In particular, StyleGAN is one of the most popular choices for this task. Not only does it achieve state-of-the-art visual fidelity and diversity, but it also demonstrates fantastic editing capabilities due to an organically formed disentangled latent space. Using this property, many methods demonstrate realistic editing abilities over StylGAN’s latent space , such as changing facial orientations, expressions, or age, by traversing the learned manifold.
While impressive, these edits are performed strictly in the generator’s latent space, and cannot be applied to real images that are out of its domain. Hence, editing a real image starts with finding its latent representation. This process, called GAN inversion, has recently drawn considerable attention . Early attempts inverted the image to — StyleGAN’s native latent space. However, Abdal et al. have shown that inverting real images to this space results in distortion, i.e. a dissimilarity between the given and generated images, causing artifacts such as identity loss, or an unnatural appearance. Therefore, current inversion methods employ an extended latent space, often denoted as , which is more expressive and induces significantly less distortion .
However, even though employing codes from potentially produces great visual quality even for out-of-domain images, these codes suffer from weaker editability, since they are not from the generator’s trained domain. Tov et al. define this conflict as the distortion-editability tradeoff, and show that the closer the codes are to , the better their editability is. Indeed, recent works suggest a compromise between edibility and distortion, by picking latent codes in which are more editable.
In this paper, we introduce a novel approach to mitigate the distortion-editability trade-off, allowing convincing edits on real images that are out-of-distribution. Instead of projecting the input image into the learned manifold, we augment the manifold to include the image by slightly altering the generator, in a process we call pivotal tuning. This adjustment is analogous to shooting a dart and then shifting the board itself to compensate for a near hit.
Since StyleGAN training is expensive and the generator achieves unprecedented visual quality, the popular approach is to keep the generator frozen. In contrast, we propose producing a personalized version of the generator, that accommodates the desired input image or images. Our approach consists of two main steps. First, we invert the input image to an editable latent code, using off-the-shelf inversion techniques. This, of course, yields an image that is similar to the original, but not necessarily identical. In the second step, we perform Pivotal Tuning — we lightly tune the pretrained StyleGAN, such that the input image is generated when using the pivot latent code found in the previous step (see Figure 2 for an illustration.). The key idea is that even though the generator is slightly modified, the latent code keeps its editing qualities. As can be seen in our experiments, the modified generator retains the editing capabilities of the pivot code, while achieving unprecedented reconstruction quality. As we demonstrate, the pivotal tuning is a local operation in the latent space, shifting the identity of the pivotal region to the desired one with minimal repercussions. To minimize side-effects even further, we introduce a regularization term, enforcing only a surgical adaptation of the latent space. This yields a version of the generator StyleGAN that can edit multiple target identities without interference.
In essence, our method extends the high quality editing capabilities of the pretrained StyleGAN to images that are out of its distribution, as demonstrated in Figure 1. We validate our approach through quantitative and qualitative results, and demonstrate that our method achieves state-of-the-art results for the task of StyleGAN inversion and real image editing. In Section 4, we show that not only do we achieve better reconstruction, but also superior editability. We show this through the utilization of several existing editing techniques, and achieve realistic editing even on challenging images. Furthermore, we confirm that using our regularization restricts the pivotal tuning side effect to be local, with negligible effect on distant latent codes, and that pivotal tuning can be applied for multiple images simultaneously to incorporate several identities into the same model (see Figure 3). Finally, we show through numerous challenging examples that our pivotal tuning-based inversion approach achieves completely automatic, fast, faithful, and powerful editing capabilities.
Related Work
Most real-life applications require control over the generated image. Such control can be obtained in the unconditional setting, by first learning the manifold, and then realizing image editing through latent space traversal. Many works have examined semantic directions in the latent spaces of pre-trained GANs. Some using full-supervision in the form of semantic labels , others find meaningful directions in a self-supervised fashion, and finally recent works present unsupervised methods to achieve the same goal , requiring no manual annotations.
More specifically for StyleGAN, Shen et al. use supervision in the form of facial attribute labels to find meaningful linear directions in the latent space. Similar labels are used by Abdal et al. to train a mapping network conditioned on these labels. Harkonen et al. identify latent directions based on Principal Component Analysis (PCA). Shen et al. perform eigenvector decomposition on the generator’s weights to find edit directions without additional supervision. Collins et al. borrow parts of the latent code of other samples to produce local and semantically aware edits. Wu et al. discover disentangled editing controls in the space of channel wise style parameters. Other works focus on facial editing, as they utilize a prior in the form of a 3D morphable face model. Most recently, Patashnik et al. utilize a contrastive language-image pre-training (CLIP) models to explore new editing capabilities. In this paper, we demonstrate our inversion approach by utilizing these editing methods as downstream tasks. As seen in Section 4, our PTI process induces higher visual quality for several of these popular approaches.
2 GAN inversion
As previously mentioned, in order to edit a real image using latent manipulation, one must perform GAN inversion , meaning one must find a latent vector from which the generator would generate the input image. Inversion methods can typically be divided into optimization-based ones — which directly optimize the latent code using a single sample , or encoder-based ones — which train an encoder over a large number of samples . Many works consider specifically the task of StyleGAN inversion, aiming at leveraging the high visual quality and editability of this generator. Abdal et al. demonstrate that it is not feasible to invert images to StyleGAN’s native latent space without significant artifacts. Instead, it has been shown that the extended is much more expressive, and enables better image preservation. Menon et al. use direct optimization for the task of super-resolution by inverting a low-resolution image to space. Zhu et al. use a hybrid approach: first, an encoder is trained, then a direct optimization is performed. Richardson et al. were the first to train an encoder for inversion which was demonstrated to solve a variety of image-to-image translation tasks.
3 Distortion-editability tradeoff
Even though inversion achieves minimal distortion, it has been shown that the results of latent manipulations over inversions are inferior compared to the same manipulations over latent codes from StyleGAN’s native space . Tov et al. define this as the distortion-editability tradeoff, and design an encoder that attempts to find a ”sweet-spot” in this trade-off.
Similarly, the tradeoff was also demonstrated by Zhu et al. , who suggests an improved embedding algorithm using a novel regularization method. StyleFlow also concludes that real image editing produces significant artifacts compared to images generated by StyleGAN. Both Zhu et al. and Tov et al. achieve better editability compared to previous methods but also suffer from more distortion. In contrast, our method combines the editing quality of inversions with highly accurate reconstructions, thus mitigating the distortion-editability tradeoff.
4 Generator Tuning
Typically, editing methods avoid altering StyleGAN, in order to preserve its excellent performance. Some works, however, do take the approach we adopt as well, and tune the generator. Pidhorskyi et al. train both the encoder and the StyleGAN generator, but their reconstruction results suffer from significant distortion, as the StyleGAN tuning step is too extensive. Bau et al. propose a method for interactive editing which tunes the generator proposed by Karras et al. to reconstruct the input image. They claim, however, that directly updating the weights results in sensitivity to small changes in the input, which induces unrealistic artifacts. In contrast, we show that after directly updating the weights, our generator keeps its editing capabilities, and demonstrate this over a variety of editing techniques. Pan et al. invert images to BigGAN’s latent space by optimizing a random noise vector and tuning the generator simultaneously. Nonetheless, as we demonstrate in Section 4, optimizing a random vector decreases reconstruction and editability quality significantly for StyleGAN.
Method
Our method seeks to provide high quality editing for a real image using StyleGAN. The key idea of our approach is that due to StyleGAN’s disentangled nature, slight and local changes to its produced appearance can be applied without damaging its powerful editing capabilities. Hence, given an image, possibly is out-of-distribution in terms of appearance (e.g., real identities, extreme lighting conditions, heavy makeup, and/or extravagant hair and headwear), we propose finding its closest editable point within the generator’s domain. This pivotal point can then be pulled toward the target, with only minimal effect in its neighborhood, and negligible effect elsewhere. In this section, we present a two-step method for inverting real images to highly editable latent codes. First, we invert the given input to in the native latent space of StyleGAN, . Then, we apply a Pivotal Tuning on this pivot code to tune the pretrained StyleGAN to produce the desired image for input . The driving intuition here is that since is close enough, training the generator to produce the input image from the pivot can be achieved through augmenting appearance-related weights only, without affecting the well-behaved structure of StyleGAN’s latent space.
The purpose of the inversion step is to provide a convenient starting point for the Pivotal Tuning one (Section 3.2). As previously stated, StyleGAN’s native latent space provides the best editability. Due to this and since the distortion is diminished during Pivotal Tuning, we opted to invert the given input image to this space, instead of the more popular extension. We use an off-the-shelf inversion method, as proposed by Karras et al. . In essence, a direct optimization is applied to optimize both latent code and noise vector to reconstruct the input image , measured by the LPIPS perceptual loss function . As described in , optimizing the noise vector using a noise regularization term improves the inversion significantly, as the noise regularization prevents the noise vector from containing vital information. This means that once has been determined, the values play a minor role in the final visual appearance. Overall, the optimization defined as the following objective:
where is the generated image using a generator with weights . Note that we do not use StyleGAN’s mapping network (converting from to ). denotes the perceptual loss, is a noise regularization term and is a hyperparameter. At this step, the generator remains frozen.
2 Pivotal Tuning
Applying the latent code obtained in the inversion, produces an image that is similar to the original one , but may yet exhibit significant distortion. Therefore, in the second step, we unfreeze the generator and tune it to reconstruct the input image given the latent code obtained in the first step, which we refer to as the pivot code . As we demonstrate in Section 4, it is crucial to use the pivot code, since using random or mean latent codes lead to unsuccessful convergence. Let be the generated image using and the tuned weights . We fine tune the generator using the following loss term:
where the generator is initialized with the pretrained weights . At this step, is constant. The pivotal tuning can trivially be extended to images , given the inversion latent codes :
Once the generator is tuned, we can edit the input image using any choice of latent-space editing techniques, such as those proposed by Shen et al. or Harkonen et al. . Numerous results are demonstrated in Section 4.
3 Locality Regularization
As we demonstrate in Section 4, applying pivotal tuning on a latent code indeed brings the generator to reconstruct the input image in high accuracy, and even enables successful edits around it. At the same time, as we demonstrate in Section 4.3, Pivotal tuning induces a ripple effect — the visual quality of images generated by non-local latent codes is compromised. This is especially true when tuning for a multitude of identities (see Figure. 14). To alleviate this side effect, we introduce a regularization term, that is designed to restrict the PTI changes to a local region in the latent space. In each iteration, we sample a normally distributed random vector and use StyleGAN’s mapping network to produce a corresponding latent code . Then, we interpolate between and the pivotal latent code using the interpolation parameter , to obtain the interpolated code :
Finally, we minimize the distance between the image generated by feeding as input using the original weights and the image generated using the currently tuned ones :
This can be trivially extended to random latent codes:
where , , are constant positive hyperparameters. Additional discussion regarding the effects of different values can be found in the Supplementary Materials.
Experiments
In this section, we justify the design choices made and evaluate our method. For all experiments we use the StyleGAN generator . For facial images, we use a generator pre-trained over the FFHQ dataset , and we use the CelebA-HQ dataset for evaluation. In addition, we have also collected a handful of images of out-of-domain and famous figures, to highlight our identity preservation capabilities, and the unprecedented extent of images we can handle that could not be edited until now.
We start by qualitatively and quantitatively comparing our approach to current inversion methods, both in terms of reconstruction quality and the quality of downstream editing. We use the direct optimization scheme proposed by Karras et al. to invert real images to space, which we denote by SG2. A similar optimization is used to invert to the extended space , denoted by SG2 . We also compare to e4e, the encoder designed by Tov et al. , which uses the space but seeks to remain relatively close to . Each baseline inverts to a different part of the latent space, demonstrating the different aspects of the distortion-editability trade-off. Note that we do not include Richardson et al. in our comparisons, since Tov et al. have convincingly shown editing superiority, rendering this comparison redundant.
Qualitative evaluation. Figures 4 and 5 present a qualitative comparison of visual quality of inverted images. As can be seen, even before considering editability, our method achieves superior reconstruction results for all examples, especially for out-of-domain ones, as our method is the only one to successfully reconstruct challenging details such as face painting or hands (Figure 4). Our method is also capable of reconstructing fine-details which most people are sensitive to, such as the make-up, lighting, wrinkles, and more (Figure 5). For more visual results, see the Supplementary Materials.
Quantitative evaluation. For quantitative evaluation, we employ the following metrics: pixel-wise distance using , perceptual similarity using , structural similarity using , and identity similarity by employing a pretrained face recognition network . The results are shown in Table 1. As can be seen, the results align with our qualitative evaluation as we achieve the best score for each metric by a substantial margin.
2 Editing Quality
Editing a facial image should preserve the original identity while performing a meaningful and visually plausible modification. However, it has been shown that using less editable embedding spaces, such as , results in better reconstruction, but also in less meaningful editing compared to the native space. For example, using the same latent edit, rotating a face in space results in a higher rotation angle compared to . Hence, in cases of minimal effective editing, the identity may seem to be preserved rather well. Therefore, we evaluate editing quality on two axes: identity preservation and editing magnitude.
Qualitative evaluation. We use the popular GANSpace and InterfaceGAN methods for latent-based editing. These approaches are orthogonal to ours, as they require the use of an inversion algorithm to edit real images. As can be expected, the -based method preserves the identity rather well, but fails to perform significant edits, the -based one is able to perform the edit, but loses the identity, and e4e provides a compromise between the two. In all cases, our method preserves identity the best and displays the same editing quality as for -based inversions. Figure 6 presents an editing comparison over the CelebA-HQ dataset. We also investigate our performance using images of other iconic characters (Figures 1 and 9) and more challenging out-of-domain facial images (Figure 10). The ability to perform sequential editing is presented in Figures 12, and 13. In addition, we demonstrate our ability to invert multiple identities using the same generator in Figures 3 and 11. For more visual and uncurated results, see the Supplementary Materials. As can be seen, our method successfully performs meaningful edits, while preserving the original identity successfully.
The recent work of StyleClip demonstrates unique edits, driven by natural language. In Figures 7, and 8 we demonstrate editing results using this model, and demonstrate substantial identity preservation improvement, thus extending StyleClip’s scope to more challenging images. We use the mapper-based variant proposed by the paper, where the edits are achieved by training a mapper network to edit input latent codes. Note that the StyleClip model is trained to handle codes returned by the e4e method. Hence, to employ this model, our PTI process uses e4e-based pivots instead of ones. As can be expected, we observe that the editing capabilities of the e4e codes are preserved, while the inherent distortion caused by e4e is diminished using PTI. More results for this experiment can be found in the supplementary materials.
Quantitative evaluation results are summarized in Table 2. To measure the two aforementioned axes, we compare the effects of the same latent editing operation between the various aforementioned baselines, and the effects of editing operations that yield the same editing magnitude. To evaluate editing magnitude, we apply a single pose editing operation and measure the rotation angle using Microsoft Face API , as proposed by Zhu et al. . As the editability increases, the magnitude of the editing effect increases as well. As expected, -based inversion induces a more significant edit compared to inversion for the same editing operation. As can be seen, our approach yields a magnitude that is almost identical to ’s, surpassing e4e and inversions, which indicates we achieve high editability (first row).
In addition, we report the identity preservation for several edits. We evaluate the identity change using a pretrained facial recognition network , and the edits we report for are smile, pose, and age. We report both the mean identity preservation induced by each of these edits (second row), and the one induced by performing them sequentially one after the other (third row). Results indeed validate that our method obtains better identity similarity compared to the baselines.
Since the lack of editability might increase identity similarity, as previously mentioned, we also measure the identity similarity while performing rotation of the same magnitude. Expectedly, the identity similarity for inversion decreases significantly when using a fixed rotation angle editing, demonstrating it is less editable compared to other inversion methods. Overall, the quantitative results demonstrate the distortion-editability tradeoff, as inversion achieves better ID similarity but lower edit magnitude, and inversion achieves inferior ID similarity but higher edit magnitude. In contrast, our method preserves the identity well and provides highly editable embeddings, or in other words, we alleviate the distortion-editability trade-off.
3 Regularization
Our locality regularization restricts the pivotal tuning side effects, causing diminishing disturbance to distant latent codes. We evaluate this effect by sampling random latent codes and comparing their generated images between the original and tuned generators. Visual results, presented in Figure 14, demonstrate that the regularization significantly minimizes the change. The images generated without regularization suffer from artifacts and ID shifting, while the images generated while employing the regularization are almost identical to the original ones. We perform the regularization evaluation using a model tuned to invert identities, as the side effects are more substantial in the multiple identities case. In addition, Figure 15 presents quantitative results. We measure the reconstruction of random latent codes with and without the regularization compared to using the original pretrained generator. To demonstrate that our regularization does not decrease the pivotal tuning results, we also measure the reconstruction of the target image. As can be seen, our regularization reduces the side effects significantly while obtaining similar reconstruction for the target image.
4 Ablation study
5 Implementation details
For the initial inversion step, we use the same hyperparameters as described by Karras et al. , except for the learning rate which is changed to . We run the inversion for iterations. Then, for pivotal tuning, we further optimize for iterations with a learning rate of using the Adam optimizer. For reconstruction, we use and and for the regularization we use , , , and .
All quantitative experiments were performed on the first 1000 samples from the CelebA-HQ test set.
Our two-step inversion takes less than 3 minutes on a single Nvidia GeForce RTX 2080. The initial -space inversion step takes approximately one minute, just like the SG2 inversion does. The pivotal tuning takes less than a minute without regularization, and less than two with it. This training time grows linearly with the number of inverted identities. The SG2 inversion takes minutes for iterations. The inversion time of e4e is less than a second, as it is encoder-based and does not require optimization at inference.
Conclusions
We have presented Pivotal Tuning Inversion — an inversion method that allows using latent-based editing techniques on practical, real-life facial images. In a sense, we break the notorious trade-off between reconstruction and editability through personalization, or in other words through surgical adjustments to the generator that address the desired image specifically well. This is achieved by leveraging the disentanglement between appearance and geometry that naturally emerges from StyleGAN’s behavior.
In other words, we have demonstrated increased quality at the cost of additional computation. As it turns out, this deal is quite lucrative: Our PTI optimization boosts performance considerably, while entailing a computation cost of around three minutes to incorporate a new identity — similar to what some of the current optimization-based inversion methods require. Furthermore, we have shown that PTI can be successfully applied to several individuals. We envision this mode of editing sessions to apply, for example, to a casting team of a movie.
Nevertheless, it is still desirable to develop a trainable mapper that approximates the PTI in a short forward pass. This would diminish the current low computational cost that entails real image editing, situating StyleGAN as a practical and accessible facial editing tool for the masses. In addition to a single-pass PTI process, in the future we plan also to consider using a set of photographs of the individual for PTI. This would extend and stabilize the notion of personalization of the target individual, compared to seeing just a single example. Another research direction is to take PTI beyond the architecture of StyleGAN, for example to BigGAN or other novel generative models.
In general, we believe the presented approach of ad-hoc fine tuning a pretrained generator potentially bears merits for many other applications in editing and manipulations of specific images or other generation-based tasks in Machine Learning.
Acknowledgements
We thank Or Patashnik, Rinon Gal and Dani Lischinski for their help and useful suggestions.
References
Appendix A Locality Regularization
We show the effect of different values when using the locality regularization. Figure 17 presents quantitative evaluation and Figure 18 visually demonstrates the interpolated code . We measure both the reconstruction of the target image and the reconstruction of sampled random latent codes, denoted as in-domain, before and after the tuning. For small values (e.g., ), the image generated by the interpolated code is very similar to the image generated by the pivot code . Therefore, the regularization is limited and less effective. For high values (e.g., ) the interpolated code image is more or less equivalent to simply using a random latent . Hence, the interpolated code is less affected by the pivotal tuning, which decreases the regularization constraint. Extremely high values (e.g., ) result in extremely unrealistic images, as can be seen in Figure 18, which cause deterioration of the target image reconstruction. Overall, we get the most effective regularization using an interpolated image which is highly similar to both pivot and the random images, e.g., in Figure 18.
Appendix B Visual Results
Figures 19 and 20 demonstrate the inversion of multiple identities. To prevent the suspicion of cherry-picking, we provide uncurated editing comparison results of the first images from CelebA-HQ test set, provided in Figures 21 to 23. To further avoid picking, we perform the same three edits recurrently. Figures 24 to 27 present further comparisons of editing quality over real images of recognizable characters, and Figure 28 depicts reconstruction results. Finally, additional StyleClip editing results can be found Figure 29. All additional results show that our method achieves higher reconstruction and editing quality, even for challenging images. This enables us to preserve the original identity successfully, while still maintaining high editability.