EditGAN: High-Precision Semantic Image Editing
Huan Ling, Karsten Kreis, Daiqing Li, Seung Wook Kim, Antonio Torralba, Sanja Fidler
Introduction
AI-driven photo and image editing has the potential to streamline the workflow of photographers and content creators and to enable new levels of creativity and digital artistry . AI-based image editing tools have already found their way into consumer software in the form of neural photo editing filters, and the deep learning research community is actively developing further techniques. A particularly promising line of research uses generative adversarial networks (GANs) and either embeds images into the GAN’s latent space or works directly with GAN-generated images. Careful modifications of the latent embeddings then translate to desired changes in generated output, allowing, for example, to coherently change facial expressions in portraits , change viewpoint or shapes and textures of cars , or to interpolate between different images in a semantically meaningful manner .
Most GAN-based image editing methods fall into few categories. Some works rely on GANs conditioning on class labels or pixel-wise semantic segmentation annotations , where different conditionings lead to modifications in the output, while others use auxiliary attribute classifiers to guide synthesis and edit images. However, training such conditional GANs or external classifiers requires large labeled datasets. Therefore, these methods are currently limited to image types for which large annotated datasets are available, like portraits . Furthermore, even if annotations are available, most techniques offer only limited editing control, since these annotations usually consist only of high-level global attributes or relatively coarse pixel-wise segmentations. Another line of work focuses on mixing and interpolating features from different images , thereby requiring reference images as editing targets and usually also not offering fine control. Other approaches carefully analyze and dissect GANs’ latent spaces, finding disentangled latent variables suitable for editing , or control the GANs’ network parameters . Usually, these methods do not enable detailed editing and are often slow.
In this work, we are addressing these limitations and propose EditGAN, a novel GAN-based image editing framework that enables high-precision semantic image editing by allowing users to modify detailed object part segmentations. EditGAN builds on a recently proposed GAN that jointly models both images and their semantic segmentations based on the same underlying latent code , and requires as few as 16 labeled examples – allowing it to scale to many object classes and choices of part labels. We achieve editing by modifying the segmentation mask according to a desired edit and optimizing the latent code to be consistent with the new segmentation, thus effectively changing the RGB image. To achieve efficiency, we learn editing vectors in latent space that realize the edits, and that can be directly applied on other images, without any or only few additional optimization steps. We can thus pre-train a library of interesting edits that a user can directly utilize in an interactive tool.
We apply EditGAN on a wide range of images, including images of cars, cats, birds, and human faces, demonstrating unprecedented high-precision editing. We perform quantitative comparisons to multiple baselines and outperform them in metrics such as identity preservation, quality preservation, and target attribute accuracy, while requiring orders of magnitude less annotated training data. EditGAN is the first GAN-driven image editing framework, which simultaneously (i) offers very high-precision editing, (ii) requires only very little annotated training data (and does not rely on external classifiers), (iii) can be run interactively in real time, (iv) allows for straightforward compositionality of multiple edits, (v) and works on real embedded, GAN-generated, and even out-of-domain images.
Related Work
Image Editing and Manipulation. Image Editing has a long history in computer vision and graphics, as well as machine learning . Recently, deep generative models , in particular modern GANs , received much attention as a promising tool for efficient image editing, as it was found that latent space manipulations often lead to interpretable and predictable changes in output .
GAN-based image editing methods can be broadly sorted into a number of categories. (i) One line of work relies on the careful dissection of the GAN’s latent space, aiming to find interpretable and disentangled latent variables, which can be leveraged for image editing, in a fully unsupervised manner . Although powerful, these approaches usually do not result in any high-precision editing capabilities. The editing vectors we are learning in EditGAN would be too hard to find independently without segmentation-based guidance. (ii) Other works utilize GANs that condition on class or pixel-wise semantic segmentation labels to control synthesis and achieve editing . Hence, these works usually rely on large annotated datasets, which are often not available, and even if available, the possible editing operations are tied to whatever labels are available. This stands in stark contrast to EditGAN, which can be trained in a semi-supervised fashion with very little labeled data and where an arbitrary number of high-precision edits can be learnt. (iii) Furthermore, auxiliary attribute classifiers have been used for image manipulation , thereby still relying on annotated data and usually only providing high-level control. (iv) Image editing is often explored in the context of “interpolating” between a target and different reference image in sophisticated ways, for example by replacing certain features in a given image with features from a reference images . From the general image editing perspective, the requirement of reference images limits the broad applicability of these techniques and prevents the user from performing specific, detailed edits for which potentially no reference images are available. (v) Recently, different works proposed to directly operate in the parameter space of the GAN instead of the latent space to realize different edits . For example, essentially specialize the generator network for certain images at test time to aid image embedding or “rewrite” the network to achieve desired semantic changes in output. The drawback is that such specializations prevent the model from being used in real-time on different images and with different edits. proposed an approach that more directly analyses the parameter space of a GAN and treats it as a latent space in which to apply edits. However, the method still merely discovers edits in the network’s parameter space, rather than actively defining them like we do. It remains unclear whether their method can combine multiple such edits, as we can, considering that they change the GAN parameters themselves. (vi) Finally, another line of research targets primarily very high-level image and photo stylization and global appearance modifications .
Generally, most works only do relatively high-level and not the detailed, high-precision editing, which EditGAN targets. Hence, we consider EditGAN as complementary to this body of work.
GANs and Latent Space Image Embedding. EditGAN builds on top of DatasetGAN and SemanticGAN , which proposed to jointly model images and their semantic segmentations using shared latent codes. However, these works leveraged this model design only for semi-supervised learning, not for editing. EditGAN also relies on an encoder, together with optimization, to embed new images to be edited into the GAN’s latent space. This task in itself has been studied extensively in different contexts before, and we are building on these works. Previous papers studied encoder-based methods , used primarily optimization-based techniques , and developed hybrid approaches .
Finally, a concurrent paper shares similarities with DatasetGAN , on which our method builds, and explores an editing approach related to our EditGAN as one of its applications. However, our editing approach is methodologically different and leverages editing vectors, and also demonstrates significantly more diverse and stronger experimental results. Furthermore, shares some high-level ideas with EditGAN; however, it leverages the CLIP model and targets text-driven editing.
High-Precision Semantic Image Editing with EditGAN
EditGAN’s image generation component is StyleGAN2 , currently the state-of-the-art GAN for image synthesis. The StyleGAN2 generator maps latent codes , drawn from a multivariate Normal distribution, into realistic images. A latent code is first transformed into an intermediate code by a non-linear mapping function and then further transformed into vectors, , through learned affine transformations. These transformed latent codes are fed into synthesis blocks, whose outputs are deep feature maps.
Deep generative models such as StyleGAN2, which are trained to synthesize highly realistic images, acquire a semantic understanding of the modeled images in their high-dimensional feature space. Recently, DatasetGAN and SemanticGAN built on this insight to learn a joint distribution over images and pixel-wise semantic segmentation labels , while requiring only a handful of labeled examples. EditGAN utilizes this joint distribution to perform high-precision semantic image editing of real and synthesized images.
Both methods model by adding an additional segmentation branch to the image generator, which is a pre-trained StyleGAN . We follow DatasetGAN , which applies a simple three-layer multi-layer perceptron classifier on the layer-wise concatenated and appropriately upsampled feature maps. This classifier operates on the concatenated feature maps in a per-pixel fashion and predicts the segmentation label of each pixel.
2 Segmentation Training and Inference by Embedding Images into GAN’s Latent Space
To both train the segmentation branch and perform segmentation on a new image, we embed an image into the GAN’s latent space using an encoder and optimization. To this end, we build on previous works and train an encoder that embeds images into space, which is defined as but where the ’s are modeled independently . Our objectives to train this encoder consist of standard pixel-wise L2 and perceptual LPIPS reconstruction losses using both the real training data as well as samples from the GAN itself. For the GAN samples, we also explicitly regularize the encoder with the known underlying latent codes. In practice, we use the encoder to initialize images’ latent space embeddings and then iteratively refine the latent code via optimization, again using standard reconstruction objectives.
3 Finding Semantics in Latent Space via Segmentation Editing
The key idea of EditGAN lies in leveraging the joint distribution of images and semantic segmentations for high-precision image editing. Given a new image to be edited, we can embed it into EditGAN’s latent space, as described above (alternatively, we can also sample images from the model itself and use those). The segmentation branch will then generate the corresponding segmentation , since segmentations and RGB images share the same latent codes . Using simple interactive digital painting or labeling tools, we can now manually modify the segmentation according to a desired edit. We denote the edited segmentation mask by . Starting from the embedding of the unedited image and segmentation , we can then perform optimization within to find a new consistent with the new segmentation , while allowing the RGB output to change within the editing region.
which means that is defined by all pixels whose part segmentation labels according to either the initial segmentation or the edited one are within an edit-specific pre-specified list of part labels relevant for the edit. For example, when modifying the wheel in a photo of a car would contain all part labels related to the wheels, such as tire, spoke, and wheelhub (see Fig. 3). We use a further buffer of pixels to give the GAN freedom in modeling the transition between the edited and non-edited area. In practice, acts as a binary pixel-wise mask (see Eqs. 2 and 3 below).
To find , approximating , we use the following losses as minimization targets:
where denotes the pixel-wise cross-entropy, loss is based on the Learned Perceptual Image Patch Similarity (LPIPS) distance , and is a regular pixel-wise L2 loss. ensures that the image appearance does not change outside the region of interest, while ensures that the target segmentation is enforced within the editing region (see visualization in Fig. 3). When editing human faces, we also apply the identity loss :
with denoting the pretrained ArcFace feature extraction network and cosine-similiarity.
The final objective function for optimization then becomes:
with hyperparameters . The only “learnable” variable is the editing vector ; all neural networks are kept fixed. After optimizing with the objective function, we can use . Note that there is a certain amount of ambiguity in how the segmentation modification is realized in RGB output. We rely on the GAN generator, trained to synthesize realistic images, to modify the RGB values in the editing region in a plausible way consistent with the segmentation edit.
4 Different Ways of Editing during Inference
The latent space editing vectors obtained by optimization as described are semantically meaningful and often disentangled with other attributes. Therefore, for new images to be edited, we can embed the images into the latent space and the same editing operations can be directly performed by applying the previously learnt as without doing any optimization from scratch again. In other words, the learnt editing vectors amortize the iterative optimization that was necessary to achieve the edit initially. For well-disentangled editing operations, can be used directly as the edited image . Note that we introduced , a scalar editing coefficient, which effectively scales and controls the editing magnitude during inference. For , we do not do any editing at all, while for we manipulate the images with an effectively larger editing operation in latent space, leading to exaggerated effects.
Unfortunately, disentanglement is not always perfect and the editing vectors do not always translate perfectly to other images. We can remove editing artifacts in other regions of the image by a few additional optimization steps at test time. Specifically, we can use the exact same minimization objectives as above, using the initial prediction , obtained after applying the editing vector , as . This assumes that the editing vector still induces a plausible segmentation change when applied on other images and that artifacts only arise in RGB output. The RGB objective then removes these editing artifacts outside the editing region, while ensures that the modified segmentation stays as predicted by the editing vector.
Summarizing, we can perform image editing with EditGAN in three different modes:
Real-time Editing with Editing Vectors. For localized, well-disentangled edits we perform editing purely by applying previously learnt editing vectors with varying scales and manipulate images at interactive rates.
Vector-based Editing with Self-Supervised Refinement. For localized edits that are not perfectly disentangled with other parts of the image, we can remove editing artifacts by additional optimization at test time, while initializing the edit using the learnt editing vectors.
Optimization-based Editing. Image-specific and very large edits do not transfer to other images via editing vectors. For such operations, we perform optimization from scratch.
Experiments
We extensively evaluate EditGAN on images across four different categories: Cars ( spatial resolution), Birds (), Cats (), and Faces ().
We train our segmentation branch as described in Sec. 3.2 using 16, 16, 30, and 30 image-mask pairs as labeled training data for Faces, Cars, Birds, and Cats, respectively. We utilize very highly-detailed part segmentations from . The annotation scheme for faces is shown in Fig. 8, all others are presented in the Appendix. When editing is done purely optimization-based or when learning the editing vectors, we always perform 100 steps of optimization using Adam . For Car, Cat, and Faces, we use real images from DatasetGAN’s test set that were not part of GAN training to demonstrate editing functionality. These images are first embedded into EditGAN’s latent space via an encoder and optimization as described in Sec. 3.2. For Birds, we show editing on GAN-generated images. Model details and hyperparameters are provided in the Appendix.
1 Qualitative Results
In Fig. 4, we demonstrate our EditGAN framework when applying previously learnt editing vectors on novel images and refining with 30 steps of optimization. Our editing operations preserve high image quality and are well disentangled for all classes. We also show the ability to combine multiple different edits in Fig. 16. To the best of our knowledge, no previous methods can perform as complex and high-precision edits as we do, while preserving image quality and subject identity. In Fig. 8, we demonstrate that we can even perform extremely high-precision edits, such as rotating a car’s wheel spoke or dilating pupils. EditGAN can edit semantic parts of objects that consist of only few pixels. At the same time, we can use EditGAN to perform large-scale modifications, too: In Fig. 9, we present how we can remove the entire roof of a car or convert it to a station wagon-like vehicle, simply by modifying the segmentation mask accordingly and optimizing. It is worth noting that several of our editing operations generate plausible manipulated images unlike those appearing in the GAN training data. For example, the training data does not include cats with overly large eyes or ears. Nevertheless, we achieve such edits in a high-quality manner.
The edits in Figs. 4, 16 and 8 are based on learnt editing vectors with self-supervised refinement. However, without such refinement usually only very minor artifacts occur, as shown in Fig. 10, hence allowing for real-time high-precision semantic image editing (discussed in detail below).
Out-of-Domain Results
We demonstrate the generalization capability of EditGAN to out-of-domain data on the MetFaces data set. We use our EditGAN model trained on FFHQ , and create editing vectors using in-domain real faces. We then embed out-of-domain MetFaces partraits (with 100 steps optimization) and apply the editing vectors with 30 steps self-supervised refinement. The results are shown in Fig. 6. We find that our editing operations seamlessly translate even to such far out-of-domain examples.
2 Quantitative Results
To quantitatively measure EditGAN’s image editing capabilities, we use the smile edit benchmark introduced by MaskGAN . Faces with neutral expressions are converted into smiling faces and performance is measured by three metrics: a. Semantic Correctness: Using a pre-trained smile attribute classifier, we measure whether the faces show smiling expressions after editing. b. Distribution-level Image Quality: Frechet Inception Distance (FID) and Kernel Inception Distance (KID) are calculated between 400 edited test images and the CelebA-HD test dataset. c. Identity Preservation: Using the pretrained ArcFace feature extraction network , we measure whether the subjects’ identity is maintained when applying the edit. Specifically, we report cosine-similiarity between original and edited images. Further details can be found in the Appendix.
For our EditGAN, we simply learn a smiling editing vector using a hold-out neutral expression face image. We embed it into EditGAN, infer its pixel-wise segmentation labels, and manually modify the segmentation towards a smile. Then we perform optimization in latent space, as described above, to learn the editing vector. For the results in Tab. 2, it is applied with unit scale on new images. We do not use the identity loss (Eq. 4) in this experiment, since identity preservation is already a target metric itself. We compare our method with three strong baselines: (i) MaskGANhttps://github.com/switchablenorms/CelebAMask-HQ : It takes non-smiling images, their segmentation masks, and a target smiling segmentation mask as inputs. Note that training MaskGAN requires large annotated datasets, in contrast to us. We also compare to (ii) LocalEditinghttps://github.com/IVRL/GANLocalEditing : It clusters GAN features to achieve local editing and relies on reference images, in this case images of faces with smiling expressions. Another baseline we use is (iii) InterFaceGANhttps://github.com/genforce/interfacegan : Similar to EditGAN, InterFaceGAN aims at finding editing vectors in latent space. However, it uses auxiliary attribute classifiers, relies on large annotated datasets, and can generally not achieve the fine editing control of our EditGAN. Finally, we compare to (iv) StyleGAN2 Distillationhttps://github.com/EvgenyKashin/stylegan2-distillation , which creates an alternative approach that does not require real image embeddings and also relies on an editing-vector model to create a training dataset.
Results are reported in Tab. 2. Using less training labels, we outperform MaskGAN on all three metrics. We similarly obtain significantly stronger results than LocalEditing. In our observation, LocalEditing does not work well on real image embeddings. We further exploit a better encoder for the LocalEditing baseline, which leads to a significant improvement in attribute accuracy and ID score, but slightly worse FID & KID scores. We find that EditGAN outperforms InterFaceGAN on identity preservation and attribute classification accuracy, while InterFaceGAN reaches slightly better FID & KID scores (for the results in Tab. 2, the latent space edits learnt by InterfaceGAN are also applied with unit scale, like for EditGAN). In Fig. 11, we report a more detailed comparison to InterFaceGAN, where we apply the smile editing vectors with different scale coefficients from zero to two. As shown, when the editing vector scale is small, the identity score is high while the smiling attribute score is low, since the modification of the original images is minimal. We find that our real-time editing with editing vectors is on-par with InterFaceGAN. When we perform self-supervised refinement at test time, EditGAN outperforms InterFaceGAN. In Tab. 2, we also compare with StyleGAN2 Distillation , which achieves strong performance. However, StyleGAN2 Distillation relies on pre-trained classifiers, like InterfaceGAN, and only enables relatively high-level editing of image attributes for which large-scale annotations exit. Moreover, it distills edits into separate Pixel2PixelHD networks, such that a new network needs to be trained for each edit, limiting broad, user-interactive applicability. Hence, we consider StyleGAN2 Distillation orthogonal to our EditGAN.
We carefully measure the run time of our editing on an NVIDIA Tesla V100 GPU. Conditional optimization, given an edited segmentation mask, with 30 (60) optimization steps takes 11.4 (18.9) seconds. This operation provides us the editing vector. Application of editing vectors is almost instantaneous, taking only 0.4 seconds, therefore allowing for complex real-time interactive editing. A 10 (30) step self-supervised refinement would add an additional 4.2 (9.5) seconds.
3 Ablation Studies: Self-Supervised Refinement and Editing Vector Scale
Fig. 11 also contains a quantitative ablation study on the number of additional optimization steps done when initializing an edit with a learnt editing vector and refining with additional optimization. Generally, the more refinement steps we perform, the better the performance our model can achieve. As shown in Fig. 11, we find that further optimization can indeed slightly improve performance. Specifically, here we improve the trade-off between maintaining identity and achieving the desired semantic operation when performing editing with different scalings of the editing vector. However, performing many steps of optimization leads to a run-time vs. performance trade-off, and our results suggest that the improvement beyond 30 additional optimization steps becomes marginal.
In Fig. 10, we analyze the editing vector scale and self-supervised refinement visually and with respect to perceptual metrics. As highlighted in the zoom-in areas, small artifacts can appear due to imperfect disentanglement in latent space when applying editing operations with large scales. Self-supervised refinement successfully cleans these editing errors up. We also apply the same edit with different scales on 400 test images and measure FID with respect to 10,000 data from GAN training, inspired by the analyses in . We can clearly see that image quality degrades as measured by FID, the stronger the edit is applied. We also observe small improvements with the iterative refinement on this metric, although the difference is small. Further details are in the Appendix. We conclude that for most editing operations, real-time editing without iterative refinement already performs very well. However, to clean up artifacts and maintain highest image quality possible, self-supervised refinement with a couple of additional optimization steps is always available.
Additional experiments are presented in the Appendix.
Conclusions
Like all GAN-based image editing methods, EditGAN is limited to images that can be modeled by the GAN. This makes EditGAN’s application on, for instance, photos of vivid city scenes challenging. Although most of our high-precision edits readily transfer to other images via learnt editing vectors, we also encountered challenging edits that required iterative optimization on each example. Future research therefore includes speeding up the optimization for such edits as well as building better generative models with more disentangled latent spaces.
Summary
We propose EditGAN, a novel method for high-precision, high-quality semantic image editing. It relies on a GAN that jointly models RGB images and their pixel-wise semantic segmentations and that requires only very few annotated data for training. Editing is achieved by performing optimization in latent space while conditioning on edited segmentation masks. This optimization can be amortized into editing vectors in latent space, which can be applied on other images directly, allowing for real-time interactive editing without any or only little further optimization. We demonstrate a broad variety of editing operations on different kinds of images, achieving an unprecedented level of flexibility and freedom in terms of editing, while preserving high image quality.
Broader Impact
Where previous generative modeling-based image editing methods offer only limited high-level editing capabilities, our method provides users unprecedented high-precision semantic editing possibilities. Our proposed techniques can be used for artistic purposes and creative expression and benefit designers, photographers, and content creators . AI-driven image editing tools like ours promise to democratize high-quality image editing. Related methods have already found their way into everyday applications in the form of neural photo editing filters. On a larger scale, the ability to synthesize data with specific attributes can be leveraged in training and finetuning machine learning models.
At the same time, more precise photo editing also offers opportunities for advanced photo manipulation for nefarious purposes. The recent progress of generative models and AI-driven photo editing has profound implications on image authenticity and beyond, which is an area of active debate . As one potential way to tackle these challenges, methods for automatically validating real images and detecting manipulated or fake images are being developed by the research community . Furthermore, generative models like ours are usually only as good as the data they were trained on. Therefore, biases in the underlying datasets are still present in the synthesized images and preserved even when applying our proposed editing methods. It is therefore important to be aware of such biases in the underlying data and counteract them, for example by actively collecting more representative data or by using bias correction methods, an area of active research .
Funding Statement
This work was funded by NVIDIA. Huan Ling and Seung Wook Kim acknowledge additional revenue in the form of student scholarships from University of Toronto and the Vector Institute, which are not in direct support of this work.
References
A Model and Training Details
We first provide additional details about our EditGAN.
EditGAN uses StyleGAN2 as a backbone generative model of images. We denote the image generator as , which is trained following standard StyleGAN training, see for more information . In particular, we use the pre-trained Car, Face-FFHQ and Cat StyleGAN2 models from the official GitHub repository provided by StyleGAN2https://github.com/NVlabs/stylegan2 (Nvidia Source Code License). For Bird, we use the StyleGAN2 model trained on NABirds-48k .
The StyleGAN2 generator maps latent codes , drawn from a multivariate Normal distribution, , into realistic images. A latent code is first transformed into an intermediate code by a non-linear mapping function . is then further transformed into independent vectors, , through learned affine transformations. These transformed latent codes are fed into synthesis blocks, sometimes called style layers and denoted as . The output of these synthesis blocks are deep feature maps . These feature maps carry the information for forming the image , which is achieved by connecting them to a residual image synthesis branch. Further details and visualizations about the StyleGAN2 architecture can be found in .
A.2 Image Encoder
To embed images into the GAN’s latent space, the EditGAN framework relies on optimization, initialized by an encoder. To train this encoder we mainly follow SemanticGAN , which builds on , with further improvements.
where loss is the Learned Perceptual Image Patch Similarity (LPIPS) distance and is a standard L2 loss. We also explicitly regularize the encoder output distribution using an additional loss that utilizes samples from the GAN itself:
Here, is the previously introduced mapping function and are hyperparameters. For all classes, we set , , , , and . We use the Adam optimizer with learning rate and batch size to train the encoder. Experimentally, for the Car and Cat datasets, we first train only on samples from the GAN itself using Eq. 8 for 20,000 iterations as warm up, and then train jointly using Eq. 6 and Eq. 8 iteratively until the model converges on the training dataset.
After successful encoder training, to embed images we first use the encoder and further iteratively refine the latent code via optimization with respect to the objective (without further modifying encoder parameters ). We run optimization for 500 steps with , . We use the Adam optimizer with the lookahead technique [zhang2019lookahead] with a constant learning rate of .
A.3 Segmentation Branch
Similar to DatasetGAN , to generate segmentation maps alongside images we then train a segmentation branch with parameters . is a simple three-layer multi-layer perceptron classifier on the layer-wise concatenated and appropriately upsampled feature maps. Specifically, the lower-resolution deep feature maps in are first appropriately upsampled, for and upsampling functions , so that all feature maps have the same spatial resolution, equal to the highest resolution, and can be concatenated channel-wise. The classifier operates on the layer-wise concatenated feature maps in a per-pixel fashion and predicts the segmentation label of each pixel. It is trained via the objective
where takes as input the concatenated and appropriately upsampled feature maps . We use bilinear-upsampling operations . Furthermore, denotes the pixel-wise cross-entropy.
A.4 Learning Editing Vectors
To perform editing and learn editing vectors, we proceed as described in detail in Secs. 3.3 and 3.4 in the main text. The ArcFace feature extraction network checkpoint is taken from https://github.com/TreB1eN/InsightFace_Pytorch (MIT License).
B Experiment Details
Here, we provide additional experiment details.
In Section 4.2 of the main paper, we evaluate our model against strong baselines on the smile edit benchmark introduced by MaskGAN . Here we provide more details for completeness. Semantic Correctness: To measure whether the faces show smiling expressions after editing, a binary smile attribute classifiers is trained on the CelebA training set, using a ResNet-18 backbone. The input faces are resized into resolution of . The classifier achieives 92.2% accuracy on the CelebA testing dataset. Identity Preservation: We again use the pretrained ArcFace feature extraction network with checkpoint from https://github.com/TreB1eN/InsightFace_Pytorch (MIT License). As pointed out in main paper, we did not use the identity loss in this benchmark experiment when performing face editing. In this benchmark experiment, this facial feature extraction network is used only for evaluation purposes.
To compare with the baselines, we took the officially released MaskGANhttps://github.com/switchablenorms/CelebAMask-HQ and LocalEditinghttps://github.com/IVRL/GANLocalEditing checkpoints. Furthermore, we train an InterFaceGAN smile model using the officially released codehttps://github.com/genforce/interfacegan where we replaced the generator with a StyleGAN2 for fair comparison. At inference time when performing editing, we use the same test image embeddings for the InterFaceGAN model as we use for our EditGAN model. As mentioned in the main paper, we also use StyleGAN2 Distillation as baseline, for which we rely on the official codebasehttps://github.com/EvgenyKashin/stylegan2-distillation and train a Pix2PixHD network for the smile edit using the default hyperparameters as provided in the paper . We further show in Fig. 13 the image and initial and modified segmentation masks that were used to learn our smile editing vector.
Finally, we provide more details for Fig. 10 in the main text: For each curve, we report results with five different editing vector scale coefficients .
B.2 Additional Results: Editing Vector Scale Experiment
In the main paper, in Fig. 9, we presented another ablation study where we studied editing quality when applying edits with different editing vector scales . We analyzed editing quality both visually and quantitatively, both with and without self-supervised refinement. While in the main paper we only presented the results for Car and Cat data, here we additionally show the results on Face images, using an edit that raises eyebrows as example (Fig. 14). Similar to the other results on Car and Cat data, we find that editing by purely applying our learnt editing vector, which can be done at interactive rates, already yields virtually perfect editing results. However, we do observe almost unnoticeable entanglement with the beard. Using self-supervised refinement we can fully remove this editing artifact, if necessary.
B.3 Additional Results: Smile Edit Benchmark with more Test Images
In Tab. 1 of the main paper, we use MaskGAN’s smile edit benchmark. The FID scores are calculated between 400 edited test images and the CelebA-HD test database, which enables a fair comparison with existing approaches and directly follows the practice by MaskGAN. Although the estimates may be biased with respect to the true FID due to the limited number of test images [chong2019effectively], we expect that they nevertheless provide a fair comparison between the different methods. However, here we re-calculate FID as well as attribute accuracy and ID score using 10 times as many images, i.e. 4000 images, from the training set from MaskGAN. Notice that only MaskGAN uses this data for training, while the GANs of all other baselines, including our EditGAN, are based on the FFHQ faces data and do not use this annotated training data that MaskGAN relies on. Hence, calculating the FID using these 4000 images is advantageous for MaskGAN.
We show results in Tab. 2. Since we use different and much more data for evaluation compared to the evaluation reported in the table in the paper, the numbers are different. However, the rankings and comparisons between the methods remain the same and the conclusions are the same. In particular, EditGAN achieves the best attribute accuracies and ID scores. MaskGAN achieves a relatively low FID, but this is simply due to the unfair comparison, as discussed above. MaskGAN still performs significantly worse than InterfaceGAN and EditGAN in attribute accuracy and ID score.
C Computational Resources
Training of the underlying StyleGAN2, the encoder, and the segmentation branch, as well as optimization for embedding and editing were performed using NVIDIA Tesla V100 GPUs on an in-house GPU cluster. Overall, the project used approximately 14,000 GPU hours (according to internal GPU usage reports), of which around 3,500 GPU hours were used for the final experiments, and the rest for exploration and testing during the earlier stages of the research project.
D Additional Qualitative Results
Below, we present further qualitative results.
We first demonstrate particularly challenging editing operations where we try to disentangle semantically related parts. For example, we want to lift the right eyebrow while keeping the left eyebrow unchanged. We present the results ins Fig. 15. Furthermore, we again demonstrate the ability to combine multiple different edits in Fig 16. We also invite the reader to watch our video, which shows latent code interpolations between the edits. Finally, for all edits we perform in the main paper, we first show the image and segmentation mask pairs that were used to learn the latent space editing vectors, and then we present a few more editing results on GAN-generated images (Figs. 17-33).