High-Fidelity GAN Inversion for Image Attribute Editing
Tengfei Wang, Yong Zhang, Yanbo Fan, Jue Wang, Qifeng Chen
Introduction
Image attribute editing is the task of modifying desired attributes of a given image while preserving other details. With the rapid advancement of generative adversarial networks (GANs) , a promising direction is to manipulate images with the strong control capacity of StyleGAN . To enable real-world image editing, GAN inversion techniques have been recently explored, which aim at projecting images to the latent space of a pre-trained GAN generator.
Existing GAN inversion approaches either perform per-image optimization or learn a data-driven encoder . Optimization approaches achieve higher reconstruction accuracy by over-fitting on a single image, but the latent code may get out of GAN manifold, leading to inferior editing quality. In contrast, encoder-based GAN inversion methods are faster and show better editing performance due to knowledge learned from numerous training images. Nevertheless, their reconstruction results are usually inaccurate and of low fidelity: these methods can reconstruct a coarse layout (low-frequency patterns), but the image-specific details (high-frequency patterns) are often ignored. For example, the reconstructed face images typically possess averaged patterns that agree with the majority of training images (e.g., normal pose/expression, occlusion/shadow-free), and the details that present minority patterns (e.g., background, illumination, accessory) in training data are subject to distortion. It is highly desirable to preserve these image-specific details in reconstruction and editing with high fidelity.
Though some works tried to improve the reconstruction accuracy of encoder-based methods, their editing performance usually decreases . To analyze the limitation of existing approaches, we consider the GAN inversion problem as a lossy data compression system with a frozen decoder. According to Rate-Distortion theory , reversing a real-world image to a low-dimensional latent code would inevitably lead to information loss. As conjectured by information bottleneck theory , the lost information is primarily image-specific details as the deep compression model tends to retain common information of a domain. Based on these analyses and experimental observations, we present the Rate-Distortion-Edit trade-off for GAN inversion, which further inspires our framework.
According to this trade-off, the low-rate latent codes are insufficient for high-fidelity GAN inversion. However, it is non-trivial to improve the reconstruction accuracy by directly increasing the rate. A higher-rate latent codes can easily achieve a low distortion by overfiting on the reconstruction process, but would suffer a dramatic editing performance drop. To achieve both accuracy and editability (high-fidelity editing), we propose a novel framework that equips low-rate encoder models with distortion consultation. The consultation branch serves as a ‘cheat sheet’ for generation that only conveys the ignored image-specific information. Specifically, we leverage the distortion map between source and low-fidelity reconstructed image as a reference and project it to higher-rate latent maps. Compared with high-rate latent codes inferred from a full image, the distortion map only conveys image-specific details and can thus alleviate the aforementioned overfitting issue. The high-rate latent map and low-rate latent code are further embedded and fused in the generator via consultation fusion. Our scheme shows a clear improvement in reconstruction quality, and no test-time optimization is involved.
For attribute editing, following previous works, we perform vector arithmetic on the low-rate latent code, while the consultation is desired to bring back lost details. While the distortion consultation substantially contributes to the inversion quality, it cannot directly apply the distortion map observed on the inversion image for editing due to the misalignment between inverted and edited images. To this end, we additionally design an adaptive distortion alignment (ADA) network to adjust the distortion map with the edited images. To disentangle the alignment from the consultation encoder and stabilize the training, we impose intermediate supervision on ADA by proposing an alignment regularization with a self-supervised training scheme.
Extensive experiments show that our method significantly outperforms current approaches in terms of details preservation in both reconstructed and edited results. On account of the high-fidelity inversion capacity, our approach is robust to viewpoint and illumination fluctuation and can thus perform temporally consistent editing on videos. Our primary contributions can be summarized as follows.
We propose a distortion consultation inversion scheme that combines both high reconstruction quality and compelling editability with consultation fusion.
For high-fidelity editing, we propose the adaptive distortion alignment module with a self-supervised learning scheme. By alignment, the distortion information can be propagated well to the edited images.
Our method outperforms state-of-the-art approaches qualitatively and quantitatively on diverse image domains and videos. The framework is simple, fast and can be easily applied to GAN models.
Related Work
Existing GAN inversion approaches can be categorized into optimization-based, encoder-based, and hybrid methods. Optimization approaches can achieve high reconstruction quality but are slow for inference. used L-BFGS, and I2S adopted ADAM for solving the optimization. adopted Covariance Matrix Adaptation for gradient-free optimization. Instead of per-image optimization, learned an encoder to project images. proposed an in-domain method on real images. pSp and GHFeat proposed to embed latent codes in a hierarchical manner. Further, e4e analyzed the trade-offs between reconstruction and editing ability. improved the inversion efficiency by a shallow network with efficient heads. ReStyle projected the latent codes with iterative refinements. These methods are more efficient but fail to achieve high-fidelity reconstruction. Hybrid approaches make a compromise. initialize the optimization with the encoder output for acceleration. designed a collaborative learning scheme for encoder and optimization iterator. fine-tuned StyleGAN parameters for each image after predicting an initial latent code, which takes a few minutes for an image. Compared with previous methods, our method considerably improves the reconstruction quality of encoder models without inference-time optimization.
GAN inversion approaches can also be classified by the used latent space. space is straightforward but suffers from feature entanglement. and space in StyleGAN are more disentangled, where space extends space by using different across layers. space is proposed by transforming through the affine layers. space inverts images to the last activation layer in the non-linear mapping network. Besides StyleGAN, some works also adopts multi-scale latent codes for ProgressGAN . Nevertheless, these latent spaces would inevitably lose details in reconstructed images due to limited bit-rate (Sec. 3.1). To perform a high-fidelity inversion, we propose a distortion consultation branch to convey high-frequency image-specific information.
2 Latent Space Editing
A number of supervised and unsupervised approaches explored GAN latent space for semantic directions under the vector arithmetic. The supervised methods need off-the-shelf attribute classifiers or annotated images for specific attributes. InterfaceGAN trained SVM to learn the boundary hyperplane for each binary attribute. StyleFlow learned reversible mapping by normalizing flow and off-the-shelf classifiers. Others explored simple geometric transformation via self-supervised learning. Unsupervised approaches do not need pre-trained classifiers. GANspace performed PCA on early feature layers. Similarly, SeFa performed eigenvector decomposition of the affine layers. Some found distinguishable directions based on mutual information. LatentCLR explored directions by contrastive learning.
Approach
Given a source image and a well-trained generator , GAN inversion infers the latent code via an encoder , which is expected to faithfully reconstruct . In this section, we first analyze the bottleneck of previous inversion methods and describe our proposed distortion consultation inversion strategy. To handle the features misalignment, we present the adaptive distortion alignment modules with a self-supervised training scheme. The whole framework is illustrated in Figure 3.
Motivation. Currently, GAN inversion frameworks lie in three categories, which are optimization-based, encoder-based and hybrid methods. Despite more accurate, optimization-based and hybrid approaches are time-consuming and thus intolerable in real-time applications. Existing encoder-based methods can be illustrated by Fig. 2 (a), where the decoder is a frozen well-trained generator (e.g., StyleGAN) while the encoder learns a mapping from the source image to the latent codes. As observed in many existing works (e.g., results (a) in Fig. 2), the encoder approaches fail to faithfully reconstruct the input images, and the inversion (and editing) results are of low fidelity in terms of details. Noted the fact that the latent codes in previous methods are of (relatively) low dimension, we conjecture that the low-rate latent codes are insufficient for high-fidelity reconstruction. This conjecture is also supported by the Rate-Distortion theory , which will be reviewed in the Supplement.
To further analyze the effect of the latent rate in high-fidelity GAN inversion, we formulate the encoder-based GAN inversion as a problem of lossy data compression. In this formulation, rate can be interpreted as the dimension of latent codes (e.g., ), and distortion indicates the reconstruction quality (fidelity). A compelling inversion method is desired to produce high-fidelity images for both inversion and editing (low distortion). Nevertheless, the current dimension of latent codes is much smaller than that of images (low rate). This implies a contradiction with , which shows the low-rate latent codes are insufficient for faithful reconstruction and some information is inevitably lost. Therefore, we are motivated to design a large-rate GAN inversion system.
Challenge. However, it is non-trivial to reduce the distortion by simply increasing the latent rate. A naive idea for faithful reconstruction is to adopt a higher-rate latent code like Fig. 2 (b). This Unet-like structure is adopted by some recent image restoration works that conveys latent maps (e.g., ) to decoder. Benefited from the higher bit rate, the restoration quality is gratifying (e.g., results (b) in Fig. 2). However, we cannot apply this structure in our case since the high-dimensional latent codes are difficult to interpret and manipulate for attribute editing (e.g., results (b) in Fig. 2). Similarly, prior work also observed tradeoffs between the reconstruction and editability brought by over-fitting. The high-rate latent code is easy to overfit on the reconstruction, thereby compromising the edit performance. As the inversion is just an intermediate step to achieve the goal of editing, it is essential to balance the rate, reconstruction, and editing quality, which we call the Rate-Distortion-Edit trade-offs (Fig. 2). To this end, a delicate system design is needed.
Design. As analyzed above, with a (relatively) low-rate latent code, the GAN inversion system is subject to inevitable information loss. By analyzing the visual results of previous GAN inversion approaches (Fig. 1, Fig. 2, Fig. 4), we found that these reconstruction results can successfully preserve frequent patterns and principle attributes of the source images. In contrast, the lost information is mostly the image-specific details such as background, make-up and illumination. This observation is consistent with the Information Bottleneck theory , which hypothesizes the deep models primarily learn common patterns in the dataset while forgetting infrequent details for reconstruction.
Considering the Rate-Distortion-Edit trade-offs, now that we have prioritized the editability (with a low-rate latent code), the main concern is how to convey the lost information to improve the fidelity (lower the distortion) without compromising the edit performance. To this end, we propose a distortion consultation branch that only conveys image-specific details to enhance the reconstruction quality, which avoids the trivial solution of a simple overfitting. For editing, we still perform vector arithmetic on the low-rate latent code for its high editability. By combining the best of both worlds, the proposed approach achieves a high fidelity in both reconstruction and editing (Fig. 2).
2 Distortion Consultation Inversion (DCI)
Basic Encoder. With a basic encoder , we can obtain a low-rate latent code and initial inversion image . In this case, the generator takes as the input in each layer to obtain the feature map:
Consultation Fusion. To combine the consultation branch with the basic encoder for image generation, we adopt a layer-wise consultation fusion for latent codes and latent maps , as shown in Fig. 3. As artifacts and inaccurate details introduced by can degrade the generation quality, we design a gated fusion scheme to adaptively filter out undesired features. In layer of , is embedded to a gate map and a high-frequency details map :
where mapping functions and are convolution layers. contains the image-specific details, and facilitates the low-fidelity features obtained from (Eq. (2)) to produce high-fidelity feature maps in StyleGANFor StyleGAN2 , the fusion layer would be .:
To avoid overfitting on the inversion result, we only perform the consultation fusion in early layers of .
3 Adaptive Distortion Alignment (ADA)
4 Losses
We also impose adversarial loss to improve image quality:
where is initialized with the well-trained discriminator. In summary, the overall loss is a weighted summation of , , and . Note that the training process only involves the inversion images, and no editing direction is needed. After training, the model can generalize to diverse attribute editing explored by different methods.
Experiments
Datasets. For the human face domain, we use the FFHQ dataset for training and the CelebA-HQ dataset for evaluation. For the car domain, we use Stanford Cars for training and evaluation. For attribute editing, we adopt InterfaceGAN for face images and GANSpace for car images. Implementation details. See the Appendix.
2 Evaluation
We compare our method (with e4e as the basic encoder) with state-of-the-art encoder-based GAN inversion approaches, pSp , e4e and Restyle (with pSp and e4e as backbones, respectively). We report quantitative comparisons of the inversion performance in Table 1. The metrics are calculated on the first 1,500 images from CelebA-HQ. We also compare the proposed method with two optimization-based approaches . Our approach substantially outperforms encoder-based baselines in terms of reconstruction quality and is considerably faster than optimization-based methods when inference.
2.2 Qualitative Evaluation
Encoder baselines. We show visual results of both inversion and editing in Fig. 4. Compared with previous approaches, our method is robust to images with occlusion and extreme viewpoints. For example, the first row in Fig. 4 gives a face image occluded by the hand, and the last row demonstrates a car image with an out-of-range viewpoint. Existing methods fail to reconstruct these challenging images faithfully. They generate distorted results and suffer artifacts for both inversion and editing. In contrast, with the proposed distortion consultation scheme, our method is more robust with high-fidelity results. Besides the robustness improvement, our approach also successfully preserves more details in backgrounds (4th row), shadow (2nd row), reflect (10th row), accessory (5th row), expressions (7th and 8th rows), and appearance (9th and 11th rows). Optimization baselines. We also compare our method with optimization-based methods in Fig. 5. Note that PTI optimizes both latent codes and StyleGAN parameters, but we still report their results for better comparison. With faster inference, our method achieves a comparable or even better reconstruction quality. Also, the editing results produced by the proposed scheme successfully preserve the image-specific details of source images without compromising the edit performance.
2.3 User Study
To perceptually evaluate the editing performance, we conduct a user study in Table 2. We select the first 50 images from CelebA-HQ and perform editing on extensive attributes. We collect 1,500 votes from 30 participants. Each participant is given a triple of images (source, our editing, baseline editing) at once and asked to choose the higher-fidelity one with proper editing. The user study shows our method outperforms baselines by a large margin.
3 Ablation Study
As discussed before, the distortion consultation inversion (DCI) scheme brings back ignored image details to complement the low-rate basic encoder, thereby achieving the high-fidelity reconstruction. To validate the effectiveness of DCI, we show our inversion results in Fig. 6. With the proposed distortion consultation branch, the model is more robust to occlusion and extreme poses and keeps more details in the reconstruction results.
3.2 Effect of Adaptive Distortion Alignment
4 Application on Video Editing
Compared with image inversion and editing, the key challenge for the video counterpart is the temporal consistency of details across frames. This puts a higher demand for reconstruction fidelity since the distortion of every single image would be magnified in a video in terms of consistency and quality . We show inversion and editing results on a real video in Fig. 8. Previous low-rate inversion approaches lack robustness to pose variation, and fail to preserve the identity of the original person and suffer notable distortion in editing results. When the pose and viewpoint change across video frames, their results show inconsistent details and abrupt identity discrepancy. In contrast, the proposed method is more robust to cross-frame discrepancy (e.g., pose, viewpoint) and achieves higher fidelity for details preservation. More results in mp4 format are given in the Supplement.
Conclusion
In this work, we propose a novel GAN inversion framework that enables high-fidelity image attribute editing. With an information consultation branch, we consult the observed distortion map as a high-rate reference for lost information. This scheme enhances the basic encoder for high-quality reconstruction without compromising editability. With the adaptive distortion alignment and distortion consultation technique, our method is more robust to challenging cases such as images with occlusion and extreme viewpoints. Benefiting from the additional information of the consultation branch, the proposed method shows clear improvements in terms of image-specific details preservation (e.g., background, appearance, and illumination) for both reconstruction and editing. The proposed framework is simple to apply, and we believe it can be easily generalized to other GAN models for future work.
Limitations. One limitation of the proposed method is the difficulty in handling large misalignment cases. As the augmented data used for ADA training in our experiments does not cover extreme misalignment, ADA is possibly insufficient when editing images with large viewpoint changes (see the Supplement for failure cases).