Third Time's the Charm? Image and Video Editing with StyleGAN3

Yuval Alaluf, Or Patashnik, Zongze Wu, Asif Zamir, Eli Shechtman, Dani Lischinski, Daniel Cohen-Or

Introduction

In recent years, Generative Adversarial Networks (GANs) have revolutionized image processing. Specifically, StyleGAN generators synthesize exceedingly realistic images, and enable editing , and image-to-image translation , particularly in well-structured domains. StyleGAN architectures are notable for their semantically-rich, disentangled, and generally well-behaved latent spaces.

When it comes to video processing, new challenges arise. Editing should not only be disentangled and realistic, but also temporally consistent across frames. The texture-sticking phenomenon in StyleGAN1 and StyleGAN2 hinders the temporal consistency and realism of generated and manipulated videos. For example, when interpolating within the latent space, the hair and face typically do not move in unison. The recent StyleGAN3 architecture is specifically designed to overcome such texture-sticking, and additionally offers translation and rotation equivariance. Naturally, these unique properties make StyleGAN3 better suited for video processing than previous style-based generators. However, significant changes introduced in StyleGAN3 architecture raise many questions and new challenges. Central among these is the disentanglement of its latent spaces and the ability to accurately invert and edit real images.

In this paper, we analyze StyleGAN3, aiming at understanding and exploring its capabilities and performance. Some of the questions we attempt to answer are: How does the disentanglement of the latent representations in StyleGAN3 compare to StyleGAN2? Do the techniques devised for identifying latent editing controls still work? Inversion is a fundamental task required for editing real images, and has been extensively studied in the context of StyleGAN2 . We therefore examine how well these existing techniques can be adapted to achieve comparable performance with StyleGAN3, and, in particular, cope with the inversion of unaligned images.

In light of the translation and rotation equivariance provided by StyleGAN3, we also examine the differences between generators trained on aligned and on unaligned data. Surprisingly, we observe that both kinds of generators are comparable in terms of their ability to generate unaligned images and control their position and rotation. However, we find that using the aligned generators is preferable for tasks such as disentangled editing and inversion. We therefore leverage aligned generators in our proposed image and video inversion and editing workflows, see Fig. 1.

Applying the insights gained in our analysis and experiments over still images, we propose a novel workflow for inverting and editing real videos using StyleGAN3. Notably, we leverage the capabilities of StyleGAN3 to reduce texture sticking and expand the field of view when working on a video with a cropped subject.

Related Work

StyleGAN2 features several semantically-rich latent spaces, which have been heavily studied and exploited in the context of image manipulation and image inversion . In this work, we experiment with the image manipulation techniques from Shen et al. , Wu et al. and Patashnik et al. in the context of StyleGAN3 .

To manipulate real images they must first be projected into one of the StyleGAN latent spaces. We refer the reader to the exposition in Xia et al. for a comprehensive review of GAN inversion and the applications it enables. In this work, we leverage existing encoder-based inversion techniques for achieving more accurate inversions of images, and in particular unaligned images, using StyleGAN3. We additionally employ existing generator tuning techniques to achieve higher-fidelity reconstructions of a wide range of facial expressions, which we find to be necessary for inverting and editing videos.

In contrast to editing images with StyleGAN, few works have addressed video-based attribute editing. Most notably, Yao et al. use a pre-trained encoder and introduce a latent space transformer to achieve consistent edits of the inverted video frames. However, their method relies on facial alignment, segmentation, and Poisson blending . We aim to leverage StyleGAN3’s rotation and translation equivariance to achieve accurate and consistent video editing while reducing the overhead of previous techniques.

We refer the reader to Appendix A for additional background on StyleGAN’s latent spaces, inversion techniques, and the editing capabilities it offers.

The StyleGAN3 Architecture

To better understand the capabilities of StyleGAN3 , it is important to understand the overall structure and function of the different components comprising the architecture. First, as in StyleGAN , a simple fully-connected mapping network translates an initial latent code z∼N ⁣(0,1)512z\sim\mathcal{N}\!\left(0,1\right)^{512}, into an intermediate code ww residing in a learned latent space W\mathcal{W}.

Compared to StyleGAN2 , StyleGAN3’s synthesis network is composed of a fixed number of convolutional layers (1616), irrespective of the output image resolution. We denote by (w0,...,w15)(w_{0},...,w_{15}) the set of input codes passed to these layers. In StyleGAN3, the constant 4×44\times 4 input tensor from StyleGAN2 is replaced by Fourier features, that can be rotated and translated using four parameters (sin⁡α\sin\alpha, cos⁡α\cos\alpha, xx, yy), obtained from w0w_{0} via a learned affine layer. In the remaining layers, each wiw_{i} is fed into an independently learned affine layer, which yields modulation factors used to adjust the convolutional kernel weights.

In StyleGAN2, the space spanned by the outputs of these affine layers has been referred to as the StyleSpace , or S\mathcal{S}. In this work, we similarly define the S\mathcal{S} space of StyleGAN3, with 9,8949,894 dimensions for a 1024×10241024\times 1024 generator.

Since the translation and in-plane rotation of the synthesized images are given by explicit parameters obtained from w0w_{0}, the result may be easily adjusted by concatenating another transformation. We parameterize this transformation using three parameters (r,tx,ty)(r,t_{x},t_{y}), where rr is the rotation angle (in degrees), and tx,tyt_{x},t_{y} are the translation parameters, and denote the resulting image by:

where, by default, tx=ty=0t_{x}=t_{y}=0 and r=0r=0. This transformation can be applied even in a generator trained solely on aligned data, enabling it to generate rotated and translated images, see Fig. 2.

Conversely, generators trained on unaligned data may be “coerced” to generate roughly aligned images by setting w0w_{0} to the generator’s average latent code w‾\overline{w}, i.e, given by G((w‾,w1,...,w15);(0,0,0))G((\mkern 1.5mu\overline{\mkern-1.5muw\mkern-1.5mu}\mkern 1.5mu,w_{1},...,w_{15});(0,0,0)). Fig. 3 demonstrates this idea on two different StyleGAN3 generators trained on unaligned FFHQ and AFHQ datasets, respectively. Intuitively, this approximate alignment may be due to the fact that the average input pose in the training distribution is roughly aligned and centered, combined with the fact that the translation and rotation transformations in StyleGAN3 are mainly controlled by the first layer. It is important to note that equivariance is at the core of the StyleGAN3 design: translations or rotations in earlier layers are preserved across later layers and appear in the generated output.

Analysis

As discussed above, the latent code fed into the first layer, w0w_{0}, controls the translation and rotation of the image content. However, as illustrated in Fig. 3, w0w_{0} affects each image in a slightly different manner. For example, the leftmost human face is slightly rotated while the face in the third column is perfectly upright. This suggests that rotation is also affected by other layers of the generator, where it is entangled with other visual attributes. To examine the extent of this phenomenon, we perform two experiments, illustrated in Fig. 4. First, we examine a series of images G((w∗,w1,w∗,...,w∗))G((w^{*},w_{1},w^{*},...,w^{*})) that differ only in their randomly sampled w1w_{1} latent entry (top row). It may be seen that altering w1w_{1} affects the in-plane rotation of the face, but this change is entangled also with other attributes, such as face shape and eyes. In our second experiment, we generate a series of images G((w0,w1,w∗,...,w∗))G((w_{0},w_{1},w^{*},...,w^{*})), where w0w_{0} and w1w_{1} are held fixed, while the remaining latent entries are set to a randomly sampled code w∗w^{*}. It may be seen that with both w0w_{0} and w1w_{1} fixed, the generated images all share the same head pose. Thus, we conclude that the subsequent layers do not appear to induce any further translation or rotation, and those are determined primarily by w0w_{0} and w1w_{1}.

2 Disentanglement Analysis

To analyze the disentanglement of the different latent spaces of StyleGAN3, we follow Wu et al. and compute the DCI (disentanglement / completeness / informativeness) metrics of each latent space. To compute the above metrics we employ pre-trained attribute regressors for various attributes, as described by Wu et al. .

Observe that accurately computing attribute scores for unaligned images is challenging given that the attribute classifiers were trained solely on aligned images. To this end, we provide the classifiers with pseudo-aligned images, generated as described in Sec. 3.

We report the DCI metrics for the Z\mathcal{Z}, W\mathcal{W}, and S\mathcal{S} spaces in Tab. 1 for both the aligned and unaligned StyleGAN3 generators trained on the FFHQ dataset. S\mathcal{S} achieves the highest DCI scores across both StyleGAN3 generators, as it also does for StyleGAN2. Furthermore, while gaps in D and C between Z\mathcal{Z} and W\mathcal{W} are smaller in StyleGAN3 than they are in StyleGAN2, the gap between W\mathcal{W} and S\mathcal{S} is larger, suggesting that using S\mathcal{S} for editing may be even more beneficial in StyleGAN3.

Since most StyleGAN inversion methods invert images into the W+\mathcal{W}+ latent space, it is also beneficial to examine the DCI metrics for this extended latent space. To do so, we randomly sample a set of latent codes w∈Ww\in\mathcal{W} and concatenate them to form latent codes in W+\mathcal{W}+. However, we find the resulting generated images are unnatural (see Sec. 4.2). Moreover, applying the pre-trained DCI classifiers on such images results in inaccurate attribute scores, making the computed metrics unreliable.

Image Editing

In this section, we examine the effectiveness of various techniques for image editing with StyleGAN3, starting with the W\mathcal{W} and W+\mathcal{W}+ latent spaces, and proceeding to S\mathcal{S}.

Here we use InterFaceGAN for finding linear directions in W\mathcal{W} for aligned and unaligned StyleGAN3 generators. Editing aligned images is simple and follows the approach used in StyleGAN2 : given a randomly sampled latent code w∈Ww\in\mathcal{W}, an editing direction DD, and a step size δ\delta, the edited image is generated by Galigned(w+δD;(0,0,0))G_{\textit{aligned}}(w+\delta D;(0,0,0)), where GalignedG_{\textit{aligned}} is the aligned generator.

As for unaligned images, there are two options. First, one may simply use an unaligned generator. Yet, one problem that arises in doing so is the fact that the attribute scores needed to learn these directions are obtained from classifiers pre-trained on aligned images. The scores produced by these classifiers on unaligned images may be inaccurate, resulting in poorly-learned directions in W\mathcal{W}. To assist the pre-trained classifiers, we generate pseudo-aligned images by replacing w0w_{0} with the generator’s average latent code w‾\overline{w}, as shown in Sec. 3. The image generated by this modified latent code is then passed to the pre-trained classifier to obtain the original latent’s attribute score. Yet, another problem with using the unaligned generator is that it requires learning a separate set of directions.

A second approach to mitigate the above overhead is to generate images using the aligned generator, but apply the user-defined transformations to control the rotation and placement of the generated object. Specifically, an edited unaligned image can be synthesized by Galigned(w+δD;(r,tx,ty))G_{\textit{aligned}}(w+\delta D;(r,t_{x},t_{y})), where (r,tx,ty)(r,t_{x},t_{y}) is controlled by the user. This gives the added benefit that the same latent directions may be used to edit both aligned and unaligned images.

In Fig. 5 we provide editing results obtained using the three approaches above. Notably, it is possible to achieve comparable edits on unaligned images via both the aligned and the unaligned generators. We also find that linear directions found in the latent space of GunalignedG_{\textit{unaligned}} to be generally more entangled than those found in the latent space of GalignedG_{\textit{aligned}}. This is most notable in the “smile” direction found in GunalignedG_{\textit{unaligned}}, see row 22. We attribute this entanglement to two factors: (1) the pseudo-aligned images may still be out-of-domain with respect to the classifier trained on aligned images, resulting in less accurate attribute scores; and (2) the linear editing directions make it more challenging to attain disentangled editing.

Given the insight that unaligned images may be edited using a single aligned generator and the fact that the aligned generator produces higher quality images (as shown in StyleGAN3 ), we focus our subsequent analysis on the aligned generator.

Editing via Non-Linear Latent Paths.

Various works have demonstrated that editing images via non-linear latent paths typically results in more faithful, disentangled edits . Following these works, we now explore learning non-linear latent editing paths within the W+\mathcal{W}+ latent space using the StyleCLIP mapper technique . As shown in Fig. 6, the resulting edits are still entangled. For example, the image background typically changes across the different edits, even for local edits such as “angry”. These results lead us to explore whether editing within the S\mathcal{S} space of StyleGAN3 achieves latent edits that are more disentangled than those achievable with W\mathcal{W} and W+\mathcal{W}+.

Editing via Latent Directions in 𝒮𝒮\mathcal{S}.

Recall that our DCI analysis (Tab. 1) indicates that the S\mathcal{S} space is more disentangled and complete than the W\mathcal{W} latent spaces. Here, we examine whether this finding extends to the editing quality of these spaces, particularly in terms of editing disentanglement. To this end, we find global linear editing directions in S\mathcal{S} using StyleCLIP .

Fig. 7 demonstrates that, in the domain of human faces, editing in S\mathcal{S} results in disentangled edits for both aligned and unaligned StyleGAN3 images. Particularly, notice how the image backgrounds are much better preserved compared to the W\mathcal{W}-based editing. Further, observe that the face identity is well-preserved for unrelated edits and that local edits, such as those changing hairstyle and expression, do not alter unrelated image regions (e.g., expression is consistent across the “gender”, “hi-top fade”, and “tanned” edits). Notably, this disentanglement holds for other domains such as animal faces (AFHQv2 ) and landscapes (Landscapes HQ ). When editing animals, fur color, pose, and backgrounds are well-preserved under the various edits. Additionally, altering the landscapes preserves key contents of the original image, such as the lake (top) or road (bottom).

StyleGAN3 Inversion

In this section, we address the task of inverting a pre-trained StyleGAN3 generator GG. In other words, given a target image xx, we seek a latent code w^\hat{w} that optimally reconstructs it:

where L\mathcal{L} is the L2L_{2} or LPIPS reconstruction loss.

Motivated by the goal of employing StyleGAN3 for editing real videos, solving the inversion task via a learned encoder (as opposed to latent vector optimization) may assist in achieving better temporal consistency due to its natural smoothness and bias for learning lower frequency representations . More formally, we seek to train an encoder EE over a large set of images {xi}i=1N\{x_{i}\}_{i=1}^{N} for minimizing the objective:

where E(x)E(x) encodes an input image xx into a latent code ww.

As discussed in Sec. 3, having obtained the latent code w=E(x)w=E(x), an additional transformation may be passed to the generator to control the translation and rotation of the reconstructed image y=G(w;(r,tx,ty))y=G(w;(r,t_{x},t_{y})). Finally, some latent manipulation ff may also be applied over this latent code to obtain an edited image

To enable the encoding and editing of aligned and unaligned images (such as those found in a video sequence), our inversion scheme must support the generation of both input types. A natural first attempt at doing so is to design an encoder trained on both types of images, paired with an unaligned generator. Specifically, one can employ the training schemes of existing StyleGAN2 encoders to minimize the objective given in Eq. 7 for unaligned inputs. Yet, we find that such a training scheme struggles in capturing the high variability of the unaligned facial images, resulting in poor reconstructions, see Sec. E.3 for an ablation study of such a design.

We instead choose to leverage an aligned StyleGAN3 generator and design an encoder trained solely on aligned images. As previously shown, this scheme can then be used for editing and synthesizing both aligned and unaligned imagery. In this formulation, the encoder no longer needs to correctly capture the highly-variable placement and pose of the unaligned input images. This in turn simplifies the encoder’s training objective, allowing it to instead focus on faithfully capturing the input identity and other image features.

Given the encoder trained to reconstruct aligned images, we are left with the question of how to extend this encoding scheme to support the encoding and editing of unaligned images at inference time. Assume we have a given unaligned image xunalignedx_{\textit{unaligned}}. We begin by using an off-the-shelf facial detector to detect and align the image, resulting in an aligned version of the input, denoted by xalignedx_{\textit{aligned}}. We then predict the translation (tx,ty)(t_{x},t_{y}) and rotation rr between xalignedx_{\textit{aligned}} and xunalignedx_{\textit{unaligned}} by detecting and aligning the eyes in the two images. We refer the reader to Appendix C for details on computing these parameters.

Finally, the inversion and resulting reconstruction of the unaligned input xunalignedx_{\textit{unaligned}} are given by:

Observe that while the encoder receives the aligned image xalignedx_{\textit{aligned}}, the reconstruction is able to capture the placement and rotation of the unaligned input through the use of the extracted transformation (r,tx,ty)(r,t_{x},t_{y}). As such, our inversion scheme, although trained solely on aligned images, is able to faithfully encode both aligned and unaligned images by leveraging the unique design of StyleGAN3.

Additional Details.

In practice, we employ the pSp and e4e encoders for performing the inversion task. We additionally follow the ReStyle iterative refinement scheme from Alaluf et al. to gradually refine the predicted inversion via small number of forward passes (e.g., 33) through the encoder. Please see Sec. E.1 for additional details.

2 Inverting Images into StyleGAN3

We now compare our inversion scheme introduced above to the ReStylepSp\text{ReStyle}_{pSp} and ReStylee4e\text{ReStyle}_{e4e} encoders used for inverting StyleGAN2.

As shown in Fig. 8, our StyleGAN3 encoders attain visually comparable results to their StyleGAN2 counterparts. Observe that with StyleGAN3 we are able to faithfully reproduce the input position, even when given aligned inputs, by using our landmark-based predicted transformations.

Quantitative Evaluation.

In Tab. 2 we provide a quantitative comparison between encoder-based inversion techniques for both StyleGAN2 and StyleGAN3 generators on the human facial domain. Since StyleGAN2 is limited to encoding aligned images, we perform our evaluation on the CelebA-HQ test set. In addition to the inference time required by each inversion technique, we report the L2L_{2} distance, the LPIPS distance, identity similarity , and MS-SSIM score between the reconstructions and their sources.

Our StyleGAN3 encoders reach a slightly worse performance compared to the StyleGAN2 encoders. We believe the higher difficulty in inverting StyleGAN3 is in part due to its less well-behaved W+\mathcal{W}+ latent space. This is also supported by our experiment in Sec. 4.2, where we observed that the quality of the generated images in StyleGAN3 quickly deteriorates as we move away from the W\mathcal{W} space. We believe this quick collapse of the latent space contributes to the challenge of training inversion encoders for StyleGAN3.

Editability via Latent Space Manipulation.

We now turn to evaluating the editability of our ReStyle encoders for StyleGAN3. As illustrated in Fig. 9, our Restylee4e\text{Restyle}_{e4e} encoder achieves realistic and meaningful edits while preserving the input identity. This is in contrast to RestylepSp\text{Restyle}_{pSp}, which despite achieving high-quality reconstructions, yields visibly less editable inversions. Notably, observe that in StyleGAN3, the gap in editing quality achieved by ReStylee4e\text{ReStyle}_{e4e} compared to ReStylepSp\text{ReStyle}_{pSp} is much larger compared to in StyleGAN2. This is most evident in artifacts along the hair in the second and last rows and hints at the increased importance of inverting into well-behaved latent regions compared to in StyleGAN2.

Inverting and Editing Videos

We now extend our inversion method to encoding and editing videos. This extension introduces two central challenges. First, the reconstructed and edited video frames should be temporally consistent, which may be difficult to attain when inverting each frame independently. Second, individual frames of a human face video often feature more challenging facial expressions than those found in still image training sets, (e.g., closed eyes or a mouth open mid-speech). These challenges must be addressed, regardless of the architecture. Using StyleGAN3 is appealing because it reduces texture sticking and inherently handles varying face positions and rotations. Additionally, as we demonstrate below, StyleGAN3 may be leveraged to increase the field of view, resulting in wide-view reconstructed and edited videos, rather than close-ups of an individual. This even enables the faithful edit of attributes that partially “spill out” of the input frame. Below, we describe our end-to-end video encoding and editing pipeline, summarized in Fig. 10.

Given an input video, we begin by cropping each video frame to be compatible with the input head size expected by StyleGAN3. For achieving a stable video that looks as if it was captured from a non-moving camera, we crop a fixed bounding box across all video frames (as illustrated by the red bounding box in Fig. 10). We denote the resulting cropped images by {xi,unaligned}i=1N\{x_{i,\textit{unaligned}}\}_{i=1}^{N}. As with still images, to invert each frame using our trained encoder, we additionally align each frame (green box in Fig. 10), yielding images {xi,aligned}i=1N\{x_{i,\textit{aligned}}\}_{i=1}^{N}. We also compute the transformation (ri,tx,i,ty,i)(r_{i},t_{x,i},t_{y,i}) between each (xi,aligned,xi,unaligned)(x_{i,\textit{aligned}},x_{i,\textit{unaligned}}) pair.

Initial Video Encoding.

We use our trained ReStylee4e\text{ReStyle}_{e4e} encoder EE to obtain the initial frame inversions wi=E(xi,aligned)w_{i}=E(x_{i,\textit{aligned}}), whose unaligned reconstructions are given by,

We can additionally apply some manipulation ff to obtain an edited version of the input frame:

Latent Vector Smoothing.

Inverting each frame independently may result in inconsistencies between successive reconstructed frames. This may be caused by the pre-processing alignment, the encoder network itself, or by the manipulation applied on the inverted latent codes. To mitigate temporal discontinuities, we temporally smooth the inverted, edited latent codes f( ⁣wi)f(\!w_{i}) and the predicted transformation matrix TiT_{i} applied on the Fourier features and derived by (ri,tx,i,ty,i)(r_{i},t_{x,i},t_{y,i}), using a weighted moving average:

We find that this smoothing operation improves temporal coherence without harming the reconstruction quality.

Pivotal Tuning for Improved Reconstructions.

To further improve the frame reconstructions, we adopt the pivotal tuning inversion (PTI) method . Specifically, the initial inversions are used for fine-tuning the weights of the StyleGAN3 generator to achieve better reconstructions of the input frames . Note, while the encoder network is trained to reconstruct aligned images, we perform the PTI fine-tuning using the original unaligned images. That is, when performing PTI, losses are computed between the xi,unalignedx_{i,\textit{unaligned}} images and their refined reconstructions given by:

where GPTIG_{\textrm{PTI}} is the PTI-modified generator.

Bringing It All Together.

Having obtained the smoothed edited latent codes, their corresponding smoothed transformations, and the fine-tuned generator, we can now generate the unified edited video. Formally, the final edited ii-th frame is given by:

We provide reconstruction and editing results in Fig. 11 and in Appendix G. In addition, by training StyleGAN-NADA on GPTIG_{\textrm{PTI}} for a given video and text prompt, we can generate edited videos in various styles (e.g., a cartoon video of Obama). We refer the reader to Appendices F and G for additional details and results.

Expanding the Field of View.

We now describe how we can expand the field of view (FOV) of the video reconstruction. Denote by Δ\Delta the desired expansion (illustrated in the original frame of Fig. 10). To expand the FOV, we construct a transformation matrix TΔT_{\Delta}. For example, for a vertical expansion of the frame, we define a transform matrix corresponding to a vertical shift derived from the parameters (0,0,Δ)(0,0,\Delta). For each input frame, we then generate two images:

Finally, we adjoin to yy the added non-overlapping parts from yshifty_{\textit{shift}}, obtaining the wider output frame. Results of such an expansion are shown in Fig. 12 where we demonstrate the ability to reconstruct the entirety of the individual’s head. Notice, images generated by StyleGAN2 are aligned, and as such, attributes that we wish to edit may overflow outside the frame boundary. In addition, since the images are aligned, we must project the edited frame back to the original context. Doing so, we may obtain a mismatch between regions within the generated image boundary (which were edited) and those outside the boundary (which were untouched). In StyleGAN3, however, this expansion allows editing the desired attribute in its entirety, resulting in a full, coherent edit.

Conclusions

In this work, we have explored the competence of StyleGAN3, wondering whether indeed “the third time is the charm”. We feel the answer is still unclear and more research may be required for a definite answer. On the one hand, the ability of StyleGAN3 to control the translation and rotation of generated images opens new intriguing opportunities. The prominent example explored is the generative field of view expansion, which allows one to apply StyleGAN editing on cropped video frames in a more consistent manner, alleviating the need for cumbersome and challenging seamless stitching.

On the other hand, the benefits of StyleGAN3 do come with limitations. Generally speaking, its latent space is somewhat more entangled than that of its predecessors. This makes the inversion task more challenging, affecting the robustness of frame inversions along a video. We have shown that this may be alleviated by applying the inversion on aligned images and exploiting the transformation control to compensate for the alignments. Moreover, as we have shown, training the encoder solely on aligned images does not introduce additional overheads, and even gains higher-quality synthesis.

We have naturally focused on facial images and videos. More research is required to investigate the power of StyleGAN3 for other domains. In particular, avoiding texture-sticking may be significant in videos of outdoor scenes containing high-frequency textures, like foliage or water streams. Another intriguing direction is to consider an encoder architecture that mirrors the StyleGAN3 generator, and might also employ 1 ⁣× ⁣11\!\times\!1 convolutions and Fourier features.

References

Appendix A Background and Related Works

In StyleGAN2 numerous latent spaces have extensively been explored , and were shown to be semantically-rich and disentangled. Each of the commonly-used spaces — W,W+\mathcal{W},\mathcal{W+}, and S\mathcal{S} — is disentangled by a different degree and may therefore be better suited for different tasks. The W\mathcal{W} latent space, obtained from StyleGAN’s mapping function, has been shown to be more disentangled than the original Z\mathcal{Z} space, and is therefore better suited for image editing . To edit real images, an extension of W\mathcal{W} is needed. Commonly, W+\mathcal{W+} is used for StyleGAN inversion, in which a different latent code is inserted into each layer of the synthesis network. While W\mathcal{W} offers disentangled editing capabilities, S\mathcal{S} has been shown to be even more disentangled.

A.2 Editing Images with StyleGAN2

Owing to these rich latent spaces, StyleGAN2 has been heavily studied for achieving diverse edits over images. Early works focused on fully supervised techniques using semantic labels and facial priors for discovering latent directions. To reduce supervision, others have explored both unsupervised approaches and self-supervised approaches . To achieve more fine-grained control, many works have explored the mixing of latent codes and local-based editing via semantic maps or reference images . Finally, to achieve text-based editing some have leveraged powerful contrast language-image (CLIP) models .

A.3 Inverting Images with StyleGAN2

To apply the aforementioned editing techniques on real images, one must first achieve an accurate inversion of the GAN . Inversion methods typically directly optimize the latent vector to minimize the reconstruction error of a given image , train an encoder to map an image to its latent representation , or design a hybrid approach combining both. While most techniques keep the generator fixed, recent works have proposed tuning the generator to achieve more accurate image inversions, either via a per-image optimization or a learned hypernetwork .

Appendix B Additional Analysis

Multiple works leverage the fine-tuning of StyleGAN2 generators for various applications, such as image-to-image translation and inversion . These works show that under certain conditions, the fine-tuned generator (child) faithfully preserves key properties of the original generator (parent), due to the alignment between their latent spaces. In other words, the semantics of latent codes and directions in the latent spaces are often unchanged. Here, we examine whether this preservation holds true in StyleGAN3. As shown in Fig. 13, one can use the same editing directions on various generators fine-tuned with StyleGAN-NADA , indicating that this latent space alignment is indeed retained. Additional generations and editing results obtained with StyleGAN-NADA are illustrated in Figs. 21 and 22, respectively.

B.2 Disentanglement Analysis of 𝒲+limit-from𝒲\mathcal{W}+

\mathcal{W}+ In Sec. 4.2, we performed a quantitative analysis of the disentanglement of the different latent spaces of StyleGAN3 and computed the DCI metrics for each space. Recall, disentanglement measures the extent to which each latent channel controls a single attribute, while completeness measures the degree to which each attribute is controlled by a single latent channel. Finally, informativeness assesses the accuracy of the attribute classifiers for a given latent representation. It is of interest to quantify the disentanglement of the W+\mathcal{W}+ latent space, often used for inversion. However, we find that randomly generated images in W+\mathcal{W}+ are unnatural, making the computation of the DCI metric unreliable. In Fig. 14, we compare a collection of uncurated samples from W+\mathcal{W}+ for both StyleGAN2 and StyleGAN3 generators trained on aligned facial images. As shown, while both sets of images are unrealistic, those of StyleGAN3 contain significantly more artifacts. As such, we choose to omit the computation of the DCI metrics of W+\mathcal{W}+.

Appendix C Landmarks-Based Transformations

As described in Sec. 6, our StyleGAN3 encoders are trained solely on aligned images. To invert a given unaligned image, we first align the image and invert the resulting image using the encoder. We then generate the unaligned image reconstruction by utilizing the user-defined transformation passed to the synthesis network along with the inverted latent code.

To compute the transformation triplet (r,tx,ty)(r,t_{x},t_{y}) for a given unaligned image xunalignedx_{unaligned}, we employ an off-the-self landmark detection tool and build on the landmark parsing procedure used for creating the FFHQ dataset. Specifically, we first align the image to obtain an aligned version of the input, denoted by xalignedx_{aligned}.

We then detect the eyes of both xunalignedx_{unaligned} and xalignedx_{aligned}, and compute the rotation and the translation between the two sets. Here, the rotation is given by the angle between the lines that connect the eyes in both images. For computing translation, we rotate the aligned image and measure the vertical and horizontal distances between the left eye of the rotated aligned image and unaligned image. These three values explicitly define the user-specified transformations that are passed to the Fourier features of StyleGAN3’s synthesis network. This process is illustrated in Fig. 15.

It should be noted that not all unaligned images can be inverted. As the unaligned FFHQ StyleGAN3 generator was trained on unaligned images of a certain facial size, we can only invert faces of this size. Therefore, before computing the landmarks for a given image, we first crop it so that the face is of the size suited for StyleGAN3. This process is done similarly to the process for creating the official FFHQ-U dataset.

Appendix D Editing: Additional Details

For editing images in W\mathcal{W}, we use the official implementation of InterFaceGAN for training the linear boundaries using off-the-shelf classifiers. We apply HopeNet for pose, Rothe et al. for age, and the classifier from Lin et al. for the remaining attributes.

Appendix E StyleGAN3 Encoding Scheme

Our encoders for inverting StyleGAN3 are based on the pSp and e4e encoding schemes. We apply the same encoder architectures as those used for inverting StyleGAN2 generators. Following our insights that a pre-trained aligned StyleGAN3 generator can synthesize both aligned and unaligned images, our encoders are trained solely on aligned images, significantly simplifying the training objective of the encoder.

Due to the larger memory consumption required by StyleGAN3, our encoders are trained using a batch size of 22. To match the batch size used in the official implementations of the StyleGAN2 pSp and e4e encoders, we apply gradient accumulation to attain an effective batch size of 88 (i.e., an optimization step is performed every four batches). All encoders are trained using a single NVIDIA P40 GPU.

Training is performed using the same set of losses as used to train the StyleGAN2 encoders. Specifically, we use a weighted combination of the L2L_{2} pixel-wise loss, the LPIPS perceptual loss, and an identity-based reconstruction . The overall loss objective is given by:

where we set λl2=1\lambda_{l2}=1, λlpips=0.8\lambda_{lpips}=0.8, and λid=0.1\lambda_{id}=0.1.

For training the ReStylee4e\text{ReStyle}_{e4e}, encoder we additionally remove the progressive training scheme used in the official implementation. Instead, all 1616 latent codes are predicted simultaneously by the encoder from the start of training. We find this leads to faster convergence.

E.2 Baselines and Comparisons

In our work we compare our ReStylepSp\text{ReStyle}_{pSp} and ReStylee4e\text{ReStyle}_{e4e} StyleGAN3 encoders with their StyleGAN2 counterparts from Alaluf et al. . All encoders are trained on the aligned FFHQ dataset consisting of 70,00070,000 images. For a quantitative comparison of the encoders, we compute the reconstruction metrics on the aligned CelebA-HQ test set. Finally, for our qualitative images, we display the aligned outputs for StyleGAN2 encoders and the unaligned outputs for the StyleGAN3 encoders.

E.3 Ablation Study

If one were to design an encoder for encoding unaligned images into StyleGAN3’s latent space, a natural first attempt at doing so would be to train an encoder on unaligned images using existing encoding schemes for inverting an unaligned generator. Specifically, assume we have pairs of images {(xalignedi, xunalignedi)}i=1N\{(x^{i}_{aligned},~{}x^{i}_{unaligned})\}_{i=1}^{N}, we can train the encoder to solve the following objective:

where GG is a pre-trained unaligned generator.

Yet, an immediate challenge arises: employing the identity loss, which incorporates a pre-trained facial recognition network , is non-trivial. This facial recognition network is trained on images that are aligned and cropped to the inner facial region. Hence, applying this network on unaligned images may lead to unpredictable results. One may mitigate this by using off-the-shelf facial detectors to detect and align the images, but applying these networks during training is impractical. In Richardson et al. , the authors overcome this challenge by using a heuristic that roughly crops the face before passing the image through the facial recognition network. Yet, they assume the images are pre-aligned. When training on unaligned images, as in our case, this heuristic is no longer applicable.

To overcome this challenge we can perform the pseudo-alignment trick described in Sec. 3. Consider the latent code w=(w0,w1,...,w15)w=(w_{0},w_{1},...,w_{15}) outputted by our encoder for some unaligned input xunalignedx_{unaligned}. We can replace the first latent code w0w_{0} with the generator’s average latent code to obtain the pseudo-aligned latent representation waligned=(w‾,w1,...,w15)w_{aligned}=(\overline{w},w_{1},...,w_{15}), corresponding to the pseudo-aligned image yaligned=G(waligned;(0,0,0))y_{aligned}=G(w_{aligned};(0,0,0)).

Given the pseudo-aligned image, we are now more accurately able to compute the identity loss between xalignedx_{aligned} and yalignedy_{aligned}. Additionally, we may compute the L2L_{2} and LPIPS reconstruction losses between the original the original unaligned image xunalignedx_{unaligned} and the unaligned reconstruction yunaligned=G(w;(0,0),0)y_{unaligned}=G(w;(0,0),0).

Qualitative Comparisons.

We now compare the unaligned encoding scheme above to our aligned scheme presented in Sec. 6. As illustrated in Fig. 16, our scheme achieves superior reconstructions compared to the unaligned version. We attribute this improvement to the simpler training task of our aligned encoder: rather than needing to capture both the input identity and position, our encoder can focus on reconstructing only the former with the desired pose provided via the user-defined transformations predicted using the procedure described in Appendix C.

Appendix F Video Inversion Scheme: Additional Details

As mentioned in Sec. 7, before feeding a given frame to our encoder, we align and crop the frame. The landmark detector used for aligning the image and computing the Fourier features transformations uses either the distance between the two eyes or the distances between the eyes and the mouth. Since the distance between the eyes and mouth may change along the video, we choose to use only the distance between the eyes for all frames for this. We find that doing so gives a slightly move stable video reconstruction. Note, that the cropping procedure assumes the distance between the camera and the input face does not change along the video.

F.2 Pivotal Tuning of Videos

For inverting and editing a given video we perform a per-video fine-tuning of the StyleGAN3 generator network using the pivotal tuning technique from Roich et al. . For each video, training is performed for a total of 8,0008,000 optimization steps with a batch size of 22 using the L2L_{2} pixel-wise loss and the LPIPS loss, both with equal weight coefficients. For example, given a video consisting of 200200 frames, each frame is observed an average of 4040 times during training. During training, we do not alter the weights of the input Fourier features layer of the generator.

F.3 Latent Vector Smoothing

We refer the reader to the accompanying video results for a comparison of video inversions and reconstructions with and without the latent vector smoothing operation. As can be seen, when no latent smoothing is performed, the resulting video is unstable due to the pre-alignment step and the per-frame inversions of the encoder.

F.4 Field-of-View Expansion

In the main paper, we describe how to expand the field of view when reconstructing and editing a given input frame. In the overview example, we performed an expansion toward a single direction (e.g., extending the top of the video). Yet, it is also possible to extend the field-of-view toward multiple directions. For each direction we wish to expand the image in, we generate another image shifted toward that direction. For example, for expanding an image to the right by Δ\Delta uses the transformations parameters (0,−Δ,0)(0,-\Delta,0) while expanding an image at the bottom uses parameters (0,0,−Δ)(0,0,-\Delta). Given the generated image for each direction, we then copy the non-overlapping parts from the sifted images and join them with the original reconstructed image. Note, for cases where we wish to expand an image both horizontally and vertically, we add an additional shift in both directions for filling in the corner regions. We demonstrate results of this field-of-view expansion in Fig. 28.

Appendix G Additional Results

We provide additional results and comparisons, as follows:

Fig. 17 demonstrates non-linear editing performed in the W+\mathcal{W}+ latent space of an aligned StyleGAN3 generator using StyleCLIP’s mapping technique .

Figs. 18, 19 and 20 illustrate edits obtained across various domains using StyleCLIP’s global editing technique applied in StyleGAN3’s S\mathcal{S} space.

In Fig. 21, we demonstrate domain adaptation results obtained by fine-tuning pre-trained StyleGAN3 generators using StyleGAN-NADA across various domains.

Fig. 22 shows editing results obtained over various fine-tuned child generators showing that the alignment of latent spaces is preserved under the fine-tuning of StyleGAN3.

Figs. 23 and 24 provide additional reconstruction comparisons between our StyleGAN3 encoders and StyleGAN2 encoders from Alaluf et al. .

Fig. 25 shows additional editing results obtained over real images using the editing techniques from Shen et al. and Patashnik et al. .

Figs. 26 and 27 show video reconstruction and editing results obtained with our full StyleGAN3 encoding pipeline. In addition, we demonstrate domain adaptation results applied over the edited videos, allowing us to generate edited videos in various styles.

Fig. 28 demonstrates our field-of-view expansion technique on multiple video sequences.

Finally, we invite the reader to visit our project page where we provide full video reconstructions and edits on various inputs.