HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image Editing

Yuval Alaluf, Omer Tov, Ron Mokady, Rinon Gal, Amit H. Bermano

Introduction

Generative Adversarial Networks (GANs) , and in particular StyleGAN have become the gold standard for image synthesis. Thanks to their semantically rich latent representations, many works have utilized these models to facilitate diverse and expressive editing through latent space manipulations . Yet, a significant challenge in adopting these approaches for real-world applications is the ability to edit real images. For editing a real photo, one must first find its corresponding latent representation via a process commonly referred to as GAN inversion . While the inversion process is a well-studied problem, it remains an open challenge.

Recent works have demonstrated the existence of a distortion-editability trade-off: one may invert an image into well-behaved regions of StyleGAN’s latent space and attain good editability. However, these regions are typically less expressive, resulting in reconstructions that are less faithful to the original image. Recently, Roich et al. showed that one may side-step this trade-off by considering a different approach to inversion. Rather than searching for a latent code that most accurately reconstructs the input image, they fine-tune the generator in order to insert a target identity into well-behaved regions of the latent space. In doing so, they demonstrate state-of-the-art reconstructions while retaining a high level of editability. Yet, this approach relies on a costly per-image optimization of the generator, requiring up to a minute per image.

A similar time-accuracy trade-off can be observed in classical inversion approaches. On one end of the spectrum, latent vector optimization approaches achieve impressive reconstructions, but are impractical at scale, requiring several minutes per image. On the other end, encoder-based approaches leverage rich datasets to learn a mapping from images to their latent representations. These approaches operate in a fraction of a second but are typically less faithful in their reconstructions.

In this work, we aim to bring the generator-tuning technique of Roich et al. to the realm of interactive applications by adapting it to an encoder-based approach. We do so by introducing a hypernetwork that learns to refine the generator weights with respect to a given input image. The hypernetwork is composed of a lightweight feature extractor (e.g., ResNet ) and a set of refinement blocks, one for each of StyleGAN’s convolutional layers. Each refinement block is tasked with predicting offsets for the weights of the convolutional filters of its corresponding layer. A major challenge in designing such a network is the number of parameters comprising each convolutional block that must be refined. Naïvely predicting an offset for each parameter would require a hypernetwork with over three billion parameters. We explore several avenues for reducing this complexity: sharing offsets between parameters, sharing network weights between different hypernetwork layers, and an approach inspired by depthwise-convolutions . Lastly, we observe that reconstructions can be further improved through an iterative refinement scheme which gradually predicts the desired offsets over a small number of forward passes through the hypernetwork. By doing so, our approach, HyperStyle, essentially learns to “optimize” the generator in an efficient manner.

The relation between HyperStyle and existing generator-tuning approaches can be viewed as similar to the relation between encoders and optimization inversion schemes. Just as encoders find a desired latent code via a learned network, our hypernetwork efficiently finds a desired generator with no image-specific optimization.

We demonstrate that HyperStyle achieves a significant improvement over current encoders. Our reconstructions even rival those of optimization schemes, while being several orders of magnitude faster. We additionally show that HyperStyle preserves the appealing structure and semantics of the original latent space, allowing one to leverage off-the-shelf editing techniques on the resulting inversions, see Fig. 1. Finally, we show that HyperStyle generalizes well to out-of-domain images, such as paintings and animations, even when unobserved during the training of the hypernetwork itself. This hints that the hypernetwork does not only learn to correct specific flawed attributes, but rather learns to refine the generator in a more general sense.

Background and Related Work

Introduced by Ha et al. , hypernetworks are neural networks tasked with predicting the weights of a primary network. By training a hypernetwork over a large data collection, the primary network’s weights are adjusted with respect to specific inputs, yielding a more expressive model. Hypernetworks have been applied to a wide range of applications including semantic segmentation , 3D modeling , neural architecture search , and continual learning , among others.

Latent Space Manipulation

A widely explored application for generative models is their use for the editing of real images. Considerable effort has gone into leveraging StyleGAN for such tasks, owing to its highly-disentangled latent spaces. Many methods have been proposed for finding semantic latent directions using varying levels of supervision. These range from full-supervision in the form of semantic labels and facial priors to unsupervised approaches . Others have explored self-supervised approaches , the mixing of latent codes to produce local edits , and the use of contrastive language-image (CLIP) models to achieve new editing capabilities . Applying these methods to real images requires one to first perform an accurate inversion of the given image.

GAN Inversion

GAN inversion is the process of obtaining a latent code that can be passed to the generator to reconstruct a given image. Generally, inversion methods either directly optimize the latent vector to minimize the error for a given image , train an encoder over a large number of samples to learn a mapping from an image to its latent representation , or use a hybrid approach combining both . Among encoder-based methods, Alaluf et al. iteratively refine the predicted latent code through a small number of forward passes through the network. Our work adopts this idea and applies it to the generator weight offsets predicted by the hypernetwork. Finally, in a concurrent work, Dinh et al. also explore the use of hypernetworks for achieving higher fidelity inversions.

Distortion-Editability

Typically, latent traversal and inversion methods concern themselves with one of two spaces: W\mathcal{W}, obtained via StyleGAN’s mapping network and W+\mathcal{W}+, where each layer of the generator is assigned a different latent code wi∈Ww_{i}\in\mathcal{W}. Images inverted into W\mathcal{W} show a high degree of editability: they can be modified through latent space traversal with minimal corruption. However, W\mathcal{W} offers poor expressiveness, limiting the range of images that can be faithfully reconstructed. Therefore, many prior works invert into the extended W+\mathcal{W}+ space, achieving reduced distortion at the cost of inferior editability. Tov et al. suggest balancing the two by designing an encoder that predicts codes in W+\mathcal{W}+ residing close to W\mathcal{W}. Others have explored similar ideas for optimization .

Generator Tuning

To leverage the visual quality of a pre-trained generator, most works avoid altering the generator weights when performing the inversion. Nonetheless, some works have explored performing a per-image tuning of the generator to obtain more accurate inversions. Pan et al. invert BigGAN by randomly sampling noise vectors, selecting the one that best matches the real image, and optimizing it simultaneously with the generator weights in a progressive manner. Roich et al. and Hussien et al. invert images into a pre-trained GAN by first recovering a latent code which approximately reconstructs the target image and then fine-tuning the generator weights for improve image-specific details. Bau et al. explored the use of a neural network to predict feature modulations to improve GAN inversion. However, the aforementioned works require a lengthy optimization for every input, typically requiring minutes per image. As such, these methods are often inapplicable to real-world scenarios at scale. In contrast, we train a hypernetwork over a large set of images, resulting in a single network used to refine the generator for any given image. Importantly, this is achieved in near real-time and is more suitable for interactive settings.

Method

When solving the GAN inversion task, our goal is to identify a latent code that minimizes the reconstruction distortion with respect to a given target image xx:

where G(w;θ)G(w;\theta) is the image produced by a pre-trained generator GG parameterized by weights θ\theta, over the latent ww. L\mathcal{L} is the loss objective, usually L2L_{2} or LPIPS . Solving Eq. 1 via optimization typically requires several minutes per image. To reduce inference times, an encoder EE can be trained over a large set of images {xi}i=1N\{x^{i}\}_{i=1}^{N} to minimize:

This results in a fast inference procedure w^=E(x)\hat{w}=E(x). A latent manipulation ff can then be applied over the inverted code w^\hat{w} to obtain an edited image G(f(w^);θ)G(f(\hat{w});\theta).

Recently, Roich et al. propose injecting new identities into the well-behaved regions of StyleGAN’s latent space. Given a target image, they use an optimization process to find an initial latent w^init∈W\hat{w}_{init}\in\mathcal{W} leading to an approximate reconstruction. This is followed by a fine-tuning session where the generator weights are adjusted so that the same latent better reconstructs the specific image:

where θ^\hat{\theta} represents the new generator weights. The final reconstruction is obtained by utilizing the initial inversion and altered weights: y^=G(w^init;θ^)\hat{y}=G(\hat{w}_{init};\hat{\theta}).

2 Overview

Our method HyperStyle aims to perform the identity-injection operation by efficiently providing modified weights for the generator, as illustrated in Fig. 2. We begin with an image xx, a generator GG parameterized by weights θ\theta, and an initial inverted latent code w^init∈W\hat{w}_{init}\in\mathcal{W}. Using these weights and w^init\hat{w}_{init}, we generate the initial reconstructed image y^init=G(w^init;θ)\hat{y}_{init}=G(\hat{w}_{init};\theta). To obtain such a latent code we employ an off-the-shelf encoder .

Our goal is to predict a new set of weights θ^\hat{\theta} that minimizes the objective defined in Eq. 3. To this end, we present our hypernetwork HH, tasked with predicting these weights. To assist the hypernetwork in inferring the desired modifications, we pass as input both the target image xx and the initial, approximate image reconstruction y^init\hat{y}_{init}. The predicted weights are thus given by: θ^=H(y^init,x)\hat{\theta}=H\left(\hat{y}_{init},x\right). We train HH over a large collection of images with the goal of minimizing the distortion of the reconstructions:

Given the hypernetwork predictions, the final reconstruction can be obtained as y^=G(w^init;θ^)\hat{y}=G(\hat{w}_{init};\hat{\theta}).

Owing to the reconstruction-editability trade-off outlined in Sec. 2, the initial latent code should reside within the well-behaved (i.e., editable) regions of StyleGAN’s latent space. To this end, we employ a pre-trained e4e encoder into W\mathcal{W} that is kept fixed throughout the training of the hypernetwork. As shall be shown, by tuning around such a code, one can apply the same editing techniques as used with the original generator.

In practice, rather than directly predicting the new generator weights, our hypernetwork predicts a set of offsets with respect to the original weights. In addition, we follow ReStyle and perform a small number of passes (e.g., 55) through the hypernetwork to gradually refine the predicted weight offsets, resulting in higher-fidelity inversions.

In a sense, one may view HyperStyle as learning to optimize the generator, but doing so in an efficient manner. Moreover, by learning to modify the generator, HyperStyle is given more freedom to determine how to best project an image into the generator, even when out of domain. This is in contrast to standard encoders which are restricted to encoding into existing latent spaces.

3 Designing the HyperNetwork

The StyleGAN generator contains approximately 3030M parameters. On one hand, we wish our hypernetworks to be expressive, allowing us to control these parameters for enhancing the reconstruction. On the other hand, control over too many parameters would result in an inapplicable network requiring significant resources for training. Therefore, the design of the hypernetwork is challenging, requiring a delicate balance between expressive power and the number of trainable parameters involved.

To further reduce the number of trainable parameters, we introduce a Shared Refinement Block, inspired by the original hypernetwork . These output heads consist of independent convolutional layers used to down-sample the input feature map. They are then followed by two fully-connected layers shared across multiple generator layers, as illustrated in Fig. 3. Here, the fully-connected weights are shared across the non-toRGB layers with dimension 3×3×512×5123\times 3\times 512\times 512, i.e., the largest generator convolutional blocks. As demonstrated in Ha et al. this allows for information sharing between the output heads, yielding improved reconstruction quality. Detailed layouts of the Refinement Blocks are given in Appendix D.

Combining the Shared Refinement Blocks and per-channel predictions, our final configuration contains 2.72.7B fewer parameters (~89%89\%) than a naïve hypernetwork. We summarize the total number of parameters of different hypernetwork variants in Tab. 1. We refer the reader to Sec. 4.3 where we validate our design choices and explore additional avenues for reducing the number of parameters.

The choice of which layers to refine is of great importance. It allows us to reduce the output dimension while focusing the hypernetwork on the more meaningful generator weights. Since we invert one identity at a time, any changes to the affine transformation layers can be reproduced by a respective re-scaling of the convolution weights. Moreover, we find that altering the toRGB layers harms the editing capabilities of the GAN. We hypothesize that modifying these layers mainly alters the pixel-wise texture and color , changes that do not translate well under global edits such as pose (see Appendix B for examples). Therefore, we restrict ourselves to modifying only the non-toRGB convolutions.

Lastly, we follow Karras et al. and split the generator layers into three levels of detail — coarse, medium, fine — each controlling different aspects of the generated image. As the initial inversions tend to capture coarse details, we further restrict our hypernetwork to output offsets for the medium and fine generator layers.

4 Iterative Refinement

To further improve the inversion quality, we adopt the iterative refinement scheme suggested by Alaluf et al. . This enables us to perform several passes through our hypernetwork for a single image inversion. Each added step allows the hypernetwork to gradually refine its predicted weight offsets, resulting in stronger expressive power and a more accurate inversion.

We perform TT passes. For the first pass, we use the initial reconstruction y^0=G(w^init;θ)\hat{y}_{0}=G(\hat{w}_{init};\theta). For each refinement step t≥1t\geq 1, we predict a set of offsets Δt=H(y^t−1,x)\Delta_{t}=H(\hat{y}_{t-1},x) used to obtain the modified weights θ^t\hat{\theta}_{t} and updated reconstruction y^t=G(w^init;θ^t)\hat{y}_{t}=G(\hat{w}_{init};\hat{\theta}_{t}). The weights at step tt are defined as the accumulated modulation across all previous steps:

The number of refinement steps is set to T=5T=5 during training. Following Alaluf et al. we compute the losses at each refinement step. Note, w^init\hat{w}_{init} remains fixed during the iterative process. The final inversion y^\hat{y} is the reconstruction obtained at the last step.

5 Training Losses

Similar to encoder-based methods, our training is guided by an image-space reconstruction objective. We apply a weighted combination of the pixel-wise L2L_{2} loss and LPIPS perceptual loss . For the facial domain, we further apply an identity-based similarity loss by employing a pre-trained facial recognition network to preserve the facial identity. As suggested by Tov et al. , we apply a MoCo-based similarity loss for non-facial domains. The final loss objective is given by:

Experiments

For the human facial domain we use FFHQ for training and the CelebA-HQ test set for quantitative evaluations. On the cars domain, we use the Stanford Cars dataset . Additional results on AFHQ Wild are provided in Appendix F. We compare our results to the state-of-the-art encoders pSp , e4e , and ReStyle applied over both pSp and e4e. A visual comparison with IDInvert is provided in Appendix F. For a comparison with optimization techniques, we compare to PTI and the latent vector optimization into W+\mathcal{W}+ from Karras et al. .

1 Reconstruction Quality

We begin with a qualitative comparison, provided in Fig. 4. While optimization techniques are typically able to achieve accurate reconstructions, they come with a high computational cost. HyperStyle offers visually comparable results with an inference time several orders of magnitude faster. Furthermore, PTI may struggle when inverting a low-resolution input (2nd row), yielding a blurred reconstruction due to its inherent design of over-fitting to the target image. Our hypernetwork, meanwhile, is trained on a large image collection and is therefore less likely to re-create such resolution-based artifacts. In addition, compared to single-shot encoders (pSp and e4e), HyperStyle better captures the input identity (3rd row). When compared to the more recent ReStyle encoders, HyperStyle is still able to better reconstruct finer details such as complex hairstyles (1st row) and clothing (2nd row).

Quantitative Evaluation

In Tab. 2, we present a quantitative evaluation focusing on the time-accuracy trade-off. Along with the inference time of each method, we report the pixel-wise L2L_{2} distance, the LPIPS distance, and the MS-SSIM score between each reconstruction and source. We additionally measure identity similarity using a pre-trained facial recognition network . For HyperStyle and ReStyle , we performed multiple iterative steps until the metric scores stopped improving or until 1010 iterations were reached. For optimization, we use at most 1,5001,500 steps, while for PTI we perform at most 350350 pivotal tuning steps.

As presented in Tab. 2, HyperStyle’s performance consistently surpasses that of the encoder-based methods. Surprisingly, it even achieves results on par with the StyleGAN2 optimization , while being nearly 200200 times faster. Overall, HyperStyle demonstrates optimization-level reconstructions achieved with encoder-like inference times.

2 Editability via Latent Space Manipulations

Good inversion methods should provide not only meticulous reconstructions but also highly-editable latent codes. We thus evaluate the editability of our produced inversions. We do so by analyzing two key aspects. One aspect of interest is the range of modifications that an inverted latent can support (e.g., how much the pose can be changed). The other is how well the identity is preserved along this range.

As shown in Fig. 5, our method successfully achieves realistic and meaningful edits, while being faithful to the input identity. The inversions of optimization, pSp, and ReStylepSp\textit{ReStyle}_{pSp} reside in poorly-behaved latent regions of W+\mathcal{W}+. Therefore, their editing is less meaningful and introduces significant artifacts. For instance, in the cars domain, they struggle in making notable changes to the car color and shape. In the 4th row, these methods fail to either preserve the original identity or perform a full frontalization. On the other end of the reconstruction-editability trade-off, e4e and ReStylee4e\textit{ReStyle}_{e4e} are more editable but cannot faithfully preserve the original identity, as demonstrated in the 3rd and 4th rows. In contrast, HyperStyle and PTI, which invert into the well-behaved W\mathcal{W} space, are more robust in their editing capabilities while successfully retaining the original identity. Yet, HyperStyle requires a significantly lower inference overhead to achieve these results.

Quantitative Evaluation

Comparing the editability of inversion methods is challenging since applying the same editing step size to latent codes obtained with different methods results in different editing strengths. This would introduce unwanted bias to the identity similarity measure, as the less-edited images may tend to be more similar to the source. To address this, we edit using a range of various step sizes and plot the measured identity similarity along this range, resulting in a continuous similarity curve for each inversion method. This allows us to validate the identity preservation with respect to a fixed editing magnitude, as well as examine the range of edits supported. Ideally, an inversion method should achieve high identity similarity across a wide range of editing strengths. We measure the editing magnitude using trait-specific classifiers (HopeNet for pose and the classifier from Lin et al. for smile extent). As before, identity similarity is measured using the CurricularFace method .

As can be seen in Fig. 6, HyperStyle consistently outperforms other encoder-based methods in terms of identity preservation while supporting an equal or greater editing range. Compared to optimization-based techniques, HyperStyle achieves similar identity preservation and editing range yet does so substantially faster.

These results highlight the appealing nature of HyperStyle. With respect to other encoders, HyperStyle achieves superior reconstruction quality while providing strong editability and fast inference. Additionally, compared to optimization techniques, HyperStyle achieves comparable reconstruction and editability at a fraction of the time, making it more suitable for real-world use at scale. This places HyperStyle favorably on both the reconstruction-editability and the time-accuracy trade-off curves.

3 Ablation Study

We now validate the design choices described in Sec. 3. Results are summarized in Tab. 3. First, we investigate the choice of layers refined by the hypernetwork. We observe that training only the medium and fine non-toRGB layers achieves comparable performance, a slimmer network, and faster inference. Notably, we also find that altering toRGB layers may harm editability. Second, we find the iterative scheme to be more accurate with fewer artifacts. Finally, we validate the effectiveness of the Shared Refinement Block and the information sharing it provides. Visual comparisons of all ablations can be found in Appendix B.

Our final configuration uses shared offsets for each convolutional kernel. An important question is whether this constrains the network too strongly. To answer this, we design an alternative refinement head, inspired by separable convolutions . Rather than predicting offsets for an entire k×k×Cin×Coutk\times k\times C^{in}\times C^{out} filter in one step, we decompose it into two slimmer predictions: k×k×Cin×1k\times k\times C^{in}\times 1 and k×k×1×Coutk\times k\times 1\times C^{out}. The final offset block is then given by their product. This allows us to predict an offset for every parameter of the kernel, potentially increasing the network’s expressiveness. We observe (Tab. 3) that the increased flexibility of predicting an offset per parameter does not improve reconstruction, indicating that simpler, per-channel predictions are sufficient.

4 Additional Applications

Many works have explored fine-tuning a pre-trained StyleGAN towards semantically similar domains. This process maintains a correspondence between semantic attributes in the two latent spaces, allowing translation between domains . Yet, some features, such as facial hair or hair color, may be lost during this translation. To address this, we use HyperStyle trained on the source generator to modify the fine-tuned target generator. Namely, given an input image, we can take the weight offsets predicted with respect to the source generator and apply them to the target generator. The image in the new domain is then obtained by passing the image’s original latent code to the modified target generator.

Fig. 7 shows examples of applying weight offsets over various fine-tuned generators. As shown, when no offsets are applied, important details are lost. However, HyperStyle leads to more faithful translations preserving identity without harming the target style. Importantly, the translations are attained with no domain-specific hypernetwork training.

Editing Out-of-Domain Images

To this point, we have discussed handling images from the same domain as used for training. If our hypernetwork has indeed learned to generalize, it should not be sensitive to the domain of the input. As may be expected, standard encoders cannot handle out-of-domain images well (see Fig. 8 for e4e and Appendix F for others). By adjusting the pre-trained generator towards a given out-of-domain input, HyperStyle enables editing diverse images, without explicitly training a new generator on their domain. This points to improved expressiveness and generalization. It seems the hypernetwork does not just fix poorly reconstructed attributes but learns to adapt the generator in a more general sense. We find these results to be a promising direction for manipulating out-of-domain images without having to train new generators or perform lengthy per-image tuning.

Conclusions

We introduced HyperStyle, a novel approach for StyleGAN inversion. We leverage recent advancements in hypernetworks to achieve optimization-level reconstructions at encoder-like inference times. In a sense, HyperStyle learns to efficiently optimize the generator for a given target image. Doing so mitigates the reconstruction-editability trade-off and enables the effective use of existing editing techniques on a wide range of inputs. In addition, HyperStyle generalizes surprisingly well, even to out-of-domain images neither the hypernetwork nor the generator have seen during training. Looking forward, further broadening generalization away from the training domain is highly desirable. This includes robustness to unaligned images and unstructured domains. The former may potentially be addressed through StyleGAN3 while the latter would probably warrant training on a richer set of images. In summary, we believe this approach to be an essential step towards interactive and semantic in-the-wild image editing and may open the door for many intriguing real-world scenarios.

Acknowledgements

We would like to thank Or Patashnik, Elad Richardson, Daniel Roich, Rotem Tzaban, and Oran Lang for their early feedback and discussions. This work was supported in part by Len Blavatnik and the Blavatnik Family Foundation.

References

Appendix A Broader Impact

HyperStyle enables accurate and highly editable inversions of real images. While our tool aims to empower content creators, it can also be used to generate more convincing deep-fakes and aid in the spread of disinformation . However, powerful tools already exist for the detection of GAN-synthesized imagery , for example through frequency analysis . These tools continually evolve, which gives us hope that any potential misuse of our method can be mitigated.

Another cause for concern is the bias that generative networks inherit from their training data . Our model was similarly trained on such a biased set, and as a result, may display degraded performance when dealing with images from minority classes . However, we have demonstrated that our model successfully generalizes beyond its training set, and allows us to similarly shift the GAN beyond its original domain. These properties allow us to better preserve minority traits when compared to prior works, and we hope that this benefit would similarly enable fairer treatment of minorities in downstream tasks.

Appendix B Ablation Study: Qualitative Comparisons

In Sec. 4.3 of the main paper, we presented a quantitative ablation study to validate the design choices of our hypernetworks. We now turn to provide visual comparisons to complement this. First, we illustrate the effectiveness of the iterative refinement scheme in Fig. 9. Observe that iteratively predicting the weight offsets results in sharper reconstructions. This is best reflected in the preservation of fine details, most notably in the reconstruction of hairstyle and facial hair. In Fig. 10, we show that altering the toRGB convolutional layers harms the editability of the resulting inversions. This is most noticeable in edits requiring global changes, such as pose and age. For instance, altering the pose in the second row results in blurred edits. Additionally, when modifying age in the bottom two rows, HyperStyle succeeds in realistically altering the clothing (33rd row) and hairstyle (44th row).

Appendix C Additional Quantitative Results

Following the quantitative reconstruction metrics provided in the main paper on the human facial domain, we provide quantitative results on the cars domain and wild animals domain in Tabs. 4 and 5.

Appendix D The HyperStyle Architecture

In addition to the standard Refinement Block, we introduce a Shared Refinement Block that is shared between multiple hypernetwork layers. These Shared Refinement Blocks make use of two fully-connected layers whose weights are shared between different output heads. The first fully-connected layer transforms the 1×1×5121\times 1\times 512 tensor to a 512×512512\times 512 intermediate representation. This is followed by a per-channel fully-connected layer which maps each 1×5121\times 512 channel to a 1×1×5121\times 1\times 512 tensor, resulting in the final 1×1×512×5121\times 1\times 512\times 512 dimensional offsets. We apply these shared blocks to all generator layers with a convolutional dimension of 3×3×512×5123\times 3\times 512\times 512.

Appendix E The StyleGAN2 Architecture

To determine which network parameters are most crucial to our inversion goal, it is important to understand the overall function of the components where these parameters reside. In the case of StyleGAN2 , we consider four key components, as illustrated in Fig. 11. First, a mapping network converts the initial latent code z∼N(0,1)512z\sim\mathcal{N}\left(0,1\right)^{512}, into an equivalent code in a learned latent space w∈Ww\in\mathcal{W}. These codes are then fed into a series of affine transformation blocks, one for each of the network’s convolutional layers, which in turn predict a series of factors used to modulate the convolutional kernel weights.

Lastly, the generator itself is built from two types of convolutional blocks: feature space convolutions, which learn increasingly complex representations of the data in some high-dimensional feature space, and toRGB blocks, which utilize convolutions to map these complex representations into residuals in a planar (or RGB) space. Of these four components, we restrict ourselves to modifying only the feature-space convolutions. We outline the different generator layers and their dimensions in Tab. 8.

Appendix F Additional Qualitative Results

Finally, we provide additional results and comparisons, as follows:

Fig. 12 provides additional reconstruction comparisons on the human facial domain.

Fig. 13 contains a visual comparison between HyperStyle and IDInvert on the human facial domain.

Fig. 14 provides reconstruction comparisons on the cars domain.

Fig. 15 shows additional editing comparisons on the human facial domain obtained with StyleCLIP and InterFaceGAN .

Fig. 16 shows additional HyperStyle editing results on the human facial domain obtained with InterFaceGAN .

Fig. 17 contains additional HyperStyle editing results on the human facial domain obtained with StyleCLIP .

Fig. 18 provides additional editing comparisons on the cars domain obtained with GANSpace .

Fig. 19 contains HyperStyle editing results on the cars domain obtained with GANSpace .

Fig. 20 contains HyperStyle editing results obtained with StyleCLIP on the AFHQ Wild test set.

Fig. 21 and Fig. 22 illustrate HyperStyle’s reconstructions and edits on challenging out-of-domain images.

Fig. 23 illustrates additional domain adaptation results.

Appendix G Implementation Details

All hypernetworks employ a ResNet34 backbone pre-trained on ImageNet. The networks have a modified input layer to accommodate the 66-channel inputs. We train our networks using the Ranger optimizer with a constant learning rate of 0.00010.0001 and a batch size of 88.

When applying the iterative refinement scheme from Alaluf et al. , our hypernetworks use T=5T=5 iterative steps per batch during training. For each step tt, we compute losses between the current reconstructions and inputs. That is, losses are computed TT times per batch.

Following recent works , we set λLPIPS=0.8\lambda_{\text{LPIPS}}=0.8. For the similarity loss human facial domain, we use a pre-trained ArcFace network with λsim=0.1\lambda_{sim}=0.1. For the remaining domains, we utilize a MoCo-based loss with λsim=0.5\lambda_{sim}=0.5, as done in Tov et al. . All experiments were conducted on a single NVIDIA Tesla P40 GPU.

Appendix H Licenses

We provide the licenses of all datasets and models used in our work in Tab. 9.