JoJoGAN: One Shot Face Stylization

Min Jin Chong, David Forsyth

Introduction

A style mapper applies some fixed style to its input images (so, for example, taking faces to cartoons). This paper describes a simple procedure to learn a style mapper from a single example of the style. Our procedure allows, for example, an unsophisticated user to provide a style example, and then apply that style to their choice of image. Because stylizing face images – make me look like JoJo – is so desirable to unsophisticated users, we describe our method in the context of face images; but the method applies to anything.

To be useful, a procedure for learning a style mapper should: be easy to use; produce compelling and high quality results; require only one style reference, but accept and benefit from more; allow users to control how much style to transfer; and allow more sophisticated users to control what aspects of the style get transferred. We demonstrate with qualitative and quantitative evidence that our method meets these goals.

Learning a style mapper is hard, because the natural method – use paired or unpaired image translation – isn’t really practical. Collecting a new dataset per style is clumsy, and for many styles – Lucien Freud portraits, say – there may not be all that many examples. One might use few-shot learning techniques to fine-tune a StyleGAN by adjusting the discriminator (as in ). But these methods do not have detailed supervision from pixel-level losses and so mostly fail to capture distinct style details and diversity.

In contrast, JoJoGAN (our procedure) takes a reference image (or images – but one image is enough) and makes a paired dataset using GAN inversion and StyleGAN’s style-mixing property. This paired dataset is used to fine-tune StyleGAN using a novel direct pixel-level loss. The mechanics are straightforward: we can obtain a mapper (and so a rich supply of stylized portraits) from a single reference image in under a minute. JoJoGAN can use extreme style references (say, animal faces) successfully. Natural procedures control what aspects of the style are used and how much of the style is applied. Qualitative examples show that the resulting images look much better than alternative methods produce. Quantitative evidence strongly supports our method. Training and demo code is available at https://github.com/mchong6/JoJoGAN.

Related Work

Style transfer methods likely start with ; these are one shot methods, but do not result in style mappers in any natural way. Neural style transfer (NST) methods start with ; Johnson et al. offer a learned mapper, trained with a large dataset and Gatys et al.’s procedure to stylize . In contrast, our method uses much less data and produces much higher resolution images. A rich literature has followed, but general style transfer methods (for example ) cannot benefit from the detailed semantic and structural information captured by a GAN. Style transfer evaluation is mostly qualitative, but see . Deformable Style Transfer (DST) corrects structural errors by estimating spatial warps, then performing traditional neural style transfer; DST achieves impressive one-shot stylization, but warp estimation errors have significant effects and are hard to avoid (Figure 10).

StyleGAN remains the state-of-the art unconditional generative model due to its unique style-based architecture. StyleGAN’s AdaIN modulation layers (originally from ) have been shown to be disentangled and exhibit impressive editability . StyleGAN has also been used as a prior for numerous tasks such as superresolution and face restoration . Pinkney et al. first showed that finetuning the StyleGAN on a new dataset and performing layer swapping allows the StyleGAN to learn image to image translation with a relatively small dataset. But even obtaining a small paired dataset is hard: collection is difficult and expensive; one needs a new dataset for each new style; and in some cases (for example, Lucien Freud portrait style) there won’t be many style images in the first place. In contrast, JoJoGAN creates a paired dataset from a single style reference by manipulating a pretrained StyleGAN2 and a GAN inversion procedure, then finetunes using the created dataset.

One shot learning covers many applications (detection; classification; image synthesis), and methods remain specialized to their application. This paper focuses on one-shot image stylization, with a particular emphasis on faces.

One shot face stylization is now established. Learning a style mapper from very few examples results in overfitting problems. To control overfitting, introduce regularization terms while enforces constraints in the network’s weights. These methods need tens to hundreds of style example images; in contrast, JoJoGAN works with one. Furthermore, these methods have difficulty capturing small style details, likely because they rely on an adversarial loss. BlendGAN introduced a VGG-based style encoder and a weight blending module to learn arbitrary face stylization over a large styled faces dataset. As our comparisons show, this method fails to capture small but pertinent style details in face images. StyleGAN-NADA uses CLIP to perform zero/one shot image stylization based on text/image prompts, resulting in very strong generalization; as our comparisons show, StyleGAN-NADA fails to capture minute facial details that are important for face stylization.

Most similar to JoJoGAN is work by Zhu et al. (detailed experimental comparison in Figure 11 and Appendix); this also uses uses GAN inversion to find a corresponding real face from a reference, so creating a paired datapoint. Zhu et al. use this simple datapoint and a number of CLIP-based losses (from ). In contrast, JoJoGAN creates a large dataset of paired datapoints from a single one, and so needs only a simple pixel loss (with an optional identity loss). Zhu et al. use gradient descent inversion II2S (from ), which is slow but more accurate. In contrast, JoJoGAN uses feed forward inversion based on a simple encoder. Complex losses and slow inversion procedures mean Zhu et al. require some 1515 minutes to train on a Titan XP; in contrast, JoJoGAN require 11.

Methodology

Write TT for GAN inversion, GG for StyleGAN, ss for style parameters in StyleGAN’s S\mathcal{S}-space (notation after ; mixing in S\mathcal{S}-space works better, see Appendix 0.A.3), and θ\theta for the parameters of the vanilla StyleGAN. JoJoGAN uses four steps (Figure 2):

GAN inversion: We GAN invert the reference style image yy to obtain a style code w=T(y)w=T(y) and from that a set of ss parameters s(w)s(w).

Training set: We use ss to find a set of style codes S\mathcal{S} that are “close” to ss. Pairs (si,y)(s_{i},y) for si∈Ss_{i}\in\mathcal{S} will be our paired training set.

Finetuning: We finetune the StyleGAN to obtain θ^\hat{\theta} such that G(si;θ^)≈yG(s_{i};\hat{\theta})\approx y.

Inference: For input uu, our stylized face is G(s(T(u));θ^)G(s(T(u));\hat{\theta}) (so G∘s∘TG\circ s\circ T is our style mapper).

Remarkably, for any but extreme face style references yy, we have G(s(T(y));θ)G(s(T(y));\theta) is a realistic – rather than stylized – face image (eg Figure 2, step 1). This is likely because a GAN inverter is trained to produce codes that result in realistic faces, does not see stylized faces in training, and so fails to generalize properly – in this context, a useful property.

Step 2: Training set:

(and do so per batch). Different MM result in different stylization effects (Section 4).

Step 3: Finetuning StyleGAN:

We now assume that a properly trained style mapper will map si∈Ss_{i}\in\mathcal{S} to yy. This assumption certainly works, and is reasonable when the style mapper “reduces information” – so, for example, mapping faces with slightly different eye sizes or hair textures to the same reference image. We finetune StyleGAN to obtain

where L\mathcal{L} is a novel perceptual loss (this choice is important; Section 3.1).

Step 4: Inference:

For input uu, our stylized face is G(s(T(u));θ^)G(s(T(u));\hat{\theta}) (so G∘s∘TG\circ s\circ T is our style mapper). We could also generate random stylized samples by sampling random noise and generating with our finetuned StyleGAN.

1 Perceptual loss

The choice of loss in Equation 2 is important (Figure 4). While LPIPS is a natural choice, it produces methods that lose detail. LPIPS is built on a VGG backbone trained at a 224×224224\times 224 resolution, but StyleGAN produces 1024×10241024\times 1024 images. The standard way to handle this mismatch is to downsample the images to 256×256256\times 256 before computing LPIPS . But this downsampling means we cannot control fine-grained details, which are mostly lost. Similarly, computing LPIPS at the native 10241024 resolution leads to a complete loss of fine-grained detail as the VGG filters are not adapted to this resolution.

The pretrained StyleGAN discriminator is trained at the same resolution as the generator. The training process means that discriminator computes features that do not ignore details (otherwise the generator could produce low detail images). Discriminator features are known to stabilize GAN training when averaged over batches . We choose to use the difference in discriminator activations at particular layers, per image (details in Appendix 0.A.5). Write D(⋅)D(\cdot) for the activations; then L(G(si;θ),y)=∣ ⁣∣ ⁣D(G(si;θ))−D(y)∣ ⁣∣1\mathcal{L}(G(s_{i};\theta),y)=\mid\!\mid\!D(G(s_{i};\theta))-D(y)\mid\!\mid_{1}. A version of this loss is used in GPEN but to our knowledge, we are the first to compare it with others and show how effective it is.

Variants

Controlling Identity: Some style references distort the original identity of the inputs (Figure 4). In such cases, writing simsim for cosine similarity and FF for a pretrained face embedding network (we use ArcFace ), we use

to compel the finetuned network to preserve identity; we use it only for references that severely distort the identity and note in captions when we use it (eg. Figure 4(b)).

Controlling Style Intensity by Feature Interpolation: Feature interpolation allows us to vary the intensity of the style. Let fiAf_{i}^{A} be the layer ii intermediate feature maps from the original StyleGAN and fiBf_{i}^{B} from JoJoGAN; then we can perform continuous face stylization by using f=(1−α)fiA+αfiBf=(1-\alpha)f_{i}^{A}+\alpha f_{i}^{B} where α\alpha is the interpolation factor. Increasing α\alpha results in stronger style intensity (Figure 5).

Extreme Style References: For JoJoGAN to work, S\mathcal{S} has to consist of sis_{i} that produce sensible responses from the StyleGAN. If the style reference is (roughly) a human face, there are no problems. An extreme style reference image is one where GAN inversion produces ss that is out of distribution for the StyleGAN, for example, an image of an animal face. We are not aware of any test (other than trying) to distinguish between extreme and standard style references, but Figure 19 in the Appendix demonstrates that using ss from GAN inversion on animal faces results in poor style transfer. For extreme style references yy, rather than use s(T(y))s(T(y)) to construct S\mathcal{S}, we use the mean style code s‾=∑110000s(FC(z∼N(0,I)))\overline{s}=\sum_{1}^{10000}s(FC(z\sim\mathcal{N}(0,I))) (note this style code is the best possible estimate of s(T(y))s(T(y)) for an image yy that one does not have). With this modification, JoJoGAN works well on extreme style references (Figure 6; note how the animal head poses are controlled by the input images).

Multi-shot Stylization: JoJoGAN extends to multi-shot stylization in the natural way (use each reference to construct a Sk\mathcal{S}_{k} for each reference yky_{k}; now finetune using

Using more than one reference produces small but useful qualitative improvements in the style mapper (Figure 12)

Controlling Aspects of Style

Style transfer is intrinsically ambiguous. The output should be “like” the reference as to style, and “like” the input as to content, but the distinction between content and style is vague. JoJoGAN offers methods to choose whether (say) the output should have exaggerated eyes (like the reference) or more natural eyes (like the input). Simple control is obtained by choice of mask and by loss. More detailed control follows by careful attention to the GAN inversion.

Controlling Aspects of Style by Mask Choice and by Loss: Different choices of MM will produce significant differences in S\mathcal{S}, and so in results. Replacing too many elements of ss with random numbers may result in a JoJoGAN that maps every face to the style reference; replacing too few means finetuning sees too few examples. Furthermore, replacing elements at locations corresponding to different StyleGAN layers controls different effects (see ). Figure 8 demonstrates this choice has significant effects by displaying results from two different MM. The first gives dataset X\mathcal{X}, the second C\mathcal{C}. Both masks are chosen to maintain the input face pose and hairstyles while allowing features such as eye sizes and textures to vary, so the mask has ones in locations known to correspond to pose and zeros in those known to correspond to eye-sizes, see . But C\mathcal{C} is chosen so that the color of the input is preserved (so ones in relevant locations); and X\mathcal{X} so that color is driven by the style example. To ensure that the color of the input is preserved for the C\mathcal{C} case, we apply the loss in Equation (2) to grayscale versions of the relevant images. This means the StyleGAN is finetuned to obtain the spatial appearance of the style target, but not its colors (variants in Appendix 0.A.4)

The choice of GAN inverter matters. If the GAN inverter produces an extremely realistic face from the reference, JoJoGAN will be trained to map sis_{i} that represent highly realistic faces to the style reference, and so will tend to produce aggressively stylized faces. By the same argument, if the GAN inverter produces a somewhat stylized face from the reference, JoJoGAN will tend to produce lightly stylized faces and to preserve the features of the input face (so an input with small eyes will result in an output with small eyes, say – example in Appendix Figure 14). This effect can be used to control how much and what style is transferred by blending inverted codes.

Using two GAN inverters is clumsy in practice, but recall the mean style code is the best possible estimate of s(T(y))s(T(y)) for an image yy that one does not have), and so is the output of a (rather bad, but very fast) GAN inverter. We produce a virtual inverter V(y)V(y) by blending the code produced by our standard inverter with the mean, using the procedure of Section 3 (but a different mask MM). The blend is adjusted so that G(s(V(y));θ)G(s(V(y));\theta) has desirable properties (so, for example, to preserve the eyes of the reference, G(s(V(y));θ)G(s(V(y));\theta) should have realistic eyes). We then apply the JoJoGAN pipeline using VV rather than TT to generate training data. Using VV rather than TT in training changes the pairs (si,y)(s_{i},y) used in finetuning, and so the behavior of GG. At inference, we compute G(s(T(u);θ^)G(s(T(u);\hat{\theta}) as before. Figure 9 demonstrates the extent of our style control. In Figure 9(b), using the blended inversion gives us larger eyes and thicker lips compared to using the accurate inversion (a). Further detail on blending the inverter in Appendix 0.A.1.

Experiments

Setup: For GAN inversion, we use ReStyle . We finetune JoJoGAN for 200200 to 500500 iterations depending on the reference with Adam optimizer at a learning rate of 2×10−32\times 10^{-3}. Finetuning on an Nvidia A40 takes about 3030 to 6060 seconds.

Qualitative evaluation: A style mapper should: produce good looking outputs; faithfully transfer features from the style reference; and preserve the identity of the input. Qualitative evaluation shows JoJoGAN has these properties and vastly outperforms current methods.

Comparisons: Figure 10 shows comparisons of JoJoGAN to the state-of-the-art one/few shot stylization methods StyleGAN-NADA , BlendGAN , Ojha et al. and DST . JoJoGAN captures small details well that define the style while maintaining the identity of the input face well. JoJoGAN results are typically improved when there are multiple consistent style references. Figure 12 compares several one-shot stylizations of each of a set of examples with a multi-shot stylization using all. Notice that one-shot stylization copies effects from the style reference aggressively (as it must), whereas when there are multiple style examples, JoJoGAN is able to blend details to hew more closely to the input.

Figure 11 shows a comparison with (two examples in figure; others – except 22, for which we cannot find source – in supplementary). Note we can use only references shown in their paper, as the method is not open sourced.

Quantitative Evaluation by User study: We proceed in two stages, to reduce choice fatigue for users. From Figure 10, DST gives good results in most cases while other methods produces examples with severe problems. We therefore compare JoJoGAN to non-DST methods in a first study, and to DST in a second. In each, users see a style reference, an input face, and stylizations from the methods and are asked to choose the stylization that best captures the style reference and while preserving the original identity. The first study resulted in a total of 186186 responses from 3131 participants who overwhelmingly prefer JoJoGAN to other methods at 80.6%80.6\%; the effect is so large that no significance issues arise. The second study gathered 9696 responses from 1616 participants who prefer JoJoGAN to DST at 74%74\%.

Quantitative Evaluation by FID: FID is a metric that is widely used to evaluate the quality and diversity of generated images by comparing population statistics. FID can be used to evaluate style mappers as follows . Randomly select a reference from the style dataset and performing one shot stylization with it; now stylize a set of face images and compute the FID between the result and the original style dataset. To compute FID, we perform one shot stylization using the sketches dataset and compute FID using the test set. JoJoGAN scores well behind SOTA on this metric. We report FID for JoJoGAN for candor and show FID for SOTA comparisons in Figure 12, but point out that FID is a poor metric for style mappers. The procedure described cannot measure the fidelity with which the mapper preserves the input (for example, the FID for the completely ineffectual mapper that just produces a random sample from the style dataset would be close to zero). Further, a perfect style mapper might produce a high FID with the protocol described, because its stylized images should be biased toward the input (for example, a perfect mapper with only male input images should produce a population of sketches that is not close to the original set of sketches). Finally, the datasets used for stylization are often very small (290 in the case of the sketches dataset), and computing FID for a small dataset is dangerous due to large biases .

Failures: Using too small a S\mathcal{S} leads to problems (Appendix Figure 17), typically artifacts and missing style details. As JoJoGAN only sees a single style reference, it does not always work for all style references. One common issue JoJoGAN has is that the eye gaze direction is often driven by the reference image rather than the input. The intended behavior is to preserve the gaze direction of the original input, yet JoJoGAN copies the reference instead. Figure 13 shows results on very difficult references, illustrating visual failure modes.

References

Appendix 0.A Appendix

JoJoGAN relies on GAN inversion to create a paired dataset. We investigate the effect of using 33 different GAN inversion methods, e4e , II2S , and ReStyle in Figure 14.

Using e4e fails to accurately recreate the style reference and conveniently gives us a corresponding real face. On the other hand, ReStyle more accurately inverts the reference, giving a non-realistic face. II2S is a gradient-descent based method with a regularization term that allows us to map the style code to higher density region in the latent space. The regularization term results in a very realistic face that are somewhat inaccurate to the reference.

The different inversions give us different JoJoGAN results. Training with ReStyle leads to clean stylization that accurately preserves the features and proportions of the input face. Training with II2S on the other hand, leads to heavy stylization that borrows the shapes and proportions from the reference. However, this also leads to pretty heavy semantic changes from the input face and artifacts (note the change of identity, artifacts along the neck).

In practice, we blend the styles codes from ReStyle and the mean face. For MM, we borrow the style code from mean face at layers 77, 99 and 1111. This borrows the facial features of the mean face to the inversion. However, it is impossible to only affect the proportions of the features by simply blending coarsely at a layer level. For example, naively blending the mean face can change the expression of the inversion, e.g. from neutral to smiling or introduce artifacts. We thus have to blend at a finer scale, which we are able to do so by isolating specific facial features in the style space using RIS . Figure 15 compares the results of using different MM for blending. Note that when the blended image is more face-like (M3M3), the exaggerated features of the reference is transferred. However, significant artifacts are introduced, see M3M3 row 22. By carefully selecting MM, we can transfer the exaggerated features while avoiding artifacts, see M2M2.

A.2 Identity loss

Before computing identity loss, we grayscale the input images to prevent the identity loss from affecting the colors. The weight of the identity loss is reference dependent, but we typically choose between 2×1032\times 10^{3} to 5×1035\times 10^{3}.

A.3 Choice of style mixing space

Style mixing in Equation (1) allows us to generate more paired datapoints. It is reasonable to map faces with slight difference in textures, colors, to the same reference. As such it is pertinent that while we style mix to generate different faces, we need certain features such as identity, face pose, etc to remain the same. We study how the choice of latent space to do style mixing affects the stylization. In Figure 16 we see that style mixing in S\mathcal{S} gives better color reproduction and overall stylization effect. This is because S\mathcal{S} is more disentangled and allows us to more aggressively style mix without changing the features we want intact.

A.4 Varying dataset

Using C\mathcal{C} and X\mathcal{X} gives different stylization effects. Finetuning with X\mathcal{X} accurately reproduces the color profile of the reference while C\mathcal{C} tries to preserve the input color profile. However this is insufficient to fully preserve the colors as we see in Figure 17. Grayscaling the images before computing the loss in Equation (2) in addition to finetuning with C\mathcal{C} gives us stylization effects without altering the color profile. We show that it is necessary to use both C\mathcal{C} and grayscaling to achieve this effect and using X\mathcal{X} and grayscaling is insufficient.

A.5 Feature matching loss

For discriminator feature matching loss, we compute the intermediate activations after resblock 2,4,5,62,4,5,6.