StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets

Axel Sauer, Katja Schwarz, Andreas Geiger

Introduction

Computer graphics has long been concerned with generating photorealistic images at high resolution that allow for direct control over semantic attributes. Until recently, the primary paradigm was to create carefully designed 3D models which are then rendered using realistic camera and illumination models. A parallel line of research approaches the problem from a data-centric perspective. In particular, probabilistic generative models (Goodfellow et al., 2014; van den Oord et al., 2017; Song et al., 2021) have shifted the paradigm from designing assets to designing training procedures and datasets. Style-based GANs (StyleGANs) are a specific instance of these models, and they exhibit many desirable properties. They achieve high image fidelity (Karras et al., 2019, 2020b), fine-grained semantic control (Härkönen et al., 2020; Wu et al., 2021; Ling et al., 2021), and recently alias-free generation enabling realistic animation (Karras et al., 2021). Moreover, they reach impressive photorealism on carefully curated datasets, especially of human faces. However, when trained on large and unstructured datasets like ImageNet (Deng et al., 2009), StyleGANs do not achieve satisfactory results yet. One other problem plaguing data-centric methods, in general, is that they become prohibitively more expensive when scaling to higher resolutions as bigger models are required.

Initially, StyleGAN (Karras et al., 2019) was proposed to explicitly disentangle factors of variations, allowing for better control and interpolation quality. However, its architecture is more restrictive than a standard generator network (Radford et al., 2016; Karras et al., 2018) which seems to come at a price when training on complex and diverse datasets such as ImageNet. Previous attempts at scaling StyleGAN and StyleGAN2 to ImageNet led to sub-par results (Gwern, 2020; Grigoryev et al., 2022), giving reason to believe it might be fundamentally limited for highly diverse datasets (Gwern, 2020).

BigGAN (Brock et al., 2019) is the state-of-the-art GAN model for image synthesis on ImageNet. The main factors for BigGANs success are larger batch and model sizes. However, BigGAN has not reached a similar standing as StyleGAN as its performance varies significantly between training runs (Karras et al., 2020a) and as it does not employ an intermediate latent space which is essential for GAN-based image editing (Abdal et al., 2021; Patashnik et al., 2021; Collins et al., 2020; Wu et al., 2021). Recently, BigGAN has been superseded in performance by diffusion models (Dhariwal and Nichol, 2021). Diffusion models achieve more diverse image synthesis than GANs but are significantly slower during inference and prior work on GAN-based editing is not directly applicable. Following these arguments, successfully training StyleGAN on ImageNet has several advantages over existing methods.

The previously failed attempts at scaling StyleGAN raise the question of whether architectural constraints fundamentally limit style-based generators or if the missing piece is the right training strategy. Recent work by (Sauer et al., 2021) introduced Projected GANs which project generated and real samples into a fixed, pretrained feature space. Rephrasing the GAN setup this way leads to significant improvements in training stability, training time, and data efficiency. Leveraging the benefits of Projected GAN training might enable scaling StyleGAN to ImageNet. However, as observed by (Sauer et al., 2021), the advantages of Projected GANs only partially extend to StyleGAN on the unimodal datasets they investigated. We study this issue and propose architectural changes to address it. We then design a progressive growing strategy tailored to the latest StyleGAN3. These changes in conjunction with Projected GAN already allow surpassing prior attempts of training StyleGAN on ImageNet. To further improve results, we analyze the pretrained feature network used for Projected GANs and find that the two standard neural architectures for computer vision, CNNs and ViTs (Dosovitskiy et al., 2021), significantly improve performance when used jointly. Lastly, we leverage classifier guidance, a technique originally introduced for diffusion models to inject additional class-information (Dhariwal and Nichol, 2021).

Our contributions culminate in a new state-of-the-art on large-scale image synthesis, pushing the performance beyond existing GAN and diffusion models. We showcase inversion and editing for ImageNet classes and find that Pivotal Tuning Inversion (PTI) (Roich et al., 2021), a powerful new inversion paradigm, combines well with our model and even embeds out-of-domain images smoothly into our learned latent space. Our efficient training strategy allows us to triple the parameters of the standard StyleGAN3 while reaching prior state-of-the-art performance of diffusion models (Dhariwal and Nichol, 2021) in a fraction of their training time. It further enables us to be the first to demonstrate image synthesis on ImageNet-scale at a resolution of 102421024^{2} pixels. We will open-source our code and models upon publication.

Background

We first introduce the main building blocks of our system: the StyleGAN3 generator (Karras et al., 2021) and Projected GAN’s (Sauer et al., 2021) feature projectors and multi-scale discriminators.

StyleGAN. This section describes style-based generators in general with a focus on the latest StyleGAN3 (Karras et al., 2021). A StyleGAN generator consists of a mapping network Gm\mathbf{G}_{m} and a synthesis network Gs\mathbf{G}_{s}. First, Gm\mathbf{G}_{m} maps a normally distributed latent code z\mathbf{z} to a style code w\mathbf{w}. This style code w\mathbf{w} is then used for modulating the convolution kernels of Gs\mathbf{G}_{s} to control the synthesis process. The synthesis network Gs\mathbf{G}_{s} of StyleGAN3 starts from a spatial map defined by Fourier features (Tancik et al., 2020; Xu et al., 2021). This input then passes through NN layers of convolutions, non-linearities, and upsampling to generate an image. Each non-linearity is wrapped by an upsampling and downsampling operation to prevent aliasing. The low-pass filters used for these operations are carefully designed to balance image quality, antialiasing, and training speed. Concretely, their cutoff and stopband frequencies grow geometrically with network depth, the transition band half-widths are as wide as possible within the limits of the layer sampling rate, and only the last two layers are critically sampled, i.e., the filter cutoff equals the bandlimit. The number of layers NN is 1414, independent of the final output resolution.

Style mixing and path length regularization are methods for regularizing style-based generators. In style mixing, an image is generated by feeding sampled style codes w\mathbf{w} into different layers of Gs\mathbf{G}_{s} independently. Path length regularization encourages that a step of fixed size in latent space results in a corresponding fixed change in pixel intensity of the generated image (Karras et al., 2020b). This inductive bias leads to a smoother generator mapping and has several advantages including fewer artifacts, more predictable training behavior, and better inversion.

Progressive growing was introduced by (Karras et al., 2018) for stable training at high resolutions but (Karras et al., 2020b) found that it can impair shift-equivariance. (Karras et al., 2021) observe that texture sticking artifacts are caused by a lack of equivariance and carefully design StyleGAN3 to prevent texture sticking. Hence, in this paper, as we build on StyleGAN3, we can revisit the idea of progressive growing to improve convergence speed and synthesis quality.

Projected GAN. The original adversarial game between a generator G\mathbf{G} and a discriminator D\mathbf{D} can be extended by a set of feature projectors {Pl}\{\mathbf{P}_{l}\} (Sauer et al., 2021). The projectors map real images x\mathbf{x} and images generated by G\mathbf{G} to the discriminator’s input space. The Projected GAN objective is formulated as

where {Dl}\{\mathbf{D}_{l}\} is a set of independent discriminators operating on different feature projections. The projectors consist of a pretrained feature network F\mathbf{F}, cross-channel mixing (CCM) and cross-scale mixing (CSM) layers. The purpose of CCM and CSM is to prohibit the discriminators from focusing on only a subset of its input feature space which would result in mode collapse. Both modules employ differentiable random projections that are not optimized during GAN training. CCM mixes features across channels via random 1x1 convolutions, CSM mixes features across scales via residual random 3x3 convolution blocks and bilinear upsampling. The output of CSM is a feature pyramid consisting of four feature maps at different resolutions. Four discriminators operate independently on these feature maps. Each discriminator uses a simple convolutional architecture and spectral normalization (Miyato et al., 2018). The depth of the discriminator varies depending on its input resolution, i.e., a spatially larger feature map corresponds to a deeper discriminator. Other than spectral normalization, Projected GANs do not use additional regularization such as gradient penalties (Mescheder et al., 2018). Lastly, (Sauer et al., 2021) apply differentiable data-augmentation (Zhao et al., 2020) before F\mathbf{F} which improves Projected GAN’s performance independent of the dataset size.

(Sauer et al., 2021) evaluate several combinations of F\mathbf{F} and G\mathbf{G} and find an EfficientNet-Lite0 (Tan and Le, 2019) and a FastGAN generator (Liu et al., 2021) to work especially well. When using a StyleGAN generator, they observe that the discriminators can quickly overpower the generator for suboptimal learning rates. The authors suspect that the generator might adapt too slowly due to its design which modulates feature maps with styles learned by a mapping network.

Scaling StyleGAN to ImageNet

As mentioned before, StyleGAN has several advantages over existing approaches that work well on ImageNet. But a naïve training strategy does not yield state-of-the-art performance (Gwern, 2020; Grigoryev et al., 2022). Our experiments confirm that even the latest StyleGAN3 does not scale well, see Fig. 1. Particularly at high resolutions, the training becomes unstable. Therefore, our goal is to train a StyleGAN3 generator on ImageNet successfully. Success is defined in terms of sample quality primarily measured by inception score (IS) (Salimans et al., 2016) and diversity measured by Fréchet Inception Distance (FID) (Heusel et al., 2017). Throughout this section, we gradually introduce changes to the StyleGAN3 baseline (Config-A) and track the improvements in Table 3. First, we modify the generator and its regularization losses, adapting the latent space to work well with Projected GAN (Config-B) and for the class-conditional setting (Config-C). We then revisit progressive growing to improve training speed and performance (Config-D). Next, we investigate the feature networks used for Projected GAN training to find a well-suited configuration (Config-E). Lastly, we propose classifier guidance for GANs to provide class information via a pretrained classifier (Config-F). Our contributions enable us to train a significantly larger model than previously possible while requiring less computation than prior art. Our model is three times larger in terms of depth and parameter count than a standard StyleGAN3. However, to match the prior state-of-the-art performance of ADM (Dhariwal and Nichol, 2021) at a resolution of 5122512^{2} pixels, training the models on a single NVIDIA Tesla V100 takes 400400 days compared to the previously required 19141914 V100-days. We refer to our model as StyleGAN-XL (Fig. 2).

Training on a diverse and class-conditional dataset makes it necessary to introduce several adjustments to the standard StyleGAN configuration. We construct our generator architecture using layers of StyleGAN3-T, the translational-equivariant configuration of StyleGAN3. In initial experiments, we found the rotational-equivariant StyleGAN3-R to generate overly symmetric images on more complex datasets, resulting in kaleidoscope-like patterns.

Regularization. In GAN training, it is common to use regularization for both, the generator and the discriminator. Regularization improves results on uni-modal datasets like FFHQ (Karras et al., 2019) or LSUN (Yu et al., 2015), whereas it can be detrimental on multi-modal datasets (Brock et al., 2019; Gwern, 2020). Therefore, we aim to avoid regularization when possible. (Karras et al., 2021) find style mixing to be unnecessary for the latest StyleGAN3; hence, we also disable it. Path length regularization can lead to poor results on complex datasets (Gwern, 2020) and is, per default, disabled for StyleGAN3 (Karras et al., 2021). However, path length regularization is attractive as it enables high-quality inversion (Karras et al., 2020b). We also observe unstable behavior and divergence when using path length regularization in practice. We found that this problem can be circumvented by only applying regularization after the model has been sufficiently trained, i.e., after 200k images. For the discriminator, following (Sauer et al., 2021), we use spectral normalization without gradient penalties. In addition, we blur all images with a Gaussian filter with σ=2\sigma=2 pixels for the first 200k200k images. Discriminator blurring has been introduced in (Karras et al., 2021) for StyleGAN3-R. It prevents the discriminator from focusing on high frequencies early on, which we found beneficial across all settings we investigated.

Pretrained Class Embeddings. Conditioning the model on class information is essential to control the sample class and improve overall performance. A class-conditional variant of StyleGAN was first proposed in (Karras et al., 2020a) for CIFAR10 (Krizhevsky et al., 2009) where a one-hot encoded label is embedded into a 512-dimensional vector and concatenated with z\mathbf{z}. For the discriminator, class information is projected onto the last discriminator layer (Miyato and Koyama, 2018). We observe that Config-B tends to generate similar samples per class resulting in high IS. To quantify mode coverage, we leverage the recall metric (Kynkäänniemi et al., 2019) and find that Config-B achieves a low recall of 0.0040.004. We hypothesize that the class embeddings collapse when training with Projected GAN. Therefore, to prevent this collapse, we aim to ease optimization of the embeddings via pretraining. We extract and spatially pool the lowest resolution features of an Efficientnet-lite0 (Tan and Le, 2019) and calculate the mean per ImageNet class. The network has a low channel count to keep the embedding dimension small, following the arguments of the previous section. The embedding passes through a linear projection to match the size of z\mathbf{z} to avoid an imbalance. Both Gm\mathbf{G}_{m} and Di\mathbf{D}_{i} are conditioned on the embedding. During GAN training, the embedding and the linear projection are optimized to allow specialization. Using this configuration, we observe that the model generates diverse samples per class, and recall increases to 0.150.15 (Config-C). Note that for all configurations in this ablation, we restrict the training time to 15  V-100  days15\;V\text{-}100\;days. Hence, the absolute recall is markedly lower compared to the fully trained models. Conditioning a GAN on pretrained features was also recently investigated by (Casanova et al., 2021). In contrast to our approach, (Casanova et al., 2021) condition on specific instances, instead of learning a general class embedding.

2. Reintroducing Progressive Growing

Progressively growing the output resolution of a GAN was introduced by (Karras et al., 2018) for fast and more stable training. The original formulation adds layers during training to both G\mathbf{G} and D\mathbf{D} and gradually fades in their contribution. However, in a later work, it was discarded (Karras et al., 2020b) as it can contribute to texture sticking artifacts. Recent work by (Karras et al., 2021) finds that the primary cause of these artifacts is aliasing, so they redesign each layer of StyleGAN to prevent it. This motivates us to reconsider progressive growing with a carefully crafted strategy that aims to suppress aliasing as best as possible. Training first on very low resolutions, as small as 16216^{2} pixels, enables us to break down the daunting task of training on high-resolution ImageNet into smaller subtasks. This idea is in line with the latest work on diffusion models (Nichol and Dhariwal, 2021; Saharia et al., 2021; Dhariwal and Nichol, 2021; Ho et al., 2022). They observe considerable improvements in FID on ImageNet when using a two-stage model, i.e., stacking an independent low-resolution model and an upsampling model to generate the final image.

Commonly, GANs follow a rigid sampling rate progression, i.e., at each resolution, there is a fixed amount of layers followed by an upsampling operation using fixed filter parameters. StyleGAN3 does not follow such a progression. Instead, the layer count is set to 1414, independent of the output resolution, and the filter parameters of up- and downsampling operations are carefully designed for antialiasing under the given configuration. The last two layers are critically sampled to generate high-frequency details. When adding layers for the subsequent highest resolution, discarding the previously critically sampled layers is crucial as they would introduce aliasing when used as intermediate layers (Karras et al., 2020b, 2021). Furthermore, we adjust the filter parameters of the added layers to adhere to the flexible layer specification of (Karras et al., 2021); we refer to the supplementary for details. In contrast to (Karras et al., 2018) we do not add layers to the discriminator. Instead, to fully utilize the pretrained feature network F\mathbf{F}, we upsample both data and synthesized images to F\mathbf{F}’s training resolution (2242224^{2} pixels) when training on smaller images.

We start progressive growing at a resolution of 16216^{2} using 1111 layers. Every time the resolution increases, we cut off 22 layers and add 77 new ones. Empirically, fewer layers result in worse performance; adding more leads to increased overhead and diminishing returns. For the final stage at 102421024^{2}, we add only 55 layers as the last two are not discarded. This amounts to 3939 layers at the maximum resolution of 102421024^{2}. Instead of a fixed growing schedule, each stage is trained until FID stops decreasing. We find it beneficial to use a large batch size of 20482048 on lower resolution (16216^{2} and 32232^{2}), similar to (Brock et al., 2019). On higher resolutions, smaller batch sizes suffice (64264^{2} to 2562256^{2}: 256256, 5122512^{2} to 102421024^{2}: 128128). Once new layers are added, the lower resolution layers remain fixed to prevent mode collapse.

In our ablation study, FID improves only slightly (Config-D) compared to Config-C. However, the main advantage can be seen at high resolutions, where progressive growing drastically reduces training time. At resolution 5122512^{2}, we reach the prior state-of-the-art (FID  =3.85\;=3.85) after 22 V100-days. This reduction is in contrast to other methods such as ADM, where doubling the resolution from 2562256^{2} to 5122512^{2} pixels corresponds to increasing training time from 393393 to 19141914 V100-days to find the best performing modelNote that these settings are not directly comparable as the stem of our model is pretrained, but the values should give a general sense of the order of magnitude.. As our aim is not to introduce texture sticking artifacts, we measure EQ-TEQ\text{-}T, a metric for determining translation equivariance (Karras et al., 2021), where higher is better. Config-C yields EQ-T=55EQ\text{-}T=55, while Config-D attains EQ-T=48EQ\text{-}T=48. This only slight reduction in equivariance shows that Config-D restricts aliasing almost as well as a configuration without growing. For context, architectures with aliasing yield EQ-T∼15EQ\text{-}T\sim 15.

3. Exploiting Multiple Feature Networks

An ablation study conducted in (Sauer et al., 2021) finds that most pretrained feature networks F\mathbf{F} perform similarly in terms of FID when used for Projected GAN training regardless of training data, pretraining objective, or network architecture. However, the study does not answer if combining several F\mathbf{F} is advantageous. Starting from the standard configuration, an EfficientNet-lite0, we add a second F\mathbf{F} to inspect the influence of its pretraining objective (classification or self-supervision) and architecture (CNN or Vision Transformer (ViT) (Dosovitskiy et al., 2021)). The results in Table 3 show that an additional CNN leads to slightly lower FID. Combining networks with different pretraining objectives does not offer benefits over using two classifier networks. However, combining an EfficientNet with a ViT improves performance significantly. This result corroborates recent results in neural architecture literature, which find that supervised and self-supervised representations are similar (Grigg et al., 2021), whereas ViTs and CNNs learn different representations (Raghu et al., 2021). Combining both architectures appears to have complementary effects for Projected GANs. We do not see significant improvements when adding more networks; hence, Config-E uses the combination of EfficientNet (Tan and Le, 2019) and DeiT-base (Touvron et al., 2021).

4. Classifier Guidance for GANs

(Dhariwal and Nichol, 2021) introduced classifier guidance to inject class information into diffusion models. Classifier guidance modifies each diffusion step at time step tt by adding gradients of a pretrained classifier ∇xtlog⁡pϕ(c∣xt,t)\nabla_{\mathbf{x}_{t}}\log p_{\phi}(\mathbf{c}|x_{t},t). The best results are obtained by applying guidance on class-conditional models and scaling the classifier gradients by a constant λ>1\lambda>1. This combination indicates that our model may also profit from classifier guidance, even though it already receives class information via embeddings.

We first pass the generated image x\mathbf{x} through a pretrained classifier CLF to predict the class label cic_{i}. We then add a cross-entropy loss LCE=−∑i=0Ccilog⁡CLF(xi)\mathcal{L}_{CE}=-\sum_{i=0}^{C}c_{i}\log CLF({x}_{i}) as an additional term to the generator loss and scale this term by a constant λ\lambda. For the classifier, we use DeiT-small (Touvron et al., 2021), which exhibits strong classification performance while not adding much overhead to the training. Similar to (Dhariwal and Nichol, 2021), we observe a significant improvement in IS, indicating an increase in sample quality (Config-F). We find λ=8\lambda=8 to work well empirically. Classifier guidance only works well on higher resolutions (>322>32^{2}); otherwise, it leads to mode collapse. This is in contrast to (Dhariwal and Nichol, 2021) who exclusively guide their low-resolution model. The difference stems from how guidance is applied: we use it for model training, whereas (Dhariwal and Nichol, 2021) guide the sampling process.

Results

In this section, we first compare StyleGAN-XL to the state-of-the-art approaches for image synthesis on ImageNet. We then evaluate the inversion and editing capabilities of StyleGAN-XL. As described above, we scale our model to a resolution of 102421024^{2} pixels, which no prior work has attempted so far on ImageNet. The resolution of most images in ImageNet is lower. We therefore preprocess the data with a super-resolution network (Liang et al., 2021), see supplementary.

Both our work and (Dhariwal and Nichol, 2021) use classifier networks to guide the generator. To ensure the models are not inadvertently optimizing for FID and IS, which also utilize a classifier network, we propose random-FID (rFID). For rFID, we calculate the Fréchet distance in the pool_3 layer of a randomly initialized inception network (Szegedy et al., 2015). The efficacy of random features for evaluating generative models has been demonstrated in (Naeem et al., 2020). Furthermore, we report sFID (Nash et al., 2021) to assess spatial structure. Lastly, sample fidelity and diversity are evaluated via precision and recall (Kynkäänniemi et al., 2019).

In Table 4.1, we compare StyleGAN-XL to the currently strongest GAN model (BigGAN-deep (Brock et al., 2019)) and diffusion models (CDM (Ho et al., 2022), ADM (Dhariwal and Nichol, 2021)) on ImageNet. The values for ADM are calculated with and without additional methods (Upsampling U and Classifier Guidance G). For StyleGAN2, we report numbers by (Grigoryev et al., 2022). We find that StyleGAN-XL substantially outperforms all baselines across all resolutions in FID, sFID, rFID, and IS. An exception is recall, according to which StyleGAN-XL’s sample diversity lies between BigGAN and ADM, making progress in closing the gap between these model types. BigGAN’s sample quality is the best among all compared approaches, which comes at the price of significantly lower recall. StyleGAN-XL allows for the truncation trick to increase sample fidelity, i.e., we can interpolate a sampled style code ww with the class-wise mean style code wˉ\bar{w}. We observe that for StyleGAN-XL, truncation does not increase precision, indicating that developing novel truncation methods for high-diversity GANs is an exciting research direction for future work. Interestingly, StyleGAN-XL attains high diversity across all resolutions, which can be attributed to our progressive growing strategy. Furthermore, this strategy enables to scale to megapixel resolution successfully. Training at 102421024^{2} for a single V100-day yields a noteworthy FID of 2.82.8. At this resolution, we do not compare to baselines because of resource constraints as they are prohibitively expensive to train. visualizes generated samples at increasing resolutions. Fig. 3 visualizes generated samples at increasing resolutions. In the supplementary, we show additional interpolations and qualitative comparisons to BigGAN and ADM.

2. Inversion and Manipulation

GAN-editing methods first invert a given image into latent space, i.e., find a style code ww that reconstructs the image as faithful as possible when passed through Gs\mathbf{G}_{s}. Then, ww can be manipulated to achieve semantically meaningful edits (Goetschalckx et al., 2019; Shen et al., 2020).

Inversion. Standard approaches for inverting Gs\mathbf{G}_{s} use either latent optimization (Abdal et al., 2019; Creswell and Bharath, 2019; Karras et al., 2020b) or an encoder (Perarnau et al., 2016; Alaluf et al., 2021; Tov et al., 2021). A common way to achieve low reconstruction error is to use an extended definition of the latent space: W+\mathcal{W}+. For W+\mathcal{W}+ a separate w\mathbf{w} is chosen for each layer of Gs\mathbf{G}_{s}. However, as highlighted by (Zhu et al., 2020; Tov et al., 2021), this extended definition achieves higher reconstruction quality in exchange for lower editability. Therefore, (Tov et al., 2021) carefully design an encoder to maintain editability by mapping to regions of W+\mathcal{W}+ that are close to the original distribution of W\mathcal{W}. We follow (Karras et al., 2020b) and use the original latent space W\mathcal{W}. We find that StyleGAN-XL already achieves satisfactory inversion results using basic latent optimization. For inversion on the ImageNet validation set at 5122512^{2}, StyleGAN-XL yields PSNR=13.5\text{PSNR}=13.5 on average, improving over BigGAN at PSNR=10.8\text{PSNR}=10.8. Besides better pixel-wise reconstruction, StyleGAN-XL’s inversions are semantically closer to the target images. We measure the FID between reconstructions and targets, and StyleGAN-XL attains FID=21.7\text{FID}=21.7 while BigGAN reaches FID=47.5\text{FID}=47.5. For qualitative results, implementation details and additional metrics, we refer to the supplementary.

Given the results above, it is also possible to further refine the obtained reconstructions. (Roich et al., 2021) recently introduced pivotal tuning inversion (PTI). PTI uses an initial inverted style code as a pivot point around which the generator is finetuned. Additional regularization prevents altering the generator output far from the pivot. Combining PTI with StyleGAN-XL allows us to invert both in-domain (ImageNet validation set) and out-of-domain images almost precisely. At the same time, the generator output remains perceptually smooth, see Fig. 4.

Image Manipulation. Given the inverted images, we can leverage GAN-based editing methods (Voynov and Babenko, 2020; Härkönen et al., 2020; Shen and Zhou, 2021; Kocasari et al., 2022; Spingarn et al., 2021) to manipulate the style code w\mathbf{w}. In Fig. 5 (Left), we first invert a given source image via latent space optimization. We can then apply a manipulation directions obtained by, e.g., GANspace (Härkönen et al., 2020). Prior work (Jahanian et al., 2020) also investigates in-plane translation. This operation can be directly defined in the input grid of StyleGAN-XL. The input grid also allows performing extrapolation, see Fig. 5 (Left).

An inherent property of StyleGAN is the ability of style mixing by supplying the style codes of two samples to different layers of Gs\mathbf{G}_{s}, generating a hybrid image. This hybrid takes on different semantic properties of both inputs. Style mixing is commonly employed for instances of a single domain, i.e., combining two human portraits. StyleGAN-XL inherits this ability and, to a certain extent, even generates out-of-domain combinations between different classes, akin to counterfactual images (Sauer and Geiger, 2021). This technique works best for aligned samples, similar to StyleGAN’s originally favored setting, FFHQ. Curated examples are shown in Fig. 5 (Right).

Limitations and Future Work

Our contributions allow StyleGAN to accomplish state-of-the-art high-resolution image synthesis on ImageNet. Furthermore, applying it to big and small unimodal datasets is straightforward, and we also achieve state-of-the-art performance on FFHQ and Pokemon at resolution 102421024^{2}, see supplementary. Exploring new editing methods and dataset generation (Chai et al., 2021; Li et al., 2022) using StyleGAN-XL are exciting future avenues. Furthermore, future work may tackle an even larger megapixel dataset. However, a larger yet diverse dataset is not available so far. Current large-scale, high-resolution datasets are of single object classes or contain many similar images (Zhang et al., 2020; Fregin et al., 2018; Perot et al., 2020). In the following, we discuss limitations of the current model, which should be addressed in the future.

Architectural Limitations. First, StyleGAN-XL is three times larger than StyleGAN3, constituting a higher computational overhead when used as a starting point for finetuning. Therefore, it will be worth exploring GAN distillation methods (Chang and Lu, 2020) that trade-off performance for model size. Second, we find StyleGAN3, and consequently, StyleGAN-XL, harder to edit, e.g., high-quality edits via W\mathcal{W} are noticeably easier to achieve with StyleGAN2. As already observed in (Karras et al., 2021), StyleGAN3’s semantic controllability is reduced for the sake of equivariance. However, techniques using the StyleSpace (Wu et al., 2021), e.g., StyleMC (Kocasari et al., 2022), tend to yield better results in our experiments, confirming the findings of concurrent work by (Alaluf et al., 2022). Furthermore, we remark that our framework can also easily be used with StyleGAN2 layers.

References

Appendix A Preprocessing ImageNet

An initial challenge is the lack of high-resolution data; the mean resolution of ImageNet is 469×387469\times 387. Similar to the procedure used for generating CelebA-HQ(Karras et al., 2018), we preprocess the whole dataset with SwinIR-Large (Liang et al., 2021), a recent model for real-world image super-resolution. Of course, a trivial way of achieving good performance on this dataset would be to draw samples from a 2562256^{2} generative model and passing it through SwinIR. However, SwinIR adds significant computational overhead as it is 6060 times slower than our upsampling stack. Furthermore, this way, StyleGAN-XL’s weights can be used for initialization when finetuning on other high-resolution datasets. Lastly, combining StyleGAN-XL and SwinIR would impair translation equivariance.

Appendix B Classes of unaligned Humans

We observe that ADM (Dhariwal and Nichol, 2021) generates more convincing human faces than StyleGAN-XL and BigGAN. Both GANs can synthesize realistic faces; however, the main challenge in this setting is that the dataset is unstructured, and the humans are not aligned. (Brock et al., 2019) remarked the particular challenge of classes containing details to which human observers are more sensitive. We show examples in Fig. 7.

Appendix C Inference Speed

GANs generate samples in a single forward pass, unlike diffusion models that must be applied several hundred or thousand times to generate a sample. Table C compares StyleGAN-XL to ADM. We find that StyleGAN-XL is several orders of magnitude faster. In defense of diffusion models, speeding up their sampling is an active area of research, and novel techniques (Watson et al., 2021) may be able to reduce this gap in the future.

Appendix D Results on Unimodal Datasets

StyleGAN-XL is designed to enable training on large and diverse datasets. However, applying it to big and small unimodal datasets is straightforward. In contrast to the configuration for ImageNet, we begin with ten layers at the lowest stage and add two layers per resolution stage. Furthermore, we do not employ classifier guidance. Table D reports the results for both datasets at resolution 102421024^{2}, StyleGAN-XL achieves state-of-the-art performance on both.

Appendix E Additional Qualitative Results

In the following, we present additional qualitative results. Fig. 8 shows additional interpolations between samples from different classes. Fig. 10 and Fig. 11 show samples on FFHQ 102421024^{2} and Pokemon 102421024^{2} respectively. Lastly, we compare BigGAN, ADM, and StyleGAN-XL on different ImageNet classes. For a fair comparison, we do not use truncation or classifier guidance. Instead, we show images with the largest logits given by a VGG16 which corresponds to individual image quality.

Appendix F Implementation details

Inversion. Following (Karras et al., 2020b), we use basic latent optimization in W\mathcal{W} for inversion. Given a target image, we first compute its average style code wˉ\bar{\mathbf{w}} by running 1000010000 random latent codes z\mathbf{z} and target specific class samples c\mathbf{c} through the mapping network. As the class label of the target image is unknown, we pass it to a pretrained classifier. We then use the classifier logits as a multinomial distribution to sample c\mathbf{c}. In our experiments, we use Deit-base (Touvron et al., 2021) as a classifier, but other choices are possible. At the beginning of optimization , we initialize w=wˉ\mathbf{w}=\bar{\mathbf{w}}. The components of w\mathbf{w} are the only trainable parameters. The optimization runs for 1000 iterations using the Adam optimizer (Kingma and Ba, 2015) with default parameters. We optimize the LPIPS (Zhang et al., 2018) distance between the target image and the generated image. For StyleGAN-XL, the maximum learning rate is λmax=0.05\lambda_{max}=0.05. It is ramped up from zero linearly during the first 50 iterations and ramped down to zero using a cosine schedule during the last 250 iterations. For BigGAN, we empirically found λmax=0.001\lambda_{max}=0.001 and a ramp-down over the last 750 iterations to yield the best results. All inversion experiments are performed at resolution 5122512^{2} and computed on 5k5k images (1010% of the validation set). We report the results in Table F and show qualitative results in Fig. 9.

Training StyleGAN3 on ImageNet. For training StyleGAN3, we use the official PyTorch implementationhttps://github.com/NVlabs/stylegan3.git. The results in Fig. 1 are computed with the StyleGAN3-R configuration on resolution 2562256^{2} until the discriminator has seen 1010 million images. We find that StyleGAN3-R and StyleGAN3-T converge to similar FID without any changes to their training paradigm. The run with the best FID score was selected from three runs with different random seeds. We use a channel base of 1638416384 and train on 88 GPUs with total batch size 256256, γ=0.256\gamma=0.256. The remaining settings are chosen according to the default configuration of the code release. For the ablation study in Table 3 , we use the StyleGAN3-T configuration as baseline since StyleGAN-XL builds upon the translational-equivariant layers of StyleGAN3. We train on 44 GPUs with total batch size 256256 and batch size 3232 per GPU, γ=0.25\gamma=0.25, and disable augmentation.

Training & Evaluation. For all our training runs, we do not use data amplification via x-flips following (Karras et al., 2020b). Furthermore, we evaluate all metrics using the official StyleGAN3 codebase. For the baseline values in Table 4.1 we report the numbers of (Dhariwal and Nichol, 2021). The official codebase of ADMhttps://github.com/openai/guided-diffusion provides files containing 5050k samples for ADM and BigGAN. We utilize the provided samples to compute rFID. Following (Dhariwal and Nichol, 2021), we compute precision and recall between 1010k real samples and 5050k generated samples. Table F reports the results on ImageNet at lower resolutions.

Layer configurations. We start progressive growing at resolution 16216^{2} using 1111 layers. The layer specifications are computed according to (Karras et al., 2021) and remain fixed for the remaining training. For the next stage, at resolution 32232^{2}, we discard the last 22 layers and add 77 new ones. The specifications for the new layers are computed according to (Karras et al., 2021) for a model with resolution 32232^{2} and 1616 layers. Continuing this strategy up to resolution 102421024^{2} yields the flexible layer specification of StyleGAN-XL in Fig. 15.