Freeze the Discriminator: a Simple Baseline for Fine-Tuning GANs

Sangwoo Mo, Minsu Cho, Jinwoo Shin

Introduction

Generative adversarial networks (GANs) have shown a remarkable success across a broad range of applications in computer vision, graphics, and machine learning, e.g., image generation , image-to-image translation , and video-to-video synthesis . Current state-of-the-art GANs, however, often require a large amount of training data and heavy computational resources, which thus limits the applicability of GANs in practical scenarios. Numerous techniques have been proposed to overcome this limitation, e.g., transferring knowledge of a well-trained source model , learning meta-knowledge for quick adaptation to a target domain , using an auxiliary task to facilitate training , improving an inference procedure of suboptimal models , using an expressive prior distribution , actively choosing samples to give supervision for conditional generation , or actively sampling mini-batches for training .

Among the approaches, transfer learning is arguably the most promising way to training models under limited data and resources. Indeed, most of the recent success in deep learning is built upon strong backbones pre-trained on large datasets in supervised or self-supervised ways. Following the success of transferring classifiers in recognition tasks, one can also consider utilizing well-trained GAN backbones for downstream generation tasks. While several methods propose such transfer-learning approaches to training GANs , they are often prone to overfitting with limited training data or not robust in learning a significant distribution shift .

In this paper, we propose a simple yet effective baseline for transfer learning of GANs. In particular, we show that simple fine-tuning of GANs (both generator and discriminator) with frozen lower layers of the discriminator performs surprisingly well (see Figure 1). Intuitively, the lower layers of the discriminator learn generic features of images while the upper layers learn to classify whether the image is real or fake based on the extracted features. We remark that this dichotomous view of a feature extractor and a classifier (and freezing the feature extractor for fine-tuning) is not new; it has been widely used for training classifiers . We confirm that this view is also useful for GANs, and set its proper baseline for transfer learning of GANs.

We demonstrate the effectiveness of the simple baseline, dubbed FreezeD, using various architectures and datasets. For unconditional GANs, we fine-tune the StyleGAN architecture, which is pre-trained on FFHQ , onto Animal Face and Anime Face datasets, and for conditional GANs, fine-tune the SNGAN-projection architecture, which is pre-trained on ImageNet , onto Oxford Flower , CUB-200-2011 , and Caltech-256 datasets. FreezeD outperforms previous techniques for all experiment settings, e.g., improving the FID score from 64.28 of fine-tuning to 61.46 (-4.4%) on ‘Dog’ class of Animal Face dataset.

Methods

The goal of GANs is to learn a generator (and a corresponding discriminator) to match with a target data distribution. In transfer learning, we assume one can utilize a pre-trained source generator (and a corresponding discriminator) trained on the source data distribution to improve the target generator. See for the survey of GANs.

We first briefly review previous methods for transfer learning of GANs.

Fine-tuning : The most intuitive and effective way to transferring knowledge is fine-tuning; initialize the parameters of target models as the pre-trained weights of the source models. The authors report that fine-tuning both the generator and the discriminator indeed shows the best performance.It is more crucial for our case, as we use stronger source models. However, fine-tuning often suffer from overfitting; hence one needs a proper regularization.

Scale/shift : Since naïve fine-tuning is prone to overfitting, scale/shift suggest to update the normalization layers only (e.g., batch normalization (BN) ) while fixing all other weights. However, it often shows inferior results due to its restriction, especially when there is a significant shift between the source and the target distribution.

Generative latent optimization (GLO) : Since GAN loss is given by the discriminator, which can be unreliable for limited data, GLO suggests fine-tuning the generator with supervised learning, where the loss is given by the sum of the L1 loss and the perceptual loss . Here, GLO jointly optimizes the generator and the latent codes to avoid overfitting; one latent code (and its corresponding generated sample) matches one real sample; hence, the generator can generalize samples by interpolation. While GLO improves the stability, it tends to produce blurry images due to the lack of adversarial loss (and prior knowledge of the source discriminator).

MineGAN : To avoid overfitting of the generator, MineGAN suggests to fix the generator and modify the latent codes. To this end, MineGAN train a miner network that transforms the latent code to another latent code. While this importance-sampling-like approach can be effective when the source distribution and the target distributions share support, it may not be generalized when their supports are disjointed.

We now introduce a simple baseline, FreezeD, which outperforms the previous methods despite its simplicity, and suggest two other methods for possible future directions, which may give further improvement. We remark that our goal is not to advocate the state-of-the-art but to set a simple and effective baseline. By doing so, we hope to encourage new techniques that outperform the proposed baseline.

FreezeD (our proposed baseline): We find that simply freezing the lower layers of the discriminator and only fine-tune the upper layers performs surprisingly well. We call this simple yet effective baseline as FreezeD, and will demonstrate its consistent gain over the previous methods in the experimental section.

L2-SP : In addition to the prior methods, we test L2-SP, which is known to be effective for the classifiers. Built upon to the fine-tuning, L2-SP regularizes the target models not to move far from the source models. In particular, it regularizes the L2-norm of the parameters of source models and target models. In our experiments, we applied L2-SP to the generator, discriminator, and both, but the results were not satisfactory. However, since freezing layers can be viewed as giving the infinite weight of L2-SP for the chosen layers and 0 for the other layers, using proper weights for each layer may perform better.

Feature distillation : We also test feature distillation, one of the most popular approaches to transfer learning of classifiers. Among the variants, we simply distill the activations of the source models and target models (initialized to the source models). We find that feature distillation shows comparable results to FreezeD while takes twice computation. Investigating more advanced techniques (e.g., ) would be an interesting and promising future direction.We observe that feature distillation shows more stable (but similar best) results than FreezeD for SNGAN-projection experiments.

Experiments

In this section, we demonstrate the effectiveness of the simple yet effective baseline, FreezeD. We conduct extensive experiments for both unconditional GANs and conditional GANs in Section 3.1 and Section 3.2, respectively.

We first demonstrate results for unconditional GANs. We use the StyleGAN architecture pre-trained on FFHQ dataset, and fine-tune it on Animal Face and Anime Face datasets. We use full 20 classes of the Animal Face dataset, and the first 10 classes among the total 1,000 classes of the Anime Face dataset. Each class contains around 100 samples. We use the public pre-trained modelhttps://github.com/rosinality/style-based-gan-pytorch of resolution 256×\times256 and fine-tune the models following the original training scheme for 50,000 iterations. We remark that the training performed successfully without progressive training by utilizing the source models.

Figure 2 visualizes the generated samples using the original weights and the fine-tuned weights on ‘Cat’ and ‘Dog’ classes in the Animal Face dataset. Notably, the same latent code shares the same semantics even after fine-tuning. See Appendix D for more qualitative results. We also evaluate the FID scores of the vanilla fine-tuning and FreezeD under Animal Face and Anime Face datasets in Table 3 and Table 3, respectively. We freeze the discriminator until layer 4. See Appendix A for the ablation study on different layers. FreezeD improves both the best performance and the stability as shown by the best and final FID scores.

We finally compare FreezeD with several previous methods, including scale/shift, GLO, MineGAN, L2-SP, and feature distillation (FD). We choose the weights of L2-SP and FD from {0.1,1,10}\{0.1,1,10\} and simply use 11 for all experiments. We follow the hyperparameters of for GLO, and use 2-layer MLP with ReLU activation for the Miner network. Table 3 presents the FID scores of each method. Feature distillation and qualitative results are in Appendix B and C, respectively. Scale/shift and L2-SP are too restrictive and thus harms diversity. GLO produces blurry images while MineGAN fails to learn the distribution shift.

2 Conditional GAN

We also demonstrate the results for conditional GANs. We use the SNGAN-projection architecture pre-trained on ImageNet dataset, and fine-tune it on Oxford Flower , CUB-200-2011 , and Caltech-256 datasets. Each dataset contains 102, 200, and 256 classes, respectively, where each class has 50-100 samples. We use the public pre-trained modelhttps://github.com/pfnet-research/sngan_projection of resolution 128×\times128 and fine-tune the networks following the original training scheme for 20,000 iterations. SNGAN-projection has a larger variance than StyleGAN, but still the trend is similar.

Figure 3 visualizes the samples generated using the model trained by fine-tuning and FreezeD. FreezeD generates more class-consistent samples than fine-tuning as shown in the 2nd and 8th rows. See Appendix E for more qualitative results. We also evaluate the FID scores of the vanilla fine-tuning and FreezeD in Table 4. We freeze the discriminator until {3, 2, 1} layers for {Oxford Flower, CUB-200-2011, Caltech-256 datasets}, respectively, as the distribution shift goes larger. See Appendix A for details. FreezeD improves both the performance and stability for most cases, but harms the stability for Oxford Flower. We find that feature distillation shows more stable results in our experiments. We leave this investigation for future work.

Conclusion

We have introduced a simple yet effective baseline, FreezeD, for transfer learning of GANs. FreezeD splits the discriminator into a feature extractor and a classifier and then fine-tune the classifier only. We demonstrate that this simple baseline clearly outperforms most of the previous methods using various architectures and datasets. Our observation raises two questions. First, the transferability of the feature extractor of the discriminator could be applied for the universal detector of generated images . Second, one can design a more sophisticated method that outperforms our proposed baseline. We hypothesize that the advanced version of feature distillation could be a promising direction.

References

Appendix A Ablation Study on Freezing Layers

We study the effect of freezing layers of the discriminator for StyleGAN and SNGAN-projection in Table 5 and Table 6, respectively. In StyleGAN, layer 4 consistently shows the best performance. However, in SNGAN-projection, layer {3, 2, 1} were the best for Oxford Flower, CUB-200-2011, and Caltech-256 datasets, respectively. It is since Caltech-256 is harder to learn compared to Oxford Flower (i.e., distribution shift is larger). Intuitively, one should less restrict the model to adapt to the large distribution shift. One can also see that FreezeD is less stable than fine-tuning for the Oxford Flower dataset. We observe that feature distillation shows better stability while showing a similar best performance in our early experiments. Investigating a more sophisticated method would be an interesting research direction.

Appendix B Comparison to Feature Distillation

We compare FreezeD with feature distillation. We linearize the activations of the ii-th layer of the discriminator, and match the activations of the source and target discriminators. Since the activation has a different size for each layer, we use the L2-norm normalized by the feature dimension. We simply use 11 for the weight of the regularizer regardless of the layer. Table 7 presents the comparison results. Feature distillation and FreezeD shows comparable results, while feature distillation is twice slower. Hence, we choose to FreezeD as the baseline for this paper.

Appendix C Qualitative Results for Prior Methods

We visualize the samples generated by the prior methods in Figure 4. Scale/shift and L2-SP generates reasonable samples, but have less diversity as measured by FID scores. GLO generates blurry images due to the lack of adversarial loss and the knowledge of source discriminator. In our experiments, MineGAN totally fails to adapt to the target distribution. Note that MineGAN assumes the source distribution covers (or at least close to) the target distribution (e.g., adult faces to child faces as in the original paper ), but cannot be applied if the distributions have disjoint support (e.g., human faces to dog faces).

Appendix D Generated Samples by StyleGAN

Appendix E Generated Samples by SNGAN-projection