Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image Synthesis

Bingchen Liu, Yizhe Zhu, Kunpeng Song, Ahmed Elgammal

Introduction

The fascinating ability to synthesize images using the state-of-the-art (SOTA) Generative Adversarial Networks (GANs) (Goodfellow et al., 2014) display a great potential of GANs for many intriguing real-life applications, such as image translation, photo editing, and artistic creation. However, expensive computing cost and the vast amount of required training data limit these SOTAs in real applications with only small image sets and low computing budgets.

In real-life scenarios, the available samples to train a GAN can be minimal, such as the medical images of a rare disease, a particular celebrity’s portrait set, and a specific artist’s artworks. Transfer-learning with a pre-trained model (Mo et al., 2020; Wang et al., 2020) is one solution for the lack of training images. Nevertheless, there is no guarantee to find a compatible pre-training dataset. Furthermore, if not, fine-tuning probably leads to even worse performance (Zhao et al., 2020).

In a recent study, it was highlighted that in art creation applications, most artists prefers to train their models from scratch based on their own images to avoid biases from fine-tuned pre-trained model. Moreover, It was shown that in most cases artists want to train their models with datasets of less than 100 images (Elgammal et al., 2020). Dynamic data-augmentation (Karras et al., 2020a; Zhao et al., 2020) smooths the gap and stabilizes GAN training with fewer images. However, the computing cost from the SOTA models such as StyleGAN2 (Karras et al., 2020b) and BigGAN (Brock et al., 2018) remain to be high, especially when trained with the image resolution on 1024×10241024\times 1024.

In this paper, our goal is to learn an unconditional GAN on high-resolution images, with low computational cost and few training samples. As summarized in Fig. 2, these training conditions expose the model to a high risk of overfitting and mode-collapse (Arjovsky & Bottou, 2017; Zhang & Khoreva, 2018). To train a GAN given the demanding training conditions, we need a generator GG that can learn fast, and a discriminator DD that can continuously provide useful signals to train GG. To address these challenges, we summarize our contribution as:

We design the Skip-Layer channel-wise Excitation (SLE) module, which leverages low-scale activations to revise the channel responses on high-scale feature-maps. SLE allows a more robust gradient flow throughout the model weights for faster training. It also leads to an automated learning of a style/content disentanglement like StyleGAN2.

We propose a self-supervised discriminator DD trained as a feature-encoder with an extra decoder. We force DD to learn a more descriptive feature-map covering more regions from an input image, thus yielding more comprehensive signals to train GG. We test multiple self-supervision strategies for DD, among which we show that auto-encoding works the best.

We build a computational-efficient GAN model based on the two proposed techniques, and show the model’s robustness on multiple high-fidelity datasets, as demonstrated in Fig. 1.

Related Works

Speed up the GAN training: Speeding up the training of GAN has been approached from various perspectives. Ngxande et al. propose to reduce the computing time with depth-wise convolutions. Zhong et al. adjust the GAN objective into a min-max-min problem for a shorter optimization path. Sinha et al. suggest to prepare each batch of training samples via a core-set selection, leverage the better data preparation for a faster convergence. However, these methods only bring a limited improvement in training speed. Moreover, the synthesis quality is not advanced within the shortened training time.

Train GAN on high resolution: High-resolution training for GAN can be problematic. Firstly, the increased model parameters lead to a more rigid gradient flow to optimize GG. Secondly, the target distribution formed by the images on 1024×10241024\times 1024 resolution is super sparse, making GAN much harder to converge. Denton et al. (2015); Zhang et al. (2017); Huang et al. (2017); Wang et al. (2018); Karras et al. (2019); Karnewar & Wang (2019); Karras et al. (2020b); Liu et al. (2020a) develop the multi-scale GAN structures to alleviate the gradient flow issue, where GG outputs images and receives feedback from several resolutions simultaneously. However, all these approaches further increase the computational cost, consuming even more GPU memory and training time.

Stabilize the GAN training: Mode-collapse on GG is one of the big challenges when training GANs. And it becomes even more challenging given fewer training samples and a lower computational budget (a smaller batch-size). As DD is more likely to be overfitting on the datasets, thus unable to provide meaningful gradients to train GG (Gulrajani et al., 2017).

Prior works tackle the overfitting issue by seeking a good regularization for DD, including different objectives (Arjovsky et al., 2017; Lim & Ye, 2017; Tran et al., 2017); regularizing the gradients (Gulrajani et al., 2017; Mescheder et al., 2018); normalizing the model weights (Miyato et al., 2018); and augmenting the training data (Karras et al., 2020a; Zhao et al., 2020). However, the effects of these methods degrade fast when the training batch-size is limited, since appropriate batch statistics can hardly be calculated for the regularization (normalization) over the training iterations.

Meanwhile, self-supervision on DD has been shown to be an effective method to stabilize the GAN training as studied in Tran et al. (2019); Chen et al. (2019). However, the auxiliary self-supervision tasks in prior works have limited using scenario and image domain. Moreover, prior works only studied on low resolution images (32232^{2} to 1282128^{2}), and without a computing resource limitation.

Method

We adopt a minimalistic design for our model. In particular, we use a single conv-layer on each resolution in GG, and apply only three (input and output) channels for the conv-layers on the high resolutions (≥512×512\geq 512\times 512) in both GG and DD. Fig. 3 and Fig. 4 illustrate the model structure for our GG and DD, with descriptions of the component layers and forward flow. These structure designs make our GAN much smaller than SOTA models and substantially faster to train. Meanwhile, our model remains robust on small datasets due to its compact size with the two proposed techniques.

For synthesizing higher resolution images, the generator GG inevitably needs to become deeper, with more conv-layers, in concert with the up-sampling needs. A deeper model with more convolution layers leads to a longer training time of GAN, due to the increased number of model parameters and a weaker gradient flow through GG (Zhang et al., 2017; Karras et al., 2017; Karnewar & Wang, 2019). To better train a deep model, He et al. design the Residual structure (ResBlock), which uses a skip-layer connection to strengthen the gradient signals between layers. However, while ResBlock has been widely used in GAN literature (Wang et al., 2018; Karras et al., 2020b), it also increases the computation cost.

We reformulate the skip-connection idea with two unique designs into the Skip-Layer Excitation module (SLE). First, ResBlock implements skip-connection as an element-wise addition between the activations from different conv-layers. It requires the spatial dimensions of the activations to be the same. Instead of addition, we apply channel-wise multiplications between the activations, eliminating the heavy computation of convolution (since one side of the activations now has a spatial dimension of 121^{2}). Second, in prior GAN works, skip-connections are only used within the same resolution. In contrast, we perform skip-connection between resolutions with a much longer range (e.g., 828^{2} and 1282128^{2}, 16216^{2} and 2562256^{2}), since an equal spatial-dimension is no longer required. The two designs make SLE inherits the advantages of ResBlock with a shortcut gradient flow, meanwhile without an extra computation burden.

Formally, we define the Skip-Layer Excitation module as:

Here x\mathbf{x} and y\mathbf{y} are the input and output feature-maps of the SLE module, the function F\mathcal{F} contains the operations on xlow\mathbf{x}_{low}, and Wi\mathbf{W}_{i} indicates the module weights to be learned. The left panel in Fig. 3 shows an SLE module in practice, where xlow\mathbf{x}_{low} and xhigh\mathbf{x}_{high} are the feature-maps at 8×88\times 8 and 128×128128\times 128 resolution respectively. An adaptive average-pooling layer in F\mathcal{F} first down-samples xlow\mathbf{x}_{low} into 4×44\times 4 along the spatial-dimensions, then a conv-layer further down-samples it into 1×11\times 1. A LeakyReLU is used to model the non-linearity, and another conv-layer projects xlow\mathbf{x}_{low} to have the same channel size as xhigh\mathbf{x}_{high}. Finally, after a gating operation via a Sigmoid function, the output from F\mathcal{F} multiplies xhigh\mathbf{x}_{high} along the channel dimension, yielding y\mathbf{y} with the same shape as xhigh\mathbf{x}_{high}.

2 Self-supervised discriminator

Our approach to provide a strong regularization for DD is surprisingly simple. We treat DD as an encoder and train it with small decoders. Such auto-encoding training forces DD to extract image features that the decoders can give good reconstructions. The decoders are optimized together with DD on a simple reconstruction loss, which is only trained on real samples:

where f\mathbf{f} is the intermediate feature-maps from DD, the function G\mathcal{G} contains the processing on f\mathbf{f} and the decoder, and the function T\mathcal{T} represents the processing on sample xx from real images IrealI_{real}.

Our self-supervised DD is illustrated in Fig. 4, where we employ two decoders for the feature-maps on two scales: f1\mathbf{f}_{1} on 16216^{2} and f2\mathbf{f}_{2} on 828^{2} . The decoders only have four conv-layers to produce images at 128×128128\times 128 resolution, causing little extra computations (much less than other regularization methods). We randomly crop f1\mathbf{f}_{1} with 18\frac{1}{8} of its height and width, then crop the real image on the same portion to get IpartI_{part}. We resize the real image to get II. The decoders produce Ipart′I^{\prime}_{part} from the cropped f1\mathbf{f}_{1}, and I′I^{\prime} from f2\mathbf{f}_{2}. Finally, DD and the decoders are trained together to minimize the loss in eq. 2, by matching Ipart′I^{\prime}_{part} to IpartI_{part} and I′I^{\prime} to II.

Such reconstructive training makes sure that DD extracts a more comprehensive representation from the inputs, covering both the overall compositions (from f2\mathbf{f}_{2}) and detailed textures (from f1\mathbf{f}_{1}). Note that the processing in G\mathcal{G} and T\mathcal{T} are not limited to cropping; more operations remain to be explored for better performance. The auto-encoding approach we employ is a typical method for self-supervised learning, which has been well recognized to improve the model robustness and generalization ability (He et al., 2020; Hendrycks et al., 2019; Jing & Tian, 2020; Goyal et al., 2019). In the context of GAN, we find that a regularized DD via self-supervision training strategies significantly improves the synthesis quality on GG, among which auto-encoding brings the most performance boost.

Although our self-supervision strategy for DD comes in the form of an auto-encoder (AE), this approach is fundamentally different from works trying to combine GAN and AE (Larsen et al., 2016; Guo et al., 2019; Zhao et al., 2016; Berthelot et al., 2017). The latter works mostly train GG as a decoder on a learned latent space from DD, or treat the adversarial training with DD as an supplementary loss besides AE’s training. In contrast, our model is a pure GAN with a much simpler training schema. The auto-encoding training is only for regularizing DD, where GG is not involved.

In sum, we employ the hinge version of the adversarial loss (Lim & Ye (2017); Tran et al. (2017)) to iteratively train our D and G. We find the different GAN losses make little performance difference, while hinge loss computes the fastest:

Experiment

Datasets: We conduct experiments on multiple datasets with a wide range of content categories. On 256×256256\times 256 resolution, we test on Animal-Face Dog and Cat (Si & Zhu, 2011), 100-Shot-Obama, Panda, and Grumpy-cat (Zhao et al., 2020). On 1024×10241024\times 1024 resolution, we test on Flickr-Face-HQ (FFHQ) (Karras et al., 2019), Oxford-flowers (Nilsback & Zisserman, 2006), art paintings from WikiArt (wikiart.org), photographs on natural landscape from Unsplash (unsplash.com), Pokemon (pokemon.com), anime face, skull, and shell. These datasets are designed to cover images with different characteristics: photo realistic, graphic-illustration, and art-like images.

Metrics: We use two metrics to measure the models’ synthesis performance: 1) Fréchet Inception Distance (FID) (Heusel et al., 2017) measures the overall semantic realism of the synthesized images. For datasets with less than 1000 images (most only have 100 images), we let GG generate 5000 images and compute FID between the synthesized images and the whole training set. 2) Learned perceptual similarity (LPIPS) (Zhang et al., 2018) provides a perceptual distance between two images. We use LPIPS to report the reconstruction quality when we perform latent space back-tracking on GG given real images, and measure the auto-encoding performance. We find it unnecessary to involve other metrics, as FID is unlikely to be inconsistent with the others, given the notable performance gap between our model and the compared ones. For all the testings, we train the models 5 times with random seeds, and report the highest scores. The relative error is less than five percent on average.

Compared Models: We compare our model with: 1) the state-of-the-art (SOTA) unconditional model, StyleGAN2, 2) a baseline model ablated from our proposed one. Note that we adopt StyleGAN2 with recent studies from (Karras et al., 2020a; Zhao et al., 2020), including the model configuration and differentiable data-augmentation, for the best training on few-sample datasets. Since StyleGAN2 requires much more computing-cost (cc) to train, we derive an extra baseline model. In sum, we compare our model with StyleGAN2 on the absolute image synthesis quality regardless of cc, and use the baseline model for the reference within a comparable cc range.

The baseline model is the strongest performer that we integrated from various GAN techniques based on DCGAN (Radford et al., 2015): 1) spectral-normalization (Miyato et al., 2018), 2) exponential-moving-average (Yazıcı et al., 2018) optimization on GG, 3) differentiable-augmentation, 4) GLU (Dauphin et al., 2017) instead of ReLU in GG. We build our model upon the baseline with the two proposed techniques: the skip-layer excitation module and the self-supervised discriminator.

Table. 1 presents the normalized cc figures of the models on Nvidia’s RTX 2080-Ti GPU, implemented using PyTorch (Paszke et al., 2017). Importantly, the slimed StyleGAN2 with 14\frac{1}{4} parameters cannot converge on the tested datasets at 102421024^{2} resolution. We compare to the StyleGAN2 with 12\frac{1}{2} parameters (if not specifically mentioned) in the following experiments.

Few-shot generation: Collecting large-scale image datasets are expensive, or even impossible, for a certain character, a genre, or a topic. On those few-shot datasets, a data-efficient model becomes especially valuable for the image generation task. In Table. 2 and Table. 3, we show that our model not only achieves superior performance on the few-shot datasets, but also much more computational-efficient than the compared methods. We save the checkpoints every 10k iterations during training and report the best FID from the checkpoints (happens at least after 15 hours of training for StyleGAN2 on all datasets). Among the 12 datasets, our model performs the best on 10 of them.

Please note that, due to the VRAM requirement for StyleGAN2 when trained on 102421024^{2} resolution, we have to train the models in Table. 3 on a RTX TITAN GPU. In practice, 2080-TI and TITAN share a similar performance, and our model runs the same time on both GPUs.

Training from scratch vs. fine-tuning: Fine-tuning from a pre-trained GAN (Mo et al., 2020; Noguchi & Harada, 2019; Wang et al., 2020) has been the go-to method for the image generation task on datasets with few samples. However, its performance highly depends on the semantic consistency between the new dataset and the available pre-trained model. According to Zhao et al., fine-tuning performs worse than training from scratch in most cases, when the content from the new dataset strays away from the original one. We confirm the limitation of current fine-tuning methods from Table. 2 and Table. 3, where we fine-tune StyleGAN2 trained on FFHQ use the Freeze-D method from Mo et al.. Among all the tested datasets, only Obama and Skull favor the fine-tuning method, making sense since the two sets share the most similar contents to FFHQ.

Module ablation study: We experiment with the two proposed modules in Table. 2, where both SLE (skip) and decoding-on-DD (decode) can separately boost the model performance. It shows that the two modules are orthogonal to each other in improving the model performance, and the self-supervised DD makes the biggest contribution. Importantly, the baseline model and StyleGAN2 diverge fast after the listed training time. In contrast, our model is less likely to mode collapse among the tested datasets. Unlike the baseline model which usually model-collapse after trained for 10 hours, our model maintains a good synthesis quality and won’t collapse even after trained for 20 hours. We argue that it is the decoding regularization on DD that prevents the model from divergence.

Training with more images: For more thorough evaluation, we also test our model on datasets with more sufficient training samples, as shown in Table. 4. We train the full StyleGAN2 for around five days on the Art and Photograph dataset with a batch-size of 16 on two TITAN RTX GPUs, and use the latest official figures on FFHQ from Zhao et al.. Instead, we train our model for only 24 hours, with a batch-size of 8 on a single 2080-Ti GPU. Specifically, for FFHQ with all 70000 images, we train our model with a larger batch-size of 32, to reflect an optimal performance of our model.

In this test, we follow the common practice of computing FID by generating 50k images and use the whole training set as the reference distribution. Note that StyleGAN2 has more than double the parameters compared to our model, and trained with a much larger batch-size on FFHQ. These factors contribute to its better performances when given enough training samples and computing power. Meanwhile, our model keeps up well with StyleGAN2 across all testings with a considerably lower computing budget, showing a compelling performance even on larger-scale datasets, and a consistent performance boost over the baseline model.

Qualitative results: The advantage of our model becomes more clear from the qualitative comparisons in Fig. 6. Given the same batch-size and training time, StyleGAN2 either converges slower or suffers from mode collapse. In contrast, our model consistently generates satisfactory images. Note that the best results from our model on Flower, Shell, and Pokemon only take three hours’ training, and for the rest three datasets, the best performance is achieved at training for eight hours. For StyleGAN2 on “shell”, “anime face”, and “Pokemon”, the images shown in Fig. 6 are already from the best epoch, which they match the scores in Table. 2 and Table. 3. For the rest of the datasets, the quality increase from StyleGAN2 is also limited given more training time.

2 More Analysis and Applications

Testing mode collapse with back-tracking: From a well trained GAN, one can take a real image and invert it back to a vector in the latent space of GG, thus editing the image’s content by altering the back-tracked vector. Despite the various back-tracking methods (Zhu et al., 2016; Lipton & Tripathi, 2017; Zhu et al., 2020; Abdal et al., 2019), a well generalized GG is arguably as important for the good inversions. To this end, we show that our model, although trained on limited image samples, still gets a desirable performance on real image back-tracking.

In Table 6, we split the images from each dataset with a training/testing ratio of 9:1, and train GG on the training set. We compute a reconstruction error between all the images from the testing set and their inversions from GG, after the same update of 1000 iterations on the latent vectors (to prevent the vectors from being far off the normal distribution). The baseline model’s performance is getting worse with more training iterations, which reflects mode-collapse on GG. In contrast, our model gives better reconstructions with consistent performance over more training iterations. Fig. 5 presents the back-tracked examples (left-most and right-most samples in the middle panel) given the real images. The smooth interpolations from the back-tracked latent vectors also suggest little mode-collapse of our GG (Radford et al., 2015; Zhao et al., 2020; Robb et al., 2020).

In addition, we show qualitative comparisons in appendix D, where our model maintains a good generation while StyleGAN2 and baseline are model-collapsed.

The self-supervision methods and generalization ability on DD: Apart from the auto-encoding training for DD, we show that DD with other common self-supervising strategies also boost GAN’s performance in our training settings. We test five self-supervision settings, as shown in Table 6, which all brings a substantial performance boost compared to the baseline model. Specifically, setting-a refers to contrastive learning which we treat each real image as a unique class and let DD classify them. For setting-b, we train DD to predict the real image’s original aspect-ratio since they are reshaped to square when fed to DD. Setting-c is the method we employ in our model, which trains DD as an encoder with a decoder to reconstruct real images. To better validate the benefit of self-supervision on DD, all the testings are conducted on full training sets with 10000 images, with a batch-size of 8 to be consistent with Table 4. We also tried training with a larger batch-size of 16, which the results are consistent to the batch-size of 8.

Interestingly, according to Table 6, while setting-c performs the best, combining it with the rest two settings lead to a clear performance downgrade. The similar behavior can be found on some other self-supervision settings, e.g. when follow Chen et al. (2019) with a ”rotation-predicting” task on art-paintings and FFHQ datasets, we observe a performance downgrade even compared to the baseline model. We hypothesis the reason being that the auto-encoding forces DD to pay attention to more areas of the input image, thus extracts a more comprehensive feature-map to describe the input image (for a good reconstruction). In contrast, a classification task does not guarantee DD to cover the whole image. Instead, the task drives DD to only focus on small regions because the model can find class cues from small regions of the images. Focusing on limited regions (i.e., react to limited image patterns) is a typical overfitting behavior, which is also widely happening for DD in vanilla GANs. More discussion can be found in appendix B.

Style mixing like StyleGAN. With the channel-wise excitation module, our model gets the same functionality as StyleGAN: it learns to disentangle the images’ high-level semantic attributes (style and content) in an unsupervised way, from GG’s conv-layers at different scales. The style-mixing results are displayed in Fig. 7, where the top three datasets are 256×256256\times 256 resolution, and the bottom three are 1024×10241024\times 1024 resolution. While StyleGAN2 suffers from converging on the bottom high-resolution datasets, our model successfully learns the style representations along the channel dimension on the “excited” layers (i.e., for feature-maps on 256×256256\times 256, 512×512512\times 512 resolution). Please refer to appendix A and C for more information on SLE and style-mixing.

Conclusion

We introduce two techniques that stabilize the GAN training with an improved synthesis quality, given sub-hundred high-fidelity images and a limited computing resource. On thirteen datasets with a diverse content variation, we show that a skip-layer channel-wise excitation mechanism (SLE) and a self-supervised regularization on the discriminator significantly boost the synthesis performance of GAN. Both proposed techniques require minor changes to a vanilla GAN, enhancing GAN’s practicality with a desirable plug-and-play property. We hope this work can benefit downstream tasks of GAN (Liu et al., 2020c; b; Elgammal et al., 2017) and provide new study perspectives for future research.

References

Appendix A Performance boost from skip-layer excitation

Here we present a more detailed ablation study for the skip-layer excitation (SLE) module. We compare between the baseline model and the baseline equipped with SLE. On four 1024×10241024\times 1024 resolutions datasets: Flower, FFHQ, Shell and Art-paintings, we record the FID performance every 10000 iterations for every model. As shown in Fig. 8, SLE brings a constant performance boost on the baseline model over all iterations.

Our key observation is, SLE speeds up the convergence of GAN, where the most noticeable effect happens at the beginning of the training. In the first 20000 iterations, the generator GG is able to converge faster and reach to a good point where the baseline model needs much more training iterations to reach. On the other hand, although SLE provides a faster convergence on GG, the overall model behavior with SLE seems follow the baseline model quite well, with a slightly better overall performance.

In other words, the lines for the two models are parallel in each sub-plot in Fig. 8. Specifically, on Shell, the model with SLE also collapsed after 60000 iterations training, just like the baseline model. And on the rest three datasets, the FID improves much slower and almost stop changing in the later half training iterations. We think such model behavior makes sense, because the SLE module neither increases the model capacity (have very few parameter increase) nor exert any explicit regularization or guidance on the training of GAN. Therefore, SLE is unlikely to make a big difference after the model reaches a good converged state.

On the other hand, SLE does a good job speeding up the convergence for GG, and improves the performance of GG. More importantly, it is SLE that enables the unsupervised style-content disentanglement for our model, in a simpler and more cost-efficient way than StyleGAN and StyleGAN2.

Appendix B Feature-extraction performance of Discriminator

Here we continue the discussion on the effectiveness of the self-supervised auto-encoding training for the discriminator DD. Specifically, we explore the relationship between the feature-extracting behavior on the discriminator DD and the synthesis performance of GAN . By feature-extracting performance, we mean how comprehensive the feature-maps extracted by DD cover the information from the input images. This feature-extracting performance can be easily checked via an auto-encoding training. In detail, we take DD trained in GAN and fix it, then train a decoder for DD which tries to reconstruct the images from the feature-maps encoded by DD. The intuition is, if DD pays attention to all the regions of an input image, and encode the image with a minimum information lost, then the decoder is easier to reconstruct the images encoded by DD. In contrast, if DD is overfitting and only focus on limited local patterns of the images, the it outputs feature-maps with lost information, thus a decoder is unable to reconstruct the images from DD’s output feature-map.

We extract the second-last layer’s activation for the decoder, which is the one for DD to determine the real/fake of an image. Table 7 shows the result, where we train all the decoders for the same 100000 iterations (all the decoders are converged). Note that such feature-extracting performance on DD does not necessarily imply a better synthesis performance for GG. Moreover, the DD from StyleGAN2 is not comparable to the DD from baseline, since they have totally different model structure and complexity.

Instead, according to Table 7, we can get some interesting information. Firstly, the GAN training is actually making DD performs worse as a feature-encoder. According to row. 3 (StyleGAN2) and row. 4 (baseline), we find that the DD after a GAN training extracts less meaningful features compared to a randomly initialized DD (col. 6 and col. 8). It means that while the GAN training leads DD to find the discriminative features between the real and fake samples, it also effective let DD to ignore quite amount of information from the input images.

Secondly, we compare the baseline model to the ones with self-supervised learning guidance (row. 4,5,6,7). It shows that the self-supervisions on DD indeed lead to a more descriptive feature-extraction compared to the randomly initialization on DD. Moreover, contrastive learning may also result in overfitting, since only a partial image (some local patterns) may be enough for the classification task. In comparison, the reconstruction task is more likely to let DD cover more information from the input images. To our surprise, combining auto-encoding training and the contrastive learning result in a worse performance on DD. It shows that the classification objective affects the auto-encoding objective and changes the behavior of DD, in a negative way.

Last but not least, we do find that a better feature-extracting performance on D result in a better synthesis performance of GAN . And it seems true for both StyleGAN2 and the baseline model. For StyleGAN2 trained on FFHQ, DD trained with more data indeed preserves more information from the input images than DD trained on only 1000 images. For our baseline model, the feature-extracting performance on D aligns well with the respective FID scores. Besides, the self-supervision methods all effectively letting DD extracts more information from the images, compared to the randomly initialization and the vanilla GAN training.

Apart from the observations, we would like to emphasize that the experiments are mostly conducted on few-shot datasets. The results does not give a full picture of the relationship between the feature-extraction performance on DD and the synthesis performance of GAN, further study on larger-scale datasets are required. However, the experiments do validate the effectiveness of the self-supervision strategies on DD for an enhanced performance of GAN, on few-shot datasets.

Appendix C Style-mixing on different resolutions

Here we present more qualitative results on the style-mix performance of our model. For the model trained on 1024×10241024\times 1024 resolution, there are three SLE layers that we can swap the feature-maps between generated samples, and there are two SLE layers for model on 256×256256\times 256 resolution.

Fig. 11 shows the results from the 1000 samples training on Art paintings and FFHQ at 1024×10241024\times 1024 resolution, and the 100 samples Obama at 256×256256\times 256 resolution. In each row, we swap the xlow\mathbf{x}_{low} in the SLE layer from the image in col. 1 to the one from each image on row. 1. The best style-mixing results is achieved when the feature-map swapping is done on all resolutions. And the most effective layer that causes the most style changes is the layer on 128128 resolution. On 256×256256\times 256 resolution, the model behaviors the same, where the SLE on lower resolution makes the most style difference.

On the Art-paintings data, the model performs well on style-mixing, where not only the coloring but also the texture can be controlled. The models transfers the style of flat or pointy brush stroke among the style-mixed synthetic images. However, the model does not perform as well on the FFHQ data. There are some cases where even the hair color can not be properly transferred. We speculate that the worse performance on FFHQ is due to the limited training sample and the dramatically varied background. The is no clear relationship between the front-end face and the background contents given the limited training samples, which confuses the model to disentangle more detailed style attributes. In contrast, Art-paintings have consistent style cues within each image and obvious connections between each object inside a scene, making it arguably easier than the FFHQ data. On the other hand, the model performs great on Obama given a even less 100 training images. It successfully transfers the style for both the face attributes and the background. Learning on 2562256^{2} resolution is a simpler task, and the model capacity is more sufficient on only 100 samples.

Appendix D More Qualitative Comparison

Appendix E Nearest images from training sets

In Table 8, we report the average LPIPS score between the generated samples from our model to their closest real samples ranked by LPIPS score. In comparison, we show the baseline as the LPIPS between real images and their randomly augmented variants (randomly horizontal flipping and random cropping with 0.80.8 spatial portion). We run each experiment 3 times with 100 randomly synthesized samples or real images, and report the lowest one. The std among the trials are usually lower than 0.0050.005. This experiment shows that, instead of memorizing the real images in the training set, our model is able to perceive the features from the real images, and generate images that are different and novel, in terms of compositions, shapes, and color patterns.

Appendix F Decoder result