MineGAN: effective knowledge transfer from GANs to target domains with few images

Yaxing Wang, Abel Gonzalez-Garcia, David Berga, Luis Herranz, Fahad Shahbaz Khan, Joost van de Weijer

Introduction

Generative adversarial networks (GANs) can learn the complex underlying distribution of image collections . They have been shown to generate high-quality realistic images and are used in many applications including image manipulation , style transfer , compression , and colorization . It is known that high-quality GANs require a significant amount of training data and time. For example, Progressive GANs are trained on 30K images and are reported to require a month of training on one NVIDIA Tesla V100. Being able to exploit these high-quality pretrained models, not just to generate the distribution on which they are trained, but also to combine them with other models and adjust them to a target distribution is a desirable objective. For instance, it might be desirable to only generate women using a GAN trained to generate men and women alike. Alternatively, one may want to generate smiling people from two pretrained generative models, one for men and one for women. The focus of this paper is on performing these operations using only a small target set of images, and without access to the large datasets used to pretrain the models.

Transferring knowledge to domains with limited data has been extensively studied for discriminative models , enabling the re-use of high-quality networks. However, knowledge transfer for generative models has received significantly less attention, possibly due to its great difficulty, especially when transferring to target domains with few images. Only recently,Wang et al. studied finetuning from a single pretrained generative model and showed that it is beneficial for domains with scarce data. However, Noguchi and Harada observed that this technique leads to mode collapse. Instead, they proposed to reduce the number of trainable parameters, and only finetune the learnable parameters for the batch normalization (scale and shift) of the generator. Despite being less prone to overfitting, their approach severely limits the flexibility of the knowledge transfer.

In this paper, we address knowledge transfer by adapting a trained generative model for targeted image generation given a small sample of the target distribution. We introduce the process of mining of GANs. This is performed by a miner network that transforms a multivariate normal distribution into a distribution on the input space of the pretrained GAN in such a way that the generated images resemble those of the target domain. The miner network has considerably fewer parameters than the pretrained GAN and is therefore less prone to overfitting. The mining step predisposes the pretrained GAN to sample from a narrower region of the latent distribution that is closer to the target domain, which in turn eases the subsequent finetuning step by providing a cleaner training signal with lower variance (in contrast to sampling from the whole source latent space as in ). Consequently, our method preserves the adaptation capabilities of finetuning while preventing overfitting.

Importantly, our mining approach enables transferring from multiple pretrained GANs, which allows us to aggregate information from multiple sources simultaneously to generate samples akin to the target domain. We show that these networks can be trained by a selective backpropagation procedure. Our main contributions are:

We introduce a novel miner network to steer the sampling of the latent distribution of a pretrained GAN to a target distribution determined by few images.

We propose the first method to transfer knowledge from multiple GANs to a single generative model.

We outperform existing competitors on a variety of settings, including transferring knowledge from unconditional, conditional, and multiple GANs.

Related work

Generative adversarial networks. GANs consists of two modules: generator and discriminator . The generator aims to generate images to fool the discriminator, while the discriminator aims to distinguish generated from real images. Training GANs was initially difficult, due to mode collapse and training instability. Several methods focus on addressing these problems , while another major line of research aims to improve the architectures to generate higher quality images . For example, Progressive GAN generates better images by synthesizing them progressively from low to high-resolution. Finally, BigGAN successfully performs conditional high-realistic generation from ImageNet .

Transfer learning for GANs. While knowledge transfer has been widely studied for discriminative models in computer vision , only a few works have explored transferring knowledge for generative models . Wang et al. investigated finetuning of pretrained GANs, leading to improved performance for target domains with limited samples. This method, however, suffers from mode collapse and overfitting, as it updates all parameters of the generator to adapt to the target domain. Recently, Noguchi and Harada proposed to only update the batch normalization parameters. Although less susceptible to mode collapse, this approach significantly reduces the adaptation flexibility of the model since changing only the parameters of the batch normalization permits for style changes but is not expected to function when shape needs to be changed. They also replaced the GAN loss with a mean square error loss. As a result, their model only learns the relationship between latent vectors and sparse training samples, requiring the input noise distribution to be truncated during inference to generate realistic samples. The proposed MineGAN does not suffer from this drawback, as it learns how to automatically adapt the input distribution. In addition, we are the first to consider transferring knowledge from multiple GANs to a single target domain.

Iterative image generation. Nguyen et al. have investigated training networks to generate images that maximize the activation of neurons in a pretrained classification network. In a follow-up approach that improves the diversity of the generated images, they use this technique to generate images of a particular class from a pretrained classifier network. In principle, these works do not aim at transferring knowledge to a new domain, and can instead only be applied to generate a distribution that is exactly described by one of the class labels of the pretrained classifier network. Another major difference is that the generation at inference time of each image is an iterative process of successive backpropagation updates until convergence, whereas our method is feedforward during inference.

Mining operations on GANs

Assume we have access to one or more pretrained GANs and wish to use their knowledge to train a new GAN for a target domain with few images. For clarity’s sake, we first introduce mining from a single GAN in Section 3.2, but our method is general for an arbitrary number of pretrained GANs, as explained in Section 3.3. Then, we show how the miners can be used to train new GANs (Section 3.4).

Let pdata(x)p_{data}(x) be a probability distribution over real data xx determined by a set of real images D\mathcal{D}, and let pz(z)p_{z}(z) be a prior distribution over an input noise variable zz. The generator GG is trained to synthesize images given z∼pz(z)z\sim p_{z}(z) as input, inducing a generative distribution pg(x)p_{g}(x) that should approximate the real data distribution pdata(x)p_{data}(x). This is achieved through an adversarial game , in which a discriminator DD aims to distinguish between real images and images generated by GG, while the generator tries to generate images that fool DD. In this paper, we follow WGAN-GP , which provides better convergence properties by using the Wasserstein loss and a gradient penalty term (omitted from our formulation for simplicity). The discriminator (or critic) and generator losses are defined as follows:

We also consider families of pretrained generators {Gi}\{G_{i}\}. Each GiG_{i} has the ability to synthesize images given input noise z∼pzi(z)z\sim p_{z}^{i}(z). For simplicity and without loss of generality, we assume the prior distributions are Gaussian, i.e. pzi(z)=N(z∣μi,Σi)p_{z}^{i}(z)=\mathcal{N}(z|\bm{\mu_{i}},\bm{\Sigma_{i}}). Each generator Gi(z)G_{i}(z) induces a learned generative distribution pgi(x)p_{g}^{i}(x), which approximates the corresponding real data distribution pdatai(x)p_{data}^{i}(x) over real data xx given by the set of source domain images Di\mathcal{D}_{i}.

2 Mining from a single GAN

We want to approximate a target real data distribution pdataT(x)p_{data}^{\mathcal{T}}(x) induced by a set of real images DT\mathcal{D}_{\mathcal{T}}, given a critic DD and a generator GG, which have been trained to approximate a source data distribution pdata(x)p_{data}(x) via the generative distribution pg(x)p_{g}(x). The mining operation learns a new generative distribution pgT(x)p_{g}^{\mathcal{T}}(x) by finding those regions in pg(x)p_{g}(x) that better approximate the target data distribution pdataT(x)p_{data}^{\mathcal{T}}(x) while keeping GG fixed. In order to find such regions, mining actually finds a new prior distribution pzT(z)p_{z}^{\mathcal{T}}(z) such that samples G(z)G(z) with z∼pzT(z)z\sim p_{z}^{\mathcal{T}}(z) are similar to samples from pdataT(x)p_{data}^{\mathcal{T}}(x) (see Fig. 1a). For this purpose, we propose a new GAN component called miner, implemented by a small multilayer perceptron MM. Its goal is to transform the original input noise variable u∼pz(u)u\sim p_{z}(u) to follow a new, more suitable prior that identifies the regions in pg(x)p_{g}(x) that most closely align with the target distribution.

Our full method acts in two stages. The first stage steers the latent space of the fixed generator GG to suitable areas for the target distribution. We refer to the first stage as MineGAN (w/o FT) and present the proposed mining architecture in Fig. 1b. The second stage updates the weights of the generator via finetuning (GG is no longer fixed). MineGAN refers to our full method including finetuning.

Miner MM acts as an interface between the input noise variable and the generator, which remains fixed during training. To generate an image, we first sample u∼pz(u)u\sim p_{z}(u), transform it with MM and then input the transformed variable to the generator, i.e. G(M(u))G(M(u)). We train the model adversarially: the critic DD aims to distinguish between fake images output by the generator G(M(u))G(M(u)) and real images xx from the target data distribution pdataT(x)p^{\mathcal{T}}_{data}(x). We implement this with the following modification on the WGAN-GP loss:

The parameters of GG are kept unchanged but the gradients are backgropagated all the way to MM to learn its parameters. This training strategy will gear the miner towards the most promising regions of the input space, i.e. those that generate images close to DT\mathcal{D}_{\mathcal{T}}. Therefore, MM is effectively mining the relevant input regions of prior pz(u)p_{z}(u) and giving rise to a targeted prior pzT(z)p_{z}^{\mathcal{T}}(z), which will focus on these regions while ignoring other ones that lead to samples far off the target distribution pdataT(x)p_{data}^{\mathcal{T}}(x).

We distinguish two types of targeted generation: on-manifold and off-manifold. In the on-manifold case, there is a significant overlap between the original distribution pdata(x)p_{data}(x) and the target distribution pdataT(x)p_{data}^{\mathcal{T}}(x). For example, pdata(x)p_{data}(x) could be the distribution of human faces (both male and female) while pdataT(x)p_{data}^{\mathcal{T}}(x) includes female faces only. On the other hand, in off-manifold generation, the overlap between the two distributions is negligible, e.g. pdataT(x)p_{data}^{\mathcal{T}}(x) contains cat faces. The off-manifold task is evidently more challenging as the miner needs to find samples out of the original distribution (see Fig. 4). Specifically, we can consider the images in D\mathcal{D} to lie on a high-dimensional image manifold that contains the support of the real data distribution pdata(x)p_{data}(x) . For a target distribution farther away from pdata(x)p_{data}(x), its support will be more disjoint from the original distribution’s support, and thus its samples might be off the manifold that contains D\mathcal{D}.

3 Mining from multiple GANs

In the general case, the mining operation is applied on multiple pretrained generators. Given target data DT\mathcal{D}_{\mathcal{T}}, the task consists in mining relevant regions from the induced generative distributions learned by a family of NN generators {Gi}\{G_{i}\}. In this task, we do not have access to the original data used to train {Gi}\{G_{i}\} and can only use target data DT\mathcal{D}_{\mathcal{T}}. Fig. 1c presents the architecture of our model, which extends the mining architecture for a single pretrained GAN by including multiple miners and an additional component called selector. In the following, we present this component and describe the training process in detail.

Supersample. In traditional GAN training, a fake minibatch is composed of fake images G(z)G(z) generated with different samples z∼pz(z)z\sim p_{z}(z). To construct fake minibatches for training a set of miners, we introduce the concept of supersample. A supersample S\mathcal{S} is a set of samples composed of exactly one sample per generator of the family, i.e. S={Gi(z)∣z∼pzi(z);i=1,...,N}\mathcal{S}=\{G_{i}(z)|z\sim p_{z}^{i}(z);i=1,...,N\}. Each minibatch contains KK supersamples, which amounts to a total of K×NK\times N fake images per minibatch.

Selector. The selector’s task is choosing which pretrained model to use for generating samples during inference. For instance, imagine that D1\mathcal{D}_{1} is a set of ‘kitchen’ images and D2\mathcal{D}_{2} are ‘bedroom’ images, and let DT\mathcal{D}_{\mathcal{T}} be ‘white kitchens’. The selector should prioritize sampling from G1G_{1}, as the learned generative distribution pg1(x)p_{g}^{1}(x) will contain kitchen images and thus will naturally be closer to pdataT(x)p_{data}^{\mathcal{T}}(x), the target distribution of white kitchens. Should DT\mathcal{D}_{\mathcal{T}} comprise both white kitchens and dark bedrooms, sampling should be proportional to the distribution in the data.

We model the selector as a random variable ss following a categorical distribution parametrized by p1,...,pNp_{1},...,p_{N} with pi>0p_{i}>0, ∑pi=1\sum p_{i}=1. We estimate the parameters of this distribution as follows. The quality of each sample Gi(z)G_{i}(z) is evaluated by a single critic DD based on its critic value D(Gi(z))D(G_{i}(z)). Higher critic values indicate that the generated sample from GiG_{i} is closer to the real distribution. For each supersample S\mathcal{S} in the minibatch, we record which generator obtains the maximum critic value, i.e. arg max⁡iD(Gi(z))\operatorname*{arg\,max}_{i}D(G_{i}(z)). By accumulating over all KK supersamples and normalizing, we obtain an empirical probability value p^i\hat{p}_{i} that reflects how often generator GiG_{i} obtained the maximum critic value among all generators for the current minibatch. We estimate each parameter pip_{i} as the empirical average p^i\hat{p}_{i} in the last 1000 minibatches. Note that pip_{i} are learned during training and fixed during inference, where we apply a multinomial sampling function to sample the index.

Critic and miner training. We now define the training behavior of the remaining learnable components, namely the critic DD and miners {Mi}\{M_{i}\}, when minibatches are composed of supersamples. The critic aims to distinguish real images from fake images. This is done by looking for artifacts in the fake images which distinguish them from the real ones. Another less discussed but equally important task of the critic is to observe the frequency of occurrence of images: if some (potentially high-quality) image occurs more often among fake images than real ones, the critic will lower its score, and thereby motivate the generator to lower the frequency of occurrence of this image. Training the critic by backpropagating from all images in the supersample prevents it from assessing the frequency of occurrence of the generated images and leading to unsatisfactory results empirically. Therefore, the loss for multiple GAN mining is:

As a result of the max⁡\max operator we only backpropagate from the generated image that obtained the highest critic score. This allows the critic to assess the frequency of occurrence correctly. Using this strategy, the critic can perform both its tasks: boosting the quality of images as well as driving the miner to closely follow the distribution of the target set. Note that we initialize the single critic DD with the pretrained weights from one of the pretrained criticsWe empirically found that starting from any pretrained critic leads to similar results (see Suppl. Mat. (Sec. F)).

Conditional GANs. So far, we have only used unconditional GANs. However, conditional GANs (cGANs), which introduce an additional input variable to condition the generation to the class label, are used by the most successful approaches . Here we extend our proposed MineGAN to cGANs that condition on the batch normalization layer See Suppl. Mat. (Sec. D) for results on another type of conditioning., more concretely, BigGAN (Fig. 2 (left)). First, a label ll is mapped to an embedding vector by means of a class embedding EE, and then this vector is mapped to layer-specific batch normalization parameters. The discriminator is further conditioned via label projection . Fig. 2 (right) shows how to mine BigGANs. Alongside the standard miner MzM^{z}, we introduce a second miner network McM^{c}, which maps from uu to the embedding space, resulting in a generator G(Mc(u),Mz(u))G(M^{c}(u),M^{z}(u)). The training is equal to that of a single GAN and follows Eqs. 3 and 4.

4 Knowledge transfer with MineGAN

The underlying idea of mining is to predispose the pretrained model to the target distribution by reducing the divergence between source and target distributions. The miner network contains relatively few parameters and is therefore less prone to overfitting, which is known to occur when directly finetuning the generator GG . We finalize the knowledge transfer to the new domain by finetuning both the miner MM and generator GG (by releasing its weights). The risk of overfitting is now diminished as the generative distribution is closer to the target, requiring thus a lower degree of parameter adaptation. Moreover, the training is substantially more efficient than directly finetuning the pretrained GAN , where synthesized images are not necessarily similar to the target samples. A mined pretrained model makes the sampling more effective, leading to less noisy gradients and a cleaner training signal.

Experiments

We first introduce the used evaluation measures and architectures. Then, we evaluate our method for knowledge transfer from unconditional GANs, considering both single and multiple pretrained generators. Finally, we assess transfer learning from conditional GANs. Experiments focus on transferring knowledge to target domains with few images.

Evaluation measures. We employ the widely used Fréchet Inception Distance (FID) for evaluation. FID measures the similarity between two sets in the embedding space given by the features of a convolutional neural network. More specifically, it computes the differences between the estimated means and covariances assuming a multivariate normal distribution on the features. FID measures both the quality and diversity of the generated images and has been shown to correlate well with human perception . However, it suffers from instability on small datasets. For this reason, we also employ Kernel Maximum Mean Discrepancy (KMMD) with a Gaussian kernel and Mean Variance (MV) for some experiments . Low KMMD values indicate high quality images, while high values of MV indicate more image diversity.

Baselines. We compare our method with the following baselines. TransferGAN directly updates both the generator and the discriminator for the target domain. VAE is a variational autoencoder trained following , i.e. fully supervised by pairs of latent vectors and training images. BSA updates only the batch normalization parameters of the generator instead of all the parameters. DGN-AM generates images that maximize the activation of neurons in a pretrained classification network. PPGN improves the diversity of DGN-AM by of adding a prior to the latent code via denoising autoencoder. Note that both of DGN-AM and PPGN require the target domain label, and thus we only include them in the conditional setting.

Architectures. We apply mining to Progressive GAN , SNGAN , and BigGAN . The miner has two fully connected layers for MNIST and four layers for all other experiments. More training details in Suppl. Mat. (Sec. A).

MNIST dataset. We show results on MNIST to illustrate the functioning of the miner We add quantitative results on MNIST in Suppl. Mat. (Sec. D). We use 1000 images of size 28×2828\times 28 as target data. We test mining for off-manifold targeted image generation. In off-manifold targeted generation, GG is pre-trained to synthesize all MNIST digits except for the target one, e.g. GG generates 0-8 but not 9. The results in Fig. 3 are after training only the miner, without an additional finetuning step. Interestingly, the miner manages to steer the generator to output samples that resemble the target digits, by merging patterns from other digits in the source set. For example, digit ‘9’ frequently resembles a modified 4 while ‘8’ heavily borrows from 0s and 3s. Some digits can be more challenging to generate, for example, ‘5’ is generally more distinct from other digits and thus in more cases the resulting sample is confused with other digits such as ‘3’. In conclusion, even though target classes are not in the training set of the pretrained GAN, still similar examples might be found on the manifold of the generator.

Single pretrained model. We start by transferring knowledge from a Progressive GAN trained on CelebA . We evaluate the performance on target datasets of varying size, using 1024×10241024\times 1024 images. We consider two target domains: on-manifold, FFHQ women and off-manifold, FFHQ children face . We refer as MineGAN to our full method including finetuning, whereas MineGAN (w/o FT) refers to only applying mining (fixed generator). We use training from Scratch, and the TransferGAN method of as baselines. Figure 5 shows the performance in terms of FID and KMMD as a function of the number of images in the target domain. MineGAN outperforms all baselines. For the on-manifold experiment, MineGAN (w/o FT) outperforms the baselines, and results are further improved with additional finetuning. Interestingly, for the off-manifold experiment, MineGAN (w/o FT) obtains only slightly worse results than TransferGAN, showing that the miner alone already manages to generate images close to the target domain. Fig. 4 shows images generated when the target data contains 100 training images. Training the model from scratch results in overfitting, a pathology also occasionally suffered by TransferGAN. MineGAN, in contrast, generates high-quality images without overfitting and images are sharper, more diverse, and have more realistic fine details.

We also compare here with Batch Statistics Adaptation (BSA) using the same settings and architecture, namely SNGAN . They performed knowledge transfer from a pretrained SNGAN on ImageNet to FFHQ and to Anime Face . Target domains have only 25 images of size 128×\times128. We added our results to those reported in in Fig. 6 (bottom). Compared to BSA, MineGAN (w/o FT) obtains similar KMMD scores, showing that generated images obtain comparable quality. MineGAN outperforms BSA both in KMMD score and Mean Variance. The qualitative results (shown in Fig. 6 (top)) clearly show that MineGAN outperforms the baselines. BSA presents blur artifacts, which are probably caused by the mean square error used to optimize their model.

Multiple pretrained models. We now evaluate the general case for MineGAN, where there is more than one pretrained model to mine from. We start with two pretrained Progressive GANs: one on Cars and one on Buses, both from the LSUN dataset . These pretrained networks generate cars and buses of a variety of different colors. We collect a target dataset of 200 images (of 256×256256\times 256 resolution) of red vehicles, which contains both red cars and red buses. We consider three target sets with different car-bus ratios (0.3:0.7, 0.5:0.5, and 0.7:0.3) which allows us to evaluate the estimated probabilities pip_{i} of the selector. To successfully generate all types of red vehicle, knowledge needs to be transferred from both pre-trained models.

Fig. 7 shows the synthesized images. As expected, the limited amount of data makes training from scratch result in overfitting. TransferGAN produces only high-quality output samples for one of the two classes (the class that coincides with the pretrained model) and it cannot extract knowledge from both pretrained GANs. On the other hand, MineGAN generates high-quality images by successfully transferring the knowledge from both source domains simultaneously. Table 1 (top rows) quantitatively validates that our method outperfroms TransferGAN with a significantly lower FID score. Furthermore, the probability distribution predicted by the selector, reported in Table 1 (bottom rows), matches the class distribution of the target data.

To demonstrate the scalability of MineGAN with multiple pretrained models, we conduct experiments using four different generators, each trained on a different LSUN category including Livingroom, Kitchen, Church, and Bridge. We consider two different off-manifold target datasets, one with Bedroom images and one with Tower images, both containing 200 images. Table 1 (top) shows that our method obtains significantly better FID scores even when we choose the most relevant pretrained GAN to initialize training for TransferGAN. Table 1 (bottom) shows that the miner identifies the relevant pretrained models, e.g. transferring knowledge from Bridge and Church for the target domain Tower. Finally, Fig. 7 (right) provides visual examples.

2 Knowledge transfer from conditional GANs

Here we transfer knowledge from a pretrained conditional GAN (see Section 3.3). We use BigGAN , which is trained using ImageNet , and evaluate on two target datasets: on-manifold (ImageNet: cock, tape player, broccoli, fire engine, harvester) and off-manifold (Places365 : alley, arch, art gallery, auditorium, ballroom). We use 500 images per category. We compare MineGAN with training from scratch, TransferGAN , and two iterative methods: DGN-AM and PPGN We were unable to obtain satisfactory results with BSA in this setting (images suffered from blur artifacts) and have excluded it here.. It should be noted that both DGN-AM and PPGN are based on a less complex GAN (equivalent to DCGAN ). Therefore, we expect these methods to exhibit results of inferior quality, and so the comparison here should be interpreted in the context of GAN quality progress. However, we would like to stress that both DGN-AM and PPGN do not aim to transfer knowledge to new domains. They can only generate samples of a particular class of a pretrained classifier network, and they have no explicit loss ensuring that the generated images follow a target distribution.

Fig. 8 shows qualitative results for the different methods. As in the unconditional case, MineGAN produces very realistic results, even for the challenging off-manifold case. Table 2 presents quantitative results in terms of FID and KMMD. We also indicate whether each method uses the label of the target domain class. Our method obtains the best scores for both metrics, despite not using target label information. PPGN performs significantly worse than our method. TransferGAN has a large performance drop for the off-manifold case, for which it cannot use the target label as it is not in the pretrained GAN (see for details).

Another important point regarding DGN-AM and PPGN is that each image generation during inference is an iterative process of successive backpropagation updates until convergence, whereas our method is feedforward. For this reason, we include in Table 2 the inference running time of each method, using the default 200 iterations for DGN-AM and PPGN. All timings have been computed with a CPU Intel Xeon E5-1620 v3 @ 3.50GHz and GPU NVIDIA RTX 2080 Ti. We can clearly observe that the feedforward methods (TransferGAN and ours) are three orders of magnitude faster despite being applied on a more complex GAN .

Conclusions

We presented a model for knowledge transfer for generative models. It is based on a mining operation that identifies the regions on the learned GAN manifold that are closer to a given target domain. Mining leads to more effective and efficient fine tuning, even with few target domain images. Our method can be applied to single and multiple pretrained GANs. Experiments with various GAN architectures (BigGAN, Progressive GAN, and SNGAN) on multiple datasets demonstrated its effectiveness. Finally, we demonstrated that MineGAN can be used to transfer knowledge from multiple domains.

Acknowledgements. We acknowledge the support from Huawei Kirin Solution, the Spanish projects TIN2016-79717-R and RTI2018-102285-A-I0, the CERCA Program of the Generalitat de Catalunya, and the EU Marie Sklodowska-Curie grant agreement No.6655919.

References

Appendix A Architecture and training details

MNIST dataset. Our model contains a miner, a generator and a discriminator. For both unconditional and conditional GANs, we use the same framework to design the generator and discriminator. The miner is composed of two fully connected layers with the same dimensionality as the latent space ∣z∣|z|. The visual results are computed with ∣z∣=16|z|=16; we found that the quantitative results improved for larger ∣z∣|z| and choose ∣z∣=128|z|=128.

We randomly initialize the weights of each miner following a Gaussian distribution centered at 0 with 0.01 standard deviation, and optimize the model using Adam with batch size of 64. The learning rate of our model is 0.0004, with exponential decay rates of (β1,β2)=(0.5,0.999)\left(\beta_{1},\beta_{2}\right)=\left(0.5,0.999\right).

In the conditional MNIST case, label cc is a one-hot vector. This differs from the conditioning used for BigGAN explained in Section 3.3.

Here we extend MineGAN to this type of conditional models by considering each possible conditioning as an independently pretrained generator and using the selector to predict the conditioning label. Given a conditional generator G(c,z)G(c,z), we consider G(i,z)G(i,z) as GiG_{i} and apply the presented MineGAN approach for multiple pretrained generators on the family {G(i,z)∣ i=1,...,N}\{G(i,z)|\text{ }i=1,...,N\}. The resulting selector now chooses among the NN classes of the model rather than among NN pretrained models, but the rest of the MineGAN training remains the same, including the training of NN independent miners.

CelebA Women, FFHQ Children and LSUN (Tower and Bedroom) datasets. We design the generator and discriminator based on Progressive GANs . Both networks use a multi-scale technique to generate high-resolution images. Note we use a simple miner for tasks with dense and narrow source domains. The miner comprises out of four fully connected layers (8-64-128-256-512), each of which is followed by a ReLU activation and pixel normalization except for last layer. We use a Gaussian distribution centered at 0 with 0.01 standard deviation to initialize the miner, and optimize the model using Adam with batch size of 4. The learning rate of our model is 0.0015, with exponential decay rates of (β1,β2)=(0,0.99)\left(\beta_{1},\beta_{2}\right)=\left(0,0.99\right).

FFHQ Face and Anime Face datasets. We use the same network as , namely SNGAN. The miner consists of three fully connected layers (8-32-64-128). We randomly initialize the weights following a Gaussian distribution centered at 0 with 0.01 standard deviation. For this additional set of experiments, we use Adam with a batch size of 8, following a hyper parameter learning rate of 0.0002 and exponential decay rate of (β1,β2)=(0,0.9)\left(\beta_{1},\beta_{2}\right)=\left(0,0.9\right).

Imagenet and Places365 datasets. We use the pretrained BigGAN . We ignore the projection loss in the discriminator, since we do not have access to the label of the target data. We employ a more powerful miner in order to allocate more capacity to discover the regions related to target domain. The miner consists of two sub-networks: miner MzM^{z} and miner McM^{c}. Both MzM^{z} and McM^{c} are composed of four fully connected layers of sizes (128, 128)-(128, 128)-(128, 128)-(128, 128)-(128, 120) and (128, 128)-(128, 128)-(128, 128)-(128, 128)-(128, 128), respectively. We use Adam with a batch size of 256, and learning rates of 0.0001 for the miner and the generator and 0.0004 for the discriminator. The exponential decay rates are (β1,β2)=(0,0.999)\left(\beta_{1},\beta_{2}\right)=\left(0,0.999\right). We randomly initialize the weights following a Gaussian distribution centered at 0 with 0.01 standard deviation.

The input of BigGAN is a random latent vector and a class label that is mapped to an embedding space. We therefore have two miner networks, the original one that maps to the input latent space (MzM^{z}) and a new one that maps to the latent class embedding (McM^{c}). Note that since we have no class label, the miner (McM^{c}) needs to learn what distribution over the embeddings best represents the target data.

Appendix B Evaluation metric details

Similarly to , we compute FID between 10,000 randomly generated images and 10,000 real images, if possible. When the number of target image exceeds 10,000, we randomly select a subset containing only 10,000. On the other hand, if the target set contains fewer images, we cap the amount of randomly generated images to this number to compute the FID. Note we also consider KMMD to evaluate the distance between generated images and real images since FID suffers from instability on small datasets.

Appendix C Ablation study

In this section, we evaluate the effect of each independent contribution to MineGAN and their combinations.

Selection strategies. We ablate the use of the max⁡\max operation in Eqs. (5) and (6), and replace it with a mean operation. We call this setting MineGAN (mean). In this case the backpropagation is not only performed for the image with the highest critic score but for all images. We hypothesized that the max⁡\max operation is necessary to correctly estimate the probabilities used by the selector. We report the results in Table 3, where we refer to our original model as MineGAN (max⁡\max). The probability distribution predicted by the selector indicates that MineGAN (mean) equally chooses both pretrained models on Car and Bus, while MineGAN successfully estimates the class distribution of the target data.

We performed an ablation study for different variants of the miner based on pretrained BigGAN. The miner always contains only fully connected layers, but we experiment with varying the number of layers. In Fig. 9, we show the results on off-manifold target class Arch from Places365 . Using more layers increases the performance of the method. Besides, we find that the results of both 3 and 4 layers are similar, indicating that adding additional layers would only result in slight improvements for our model. Therefore, in this paper, we use miners with 4 fully connected layers.

Appendix D MNIST experiment

We expand the MNIST experiments presented in Section 5.1 by providing a quantitative evaluation and including results on conditional GANs. As evaluation measures, we use FID (Section 5) and classifier error . To compute classifier error, we first train a CNN classifier on real training data to distinguish between multiple classes (e.g. digit classifier). Then, we classify the generated images that should belong to a particular class and measure the error as the percentage of misclassified images. This gives us an estimation of how realistic and accurate the generated images are in the context of targeted generation.

Table 4 presents the results for both unconditional and conditional models, using a noise length of ∣z∣=128|z|=128. The relatively low error values indicate that the miner manages to identify the correct regions for generating the target digits. The conditional model offers better results than the unconditional one by selecting the target class more often. We can also observe that the off-manifold task is more difficult than the on-manifold task, as indicated by the higher evaluation scores. However, the off-manifold scores are still reasonably low, indicating that the miner manages to find suitable regions from other digits by mining local patterns shared with the target. Overall, these results indicate the effectiveness of mining on MNIST for both types of targeted image generation. In addition, in Fig. 10 we have added a visualization for the off-manifold MNIST classes which were not already shown in Fig. 2.

Appendix E Further results on CelebA

We provide additional results for the on-manifold experiment CelebA→\rightarrowFFHQ women in Fig. 11, and the off-manifold CelebA→\rightarrowFFHQ children in Fig. 12. In addition, we have also performed an on-manifold experiment with CelebA→\rightarrowCelebA women, whose results are provided in Fig. 13.

Appendix F Further results for LSUN

We provide additional results for the experiment ({\{Bus, Car}\}) →\rightarrow Red vehicles in Fig. 16 and for the experiment {\{Bedroom, Bridge, Church, Kitchen}\} →\rightarrow Tower/Bedroom in Fig. 17. When applying MineGAN to multiple pretrained GANs, we use one of the domains to initialize the weights of the critic. In Fig. 17 we used Church to initialize the critic in case of the target set Tower, and Kitchen to initialize the critic for the target set Bedroom. We found this choice to be of little influence on the final results. When using Kitchen to initialize the critic for target set Tower results change from 62.4 to 61.7. When using Church to initialize the critic for target set Bedroom results change from 54.7 to 54.3.