Labels4Free: Unsupervised Segmentation using StyleGAN

Rameen Abdal, Peihao Zhu, Niloy Mitra, Peter Wonka

Introduction

Given the high quality and photo-realistic results of current generative adversarial networks (GANs), we are witnessing their widespread adaptation for many applications. Examples include various image and video editing tasks, image inpainting , local image editing , low bit-rate video conferencing , image super resolution , image colorization , and extracting 3D models .

While originally conjectured that GANs are merely great at memorizing the training data, recent work in GAN-based image editing demonstrates that GANs learn non-trivial semantic information about a class of objects, e.g., faces or cars. For example, GANs are able to learn the concept of pose and they can show the same, or at least a very similar looking, object with different orientations. Even though the background changes in subtle ways during this editing operations, in this paper we explore to what extent the underlying generator network actually learns the distinction between foreground and background, and then to encourage it to disentangle foreground and background, without explicit mask-level supervision. As an important byproduct, we can extract information from an unsupervised GAN that is useful for general object segmentation. For example, we are able to create a large synthetic dataset for training a state-of-the-art segmentation network and then segment arbitrary face (or horse, car, cat) images into foreground and background without requiring any manually assigned labels (see Fig. 1).

Our implementation is based on StyleGAN , generally considered the state-of-the-art for GANs trained on individual object classes. Our solution is built on two ideas. First, based on our analysis of GAN-based image editing, the features generated by StyleGAN hold a lot of information useful for segmentation, and can be used towards corresponding mask synthesis. Second, the foreground and background should be largely independent and be composited in different ways. The exact coupling between foreground and background is highly non-trivial however and there are multiple ways of decoupling foreground and background that we analyze in our work.

For our solution, we propose to augment the StyleGAN generator architecture with a segmentation branch and to split the generator into a foreground and background network. This enables us to generate soft segmentation masks for the foreground object. In order to facilitate easier training, we propose a training strategy that starts from a fully trained network that only has a single generator and utilizes it towards unsupervised segmentation mask generation.

To summarize, our main contributions are:

A novel architecture modification, a loss function, and a training strategy to split StyleGAN into a foreground and background network.

Generating synthetic datasets for segmentation. Our framework can be used to create a complete dataset of labeled GAN generated images of high quality in an unsupervised manner. This dataset can then be used to train other state-of-the-art segmentation networks to yield compelling results.

Related Work

From the seminal works of Goodfellow et al. and Radford et al. , subsequent GAN research contributed to big improvements in the visual quality of the generated results. The state-of-the-art networks like ProGAN , BigGAN , StyleGAN , StyleGAN2 and StyleGAN2-ada demonstrate superior performance in terms of diversity and quality of the generated samples. While StyleGAN series by Karras et al. has demonstrated high quality and photo-realistic results on human faces using the high quality FFHQ dataset, BigGAN can produce high quality samples using complex datasets like ImageNet. In our work we build on StyleGAN2 which is the current state of the art for many smaller data-sets, including faces.

GAN Interpretability and Image Editing.

GAN interpretability has been an important aspect of the GAN research since the beginning. Some recent works in this domain study the structure of the activation and latent space. For instance GANspace simplifies the latent space of StyleGAN (WW Space) to be linear and extracts meaningful directions using PCA. StyleRig mapped StyleGAN latent space to a riggable face model. StyleFlow studied the non-linear nature of the StyleGAN latent space using normalizing flows and is able produce high-quality sequential edits on the generated and real images. On the other hand, the layer activation based methods try to understand the properties of GANs using an analysis of the activation space. A recent work StyleSpace studies the style parameters of the channels to determine various properties and editing abilities of StyleGAN.

Another interesting approach for interpretability of GANs is image embedding. Abdal et. al. demonstrated high quality embeddings into the extended WW space called W+W^{+} space for real image editing. Subsequent works try to improve upon the embedding quality, e.g. by suggesting new regularizers for the optimization. The embedding methods combined with image editing capabilities of StyleGAN has lead to many applications and even commercial software such as Adobe Photoshop’s Neural Filters . A known problem in the domain of image editing is the disentanglement of the features during an editing operation. While state-of-the-art editing frameworks achieve high quality fine grained edits using supervised and unsupervised approaches, background/ foreground aware embeddings and edits still pose a challenge. Often, the background of the scene changes a lot during the embedding itself or during sequential edits performed on the images. Later in Sec. 4, we show how our unsupervised segmentation framework leads to improved image editing.

Unsupervised Object Segmentation.

Among the few attempts in this domain are works by Xu et al. and Yassine et al. which learn a clustering function in an unsupervised setting. A more recent work, Van et al. adopts a predetermined prior in a contrastive optimization objective to learn pixel embeddings. In the GAN domain, PSeg is currently the only approach to segmenting foreground and background by reformulating and retraining the GAN using relative jittering between the composite and background generators at the cost of quality of samples.

Method

Fig. 2 shows an overview of this architecture.

We expand on the training dynamics of this network in Sec. 3.3.

Features of the pretrained generator G\mathcal{G} already contain sufficient information to segment image pixels into foreground and background. We extend the foreground generator with an Alpha Network (A\mathcal{A}) to learn the alpha mask for the foreground (see Fig. 3). In this context, we overcome multiple challenges.

First, the feature maps from lower layers of the generator need to be upsampled. To this effect, we introduce upsampling blocks in the alpha network.

Second, the number of features per pixel is quite high (e.g., 9000+9000+ for StyleGAN2). We therefore use 1×11\times 1 convolutions for compression and feature selection. This encourages dropping the several features that do not contain segmentation information.

Fourth, the features per pixel have to be processed to output an alpha value. We therefore use several non-linear activation functions in the upsampling blocks and add a sigmoid function at the end to make the output in the range . This yields a lightweight but simplified network to extract segmentation information. In particular, neighboring pixels do not interact with each other.

2 Background Generator

The challenge for the background generator is the training dynamics. Specifically, when initializing with a pre-trained generator, the method fails because the background already includes the foreground object. When pre-training the background generator on a different dataset, our attempts failed because the discriminator can easily detect that the backgrounds are out of distribution (see Supplemental (Appendix A)). We therefore adopted an approach based on the conjecture that the foreground and background image are already composited by StyleGAN.

We start out by seeking channels which are responsible for generating the background pixels in a StyleGAN image. Let G\mathcal{G} be the StyleGAN generator and, ww and nn be the latent and noise variables, respectively. In order to bootstrap the network to identify the channels, we first identify and collect generic StyleGAN backgrounds from multiple sampled images by cropping. We notice that background having, for example, a white, black and blue tinge are common in-domain backgrounds. Then we find the gradient of the objective function: ∥(G(w,n)−x)∥22\|(\mathcal{G}(w,n)-x)\|_{2}^{2} with respect to the tensors in the StyleGAN2 layers at all resolutions, where xx is the upsampled background (crops). The above process allows identifying the layers which are most responsible for deleting an object (e.g., faces, cats, horses) from a composited StyleGAN2 image.

In order to quantify the calculated gradient maps, we the calculate the sum of gradients norm over the channels of the respective layers. We found that the first layer (the constant layer and the first layer in which the W latent is injected excluding the tRGB layer ) has the maximum value of the above measure. Hence, we hypothesise that switching off (i.e., zeroing out) the identified channels would exclude the object information from the tensor representation and produce an approximate background. Fig. 4 and 5 show the sampled backgrounds following the steps described above. Note that we could produce higher quality curated backgrounds by empirically setting some channels to be active based on a threshold of the above measure. However, we noticed that such backgrounds occasionally contained traces of foreground objects. Hence, in the training phase, we adopt a safe strategy and zero out all the channels of the selected layer. We refer to this trimmed background generator as Gbg\mathcal{G}_{\text{bg}}.

3 Training Dynamics

Our unsupervised setup consists of the Alpha Network A\mathcal{A}, a pre-trained StyleGAN generator G\mathcal{G} for the foreground, a modified pre-trained StyleGAN generator Gbg\mathcal{G}_{\text{bg}} for the background, and a weak discriminator D\mathcal{D}.

We freeze the generators Gbg\mathcal{G}_{\text{bg}} and G\mathcal{G} during training and only train the Alpha Network A\mathcal{A} and the discriminator D\mathcal{D} using adversarial training. Note that we do not apply the path regularization to the discriminator D\mathcal{D} and style mixing during the training. Unlike others , our framework does not alter the generation quality of samples. The final composited image given to the Discriminator D\mathcal{D} is:

where MM is the mask predicted by the Alpha Network (i.e., M=A(z)M=\mathcal{A}(z)). We use WW space for the training, and show in Sec. 4 that the method can generalize to W+W^{+} space to handle real images projected into StyleGAN2. In order to robustly train the network and avoid degenerate solutions (e.g., A\mathcal{A} to produce all 1s), we make several changes as described next.

(a) Discriminator D\mathcal{D}: Unlike using a frozen pre-trained generator (G,Gbg\mathcal{G},\mathcal{G}_{\text{bg}}), we train the discriminator from scratch. Our method fails to converge by starting the training from a pre-trained discriminator. We hypothesize that that the pre-trained discriminator is already very strong in detecting the correlations between the foreground and the background . For example, environmental illumination, shadows and reflections in case of cars, cats, and horses. Hence to enable our network to train on independently sampled backgrounds, we use a ‘Weak’ discriminator, in the sense that it does not have a strong prior about the correlation between foregrounds and backgrounds.

(b) Truncation trick: We observed that unlike FFHQ trained StyleGAN2, the samples from the LSUN-Object trained StyleGAN2 have variable quality produced without truncation. To avoid the rejection sampling and ensure high quality samples during the training, we use truncation ψ\psi ∈\in for the generator G\mathcal{G} in case of LSUN-Object training .

(c) Regularizer: As the alpha segmentation network may choose to converge into a sparse map, we use a binary enforcing regularizer, B(M):=min⁡(M,1−M)B(M):=\min(M,1-M) to guide the training. Note that we control this regularizer to still get soft segmentation maps. The truncation trick above may result in backgrounds from Gbg\mathcal{G}_{\text{bg}} more aligned to the original distribution than the composite image from G\mathcal{G}. Hence, for the LSUN-Object, we use another regularizer C(M):=ReLU(ϕ1−1m×m∑iMi)C(M):=\text{ReLU}(\phi_{1}-\frac{1}{m\times m}\sum_{i}M_{i}) which ensures that the optimization does not converge to a trivial solution of only using the background network. Similarly, the regularizer as E(M):=ReLU(ϕ2−1m×m∑i(1−Mi))E(M):=\text{ReLU}(\phi_{2}-\frac{1}{m\times m}\sum_{i}(1-M_{i})) checks that the solution does not degenerate to only using the G\mathcal{G} network.

Results

For the task of segmentation, IOU (Intersection Over Union) and mIOU (mean Intersection Over Union) give the measure of overlap between the ground truth and predicted segments. Since there are two classes we calculate IOU for both the classes in the experiments. We also report the final mIOU. Additional metrics that we report to evaluate the segmentation are Precision (Prec), Recall (Rec), F1 Score, and accuracy (Acc). As surrogate for visual quality, we use FID in our experiments. This can only be seen as a rough approximation to visual quality and is mainly useful in conjunction with visual inspection of the results.

2 Datasets

We use StyleGAN2 pretrained on FFHQ , LSUN-Cars, LSUN-Cats, and LSUN-Horse datasets. FFHQ is a face image dataset with resolution 1024×10241024\times 1024. The facial features are aligned canonically. The LSUN-Object dataset contains various objects. It is relatively diverse in terms of poses, position, and number of objects in a scene. We use the subcategories for cars, cats, and horses.

3 Competing Methods

We compare our results with two approaches, a supervised approach and an unsupervised approach.

In the supervised setting, we use BiSeNet trained on the CelebA-Mask dataset for the evaluation of the faces and Facebook’s Detectron 2 Mask R-CNN Model ( R101-FPN ) with the ResNet101 architecture pre-trained on the MS-COCO dataset for the evaluation on LSUN-Object datasets. As our method is unsupervised, these methods are mainly suitable to judge how close we can come to supervised training. They are not direct competitors to our method. We also create a custom evaluation dataset of 10 images per class to directly compare the two approaches (See Supplemental (Appendix C)).

In the unsupervised setting, we compare our method with PSeg using the parameters in the open source GitHub repo. This method is the only published competitor to ours.

4 Comparison on Faces

We compare and evaluate segmentation results on both sampled and real images quantitatively and qualitatively. First we show qualitative results of the learned segmentation masks using the images and backgrounds sampled from the StyleGAN2 in Fig. 4. To put these results into context, we compare the segmentation masks of the learned alpha network A\mathcal{A} with the supervised BiSeNet. The figure shows results at different truncation levels. The truncation setting ψ=0.5\psi=0.5 leads to higher quality images. The setting ψ=1.0\psi=1.0 leads to lower quality images that are more difficult to segment from the composited representation. In such cases our method is able to outperform the supervised segmentation network. In Table 1, we compare the results of the unsupervised segmentation of StyleGAN2 generated images using our method with unsupervised Pseg. We train PSeg using the parameters in the Github repo on the FFHQ dataset. In the absence of a large corpus of testing data, we estimate the ground truth using a supervised segmentation network (BiSeNet). We sample 1k1k images and backgrounds at different truncation levels and compute the evaluation metrics. The quantitative results clearly show that our method is able to produce results of much higher quality than our only direct (unsupervised) competitor. The PSeg method tends to extract some random attributes from the learned images and is very sensitive to hyperparameters (See Fig. 7). In contrast, our method ensures that the original generative quality of StyleGAN is maintained. We also compare the FID of the sampled images in Table 5. The results show that PSeg drastically affects the sample quality.

In order to show the application for our method in real image editing, we use state-of-the-art image editing framework StyleFlow . One of the limitations of sequential editing using StyleGAN2 is that it has a tendency to change the background. Fig. 6 shows some edits on the real and sample images using StyleFlow and corresponding results where our method is able to preserve the background using the same edit. Notice that our method is robust to pose changes.

Performance on real face images.

One method to obtain segmentation results on real images is to project them into the StyleGAN2 latent space. We do this for 1k1k images from the CelebA-Mask dataset using the PSP encoder . We show quantitative results in Table 4 comparing the output segmentation masks with the ground truth in the dataset. Note that the Alpha Network A\mathcal{A} is not trained on the W+W^{+} space but is able to generalize well on the real images. A downside of this approach is that the projection method introduces its own projection errors affecting segmentation.

Labels4Free to obtain synthetic training data.

A better method to extend to arbitrary images is to generate synthetic training data (see Table 4). We train a UNet and BiSeNet on 10k10k sampled images and backgrounds produced by our unsupervised framework and report the scores. Note that the scores are comparable to the supervised setup. This demonstrates that our method can be applied to synthetically generate high quality training data for class specific foreground/background segmentation.

5 Comparison on other datasets

We also train our framework on LSUN-Cat, LSUN-Horse and LSUN-Car . These are more challenging datasets than FFHQ. We have identified two related problems with StyleGAN2 trained on these datasets. Firstly, the quality of the samples at lower truncation or no truncation levels is not as high as the FFHQ trained StyleGAN2. Second, these datasets can have multiple instances of the objects in a scene. StyleGAN2 does not handle such samples well. Both these factors affect the training as well as the evaluation of the unsupervised framework.

In Fig. 5 we show the qualitative results of our unsupervised segmentation approach on LSUN-Object datasets. Notice that the quality of segmentation masks produced by our framework is comparable to the supervised network. In Table 2, we calculated the quantitative results of 1k1k sampled images from StyleGAN2. Note that here we resort to rejection sampling based on the technique to reject the sample not identified by the detectron 2 model. For simplicity we also reject the multiple instance object samples. The results show that the metrics scores are comparable to a state-of-the-art supervised method. In order to directly compare the supervised and unsupervised approaches we compare the metrics on a custom dataset mentioned in Section 4.3. Table 3 shows that our supervised approach is either better or has similar segmentation capabilities as the supervised approach.

6 Training Details

Our method is faster to train than the competing methods. We train our framework on 4 RTX 2080 (24 GB) cards with a batch size of 8 using the StyleGAN pytorch implementation . Let λ1\lambda_{1} and λ2\lambda_{2} be the weights of the regualarizers B(M)B(M) and C(M)C(M), respectively. For FFHQ dataset, we run 1k1k iterations and set λ1=1.2\lambda_{1}=1.2. For LSUN-Object, we set ϕ=0.25\phi=0.25. For LSUN-Cat, we run 900 iterations and set λ1=λ2=3\lambda_{1}=\lambda_{2}=3 and ψ=0.5\psi=0.5. For LSUN-Horse, we run 500 iterations and set λ1=λ2=3\lambda_{1}=\lambda_{2}=3 and ψ=1.0\psi=1.0. For LSUN-Car, we run 250 iterations and set λ1=λ2=20\lambda_{1}=\lambda_{2}=20 and ψ=0.3\psi=0.3. Also, we set the weight of the non-saturating discriminator loss to 0.10.1. For FFHQ, LSUN-Horse and LSUN-Cat, we set the learning rate to 0.00020.0002. For LSUN-Car, we set the learning rate to 0.0020.002. Each model takes less than 30 minutes to converge.

Conclusion

We proposed a framework for unsupervised segmentation of StyleGAN generated images into a foreground and a background layer. The most important property of this segmentation is that it works entirely without supervision. To that effect, we leverage information that is already present in the layers to extract segmentation information. In the future, we would like to explore the unsupervised extraction of other information, e.g., illumination information, the segmentation into additional classes, and depth information.

Acknowledgements

This work was supported by Adobe and the KAUST Office of Sponsored Research (OSR) under Award No. OSR-CRG2017-3426.

References

Appendix A

In order to validate the importance of the in-domain backgrounds for the training of the unsupervised network, we train our framework with the backgrounds taken from the MIT places dataset. In order to do so, we replace the background generator with a random selection of an image from MIT places. During training, we not only train the alpha network but also the discriminator (as in the main method described in the paper). As discussed in the main paper, the discriminator is very good in identifying the out of the distribution images. In Table 6, we show the scores compared with BiSeNet and Detectron 2 (see Table 1 and Table 2 in the main paper) when using the MIT places for the backgrounds. Note that the scores decrease drastically.

As a second ablation study, we try to learn the Alpha mask only from features of the last layer before the output. This straightforward extension does not work well as seen in Table 7. In summary, using features from multiple layers in the generator is important to achieve higher fidelity.

Appendix B

In Fig. 9, 10 and 11, we visualize the tRGB layers at different resolutions as discussed in the experiment in Section 3.1 of the main paper. We select all the resolution tensors including and after the highlighted resolution. Here, we first normalize the tensors by x−min⁡(x)max⁡(x)−min⁡(x)\frac{x-\min(x)}{\max(x)-\min(x)}, where xx represents the tRGB tensor at a given resolution. Notice that for the face visualization in Fig. 9, the face structure is clearly noticeable at the 32×3232\times 32 resolution corresponding to the 4th layer of StyleGAN2. Other efforts in the StyleGAN-based local editing also selects early layers for the semantic manipulation of the images. These tests support that even features from earlier layers are beneficial for segmentation.

2 Multiple object segmentation

Apart from the main object, there can be multiple instances of the same object or multiple objects in the scene in correlation with the foreground object. Fig. 8 shows that our method is able to handle such cases.

Appendix C

We curated a custom dataset to evaluate the performance of our method with the supervised frameworks (BiSeNet and Detectron 2). We collected 10 images per class. These images were sampled by using the pretrained StyleGAN2 at different trucation levels and on different datasets i.e., FFHQ, LSUN-Horse, LSUN-Cat and LSUN-Car. Note that we collected a diverse set of images , e.g., diverse poses, lighting, and background instances (see Fig. 12). The images were annotatd using the LabelBox tool .