StyleGAN-Human: A Data-Centric Odyssey of Human Generation

Jianglin Fu, Shikai Li, Yuming Jiang, Kwan-Yee Lin, Chen Qian, Chen Change Loy, Wayne Wu, Ziwei Liu

Introduction

Generating photo-realistic images of clothed humans unconditionally can provide great support for downstream tasks such as human motion transfer , digital human animation , fashion recommendation , and virtual try-on . Traditional methods create dressed humans with classical graphics modeling and rendering processes . Although impressive results have been achieved, these prior works are easy to suffer from the limitation of robustness and generalizability in complex environments. Recent years, Generative Adversarial Networks (GANs) have demonstrated remarkable abilities in real-world scenarios, generating diverse and realistic images by learning from large-quantity and high-quality datasets. .

Among the GAN family, StyleGAN2 stands out in generating faces and simple objects with unprecedented image quality. A major driver behind recent advancements on such StyleGAN architectures is the prosperous discovery of “network engineering” like designing new components and loss functions . While these approaches show compelling results in generating diverse objects (e.g., faces of humans and animals), applying them to the photo-realistic generation of articulated humans in natural clothing is still a challenging and open problem.

In this work, we focus on the task of Unconditional Human Generation, with a specific aim to train a good StyleGAN-based model for articulated humans from a data-centric perspective. First, to support the data-centric investigation, collecting a large-scale, high-quality, and diverse dataset of human bodies in clothing is necessary. We propose the Stylish-Humans-HQ Dataset (SHHQ), which contains 230K230K clean full-body images with a resolution of 1024×5121024\times 512 at least and up to 2240×19202240\times 1920. The SHHQ dataset lays the foundation for extensive experiments on unconditional human generation. Second, based on the proposed SHHQ dataset, we investigate three fundamental and critical questions that were not thoroughly discussed in prior works and attempt to provide useful insights for future research on unconditional human generation.

To extract the questions that are indeed important for the community of Unconditional Human Generation, we make an extensive survey on recent literature in the field of general unconditional generation . Based on the survey, three questions that are investigated actively can be concluded as below. Question-1: What is the relationship between the data size and the generation quality? Several previous works pointed out that the quantity of training data is the primary factor to determine the strategy for improving image quality in face and other object generation tasks. In this study, we want to examine the minimum quantity of training data required to generate human images of high quality without any extensive “network engineering” effort. Question-2: What is the relationship between the data distribution and the generation quality? This question has received extensive attention and leads to a research topic dealing with data imbalance . In this study, we aim to exploit data imbalance problem in the human generation task. Question-3: What is the relationship between the scheme of data alignment and the generation quality? Different alignment schemes applied to uncurated faces and non-rigid objects show success in enhancing training performance. In this study, we seek a better data alignment strategy for human generation.

Based on the proposed SHHQ dataset and observations from our experiments, we establish a Model Zoo with three widely-adopted unconditional generation models, i.e., StyleGAN , StyleGAN2 , and alias-free StyleGAN , in both resolution of 1024×5121024\times 512 and 512×256512\times 256. Although hundreds of StyleGAN-based studies exist for face generation/editing tasks, a high-quality and public model zoo for human generation/editing with StyleGAN family is still missing. The provided model zoo is positioned to complement existing facial model zoo. We believe it has great potentials in many human-centric tasks, e.g., human editing, neural rendering, and virtual try-on.

We further construct a human editing benchmark by adapting previous editing methods based on facial models to human body models (i.e., PTI for image inversion, InterFaceGAN , StyleSpace , and SeFa for image manipulation). The impressive results in editing human clothes and attributes demonstrate the potential of the given model zoo in downstream tasks. In addition, a concurrent work, InsetGAN , is also evaluated with our baseline model, further showing the potential usage of our pre-trained human generative models.

Here is the summary of the main contributions of this paper: 1) We collect a large-scale, high-quality, and diverse dataset, Stylish-Humans-HQ (SHHQ), containing 230K230K human full-body images for unconditional human generation task. 2) We investigate three crucial questions that have aroused broad interest in the community and discuss our observation through comprehensive analysis. 3) We build a model zoo for unconditional human generation to facilitate future research. An editing benchmark is also established to demonstrate the potential of the proposed model zoo.

Related Work

Large-scale and high-quality clothed human-centric training datasets are the critical fuel for the training of StyleGAN models. A qualified dataset should conform to the following aspects: 1) Image quality: high-resolution images with rich textures offer more raw detailed semantic information to the model. 2) Data volume: the size of dataset should be sufficient to avoid generative overfitting . 3) Data coverage: the dataset should cover multiple attribute dimensions to guarantee diversity of the model, for instance, gender, clothing type, clothing texture, and human pose. 4) Data content: since this report only focuses on the generation of single full-body human, occlusion caused by other people or objects is not considered here, whereas self-occlusion is taken into account. That is, each image should contain only one complete human body.

Publicly available datasets built particularly for full human-body generation are rare, but there are several practices cooperating with DeepFashion and Market1501 . DeepFashion dataset with well-labeled attributes and diverse garment categories is satisfactory for image classification and attribute prediction, but not adequate for unconditional human generation since it emphasizes fashion items rather than human bodies. Thus the number of close-up shots of clothing is much higher than that of full-body images. Market1501 dataset fails for human generation tasks due to its low resolution (128×64128\times 64). There are some human-related datasets in other domains rather than GAN-based applications: datasets related to human parsing are limited by scalability and diversity; common datasets for virtual try-on tasks either contain only the upper body or are not public . A detailed comparison of the above datasets in terms of data scale, average resolution, attributes labeling, and proportion of full-body images across the whole dataset is listed in Table 1. In general, there is no high-quality and large-scale full human-body dataset publicly available for the generative purpose.

2 StyleGAN

In recent years, the research focus has gradually shifted to generating high-fidelity and high-resolution images through Generative Adversarial Networks . The StyleGAN generator was introduced and became the state-of-the-art network of unconditional image generation. Compared to previous GAN-based architectures , SytleGAN injects a separate attribute factor (i.e., style) into the generator to influence the appearance of generated images. Then StyleGAN2 redesigns the normalization, multi-scale scheme, and regularization method to rectify the artifacts in StyleGAN images. The latest update to StyleGAN reveals the non-ideal case of detailed textures sticking to fixed pixel locations and proposes an alias-free network.

3 Human Generation

In human generation research, most of the existing applications focus on precise control of pose and appearance by leveraging conditional VAE and U-Net or StyleGAN-related architectures . Specifically, the 3D method renders StyleGAN-generated neural textures on the parametric human models, but the results are restricted by the quantity and quality of training data. The other works preserve texture quality by spatial modulation using the extracted UV texture map, and perform pose transfer conditioned by extracted pose features. The limitation of these works is that paired data with satisfied volume is required for training. Moreover, studies trained with DeepFashion indicate that DeepFashion can produce decent results in human generation, at least for head-to-waist images. These works rely on additional network modifications and certain human priors. The above works can be summarized as “network engineering”; they require some architectural changes and certain priors. In contrast, this study probes unconditional human generation challenges from a data perspective.

4 Image Editing

Benefiting from StyleGAN, one of the significant downstream applications is image editing . A standard image editing pipeline usually involves inversion from a real image to the latent space and manipulating the embedded latent code. Existing works for image inversion can be categorized into optimization-based , encoder-based , and hybrid methods , which exploit encoders to embed images into latent space first and then refine with optimization. As for image manipulation, studies explore the capability of attribute disentanglement in the latent space with supervised and unsupervised networks. In specific, Jiang et al. proposes to use manually labeled fine-grained annotations to find non-linear manipulation directions in the StyleGAN latent space, while SeFa search for semantic directions without supervision. StyleSpace defines the style space SS and proves that it is more disentangled than WW and W+W+ space. In this report, we perform image editing on real images by the inversion method PTI and various directions through the chosen methods on our model to verify whether our model with human images could preserve the characteristics demonstrated on rigid objects.

Stylish-Humans-HQ Dataset

To investigate the key factors in unconditional human generation task from a data-centric perspective, we propose a large-scale, high-quality, and diverse dataset, Stylish-Humans-HQ (SHHQ). In this section, we first present the data collection and preprocessing (Section 3.1) , in which we construct the SHHQ dataset. Then, we analyze the data statistic (Section 3.2) of Stylish-Humans-HQ dataset to demonstrate the superiority of SHHQ compared to other datasets from a statistical perspective.

We first obtain over 500K500K raw data of human images from the Internet, covering a wide variety of races, ages and clothing styles. Some representative samples of raw data are shown in Appendix A. We preprocess the data with six factors taken into consideration (i.e., resolution , body position , body-part occlusion, human pose , multi-person, and background), which are critical for the quality of a human dataset. After the data preprocessing procedure, we obtain a clean dataset of 231,176231,176 images with high quality; see Figure 6 (a) for examples.

Resolution. We discard images with a resolution lower than 1024×5121024\times 512 (Figure 2 (a)).

Body Position. The position of the body varies widely in different images, as shown in Figure 2 (b). We design a procedure in which each person is appropriately cropped based on human segmentation , padded and resized to the same scale, and then placed in the image such that the body center is aligned. The body center is defined as the average coordinate of the entire body using segmentation.

Body-Part Occlusion. This work aims at generating full-body human images, images with any missing body parts are removed (e.g., the half-body portrait shown in Figure 2 (c)).

Human Pose. We remove images with extreme poses (e.g., lying postures, handstand in Figure 2 (d)) to ensure learnability of the data distribution. We exploit human pose estimation to detect those extreme poses.

Multi-Person Images. Some raw images contain multiple persons, such as in Figure 2 (e). In this work, our goal is to generate single-person full-body images, so we keep unoccluded single-person full-body images, and remove those with occluded people.

Background. Some images contain complicated backgrounds, requiring additional representation ability. To focus on the generation of the human body itself and eliminate the influence of various backgrounds, we use a segmentation mask to modify the image background to pure white. The edges of the mask are smoothed by Gaussian blur.

2 Data Statistics

Table 1 presents the comparison between SHHQ and other public datasets. Dataset Scale. As shown in the table, our proposed SHHQ is currently the largest dataset in scale compared to other datasets. Among them, the data volume of the SHHQ dataset is 1.61.6 times that of DeepFashion dataset and is much larger than that of others (3030 times to ATR , 77 times to Market1501 , 4.64.6 times to LIP , and 1414 times to VITON ). Resolution. Images from ATR , Market1501 , LIP , and VITON are lower in resolution, which is insufficient for our generation task, while the proposed SHHQ and DeepFashion provide high-definition images up to 2240×19202240\times 1920. Labels. All datasets beside VITON provide various labeled attributes. Specifically, DeepFashion and SHHQ label the clothing types and textures, which is useful for human generation/editing tasks. Full-body ratio denotes the proportion of full-body images in the dataset. Full-Body Ratio. Although DeepFashion offers over 146K146K images with decent resolution, only 6.8%6.8\% of them are full-body images, while SHHQ achieves a 100%100\% full-body ratio. Appendix A shows a visual comparison among several representative datasets and the proposed SHHQ dataset.

In summary, SHHQ covers the largest number of human images with high-resolution, labeled clothing attributes, and 100%100\% full-body ratio. It again confirms that our dataset is more suitable for full-body human generation than other public datasets.

Of all the datasets compared above, DeepFashion is the most relevant to our human generation task. In Figure 3, we further present the comparison of different attributes between filtered DeepFashion (removing occluded body) and SHHQ in a more detailed view. The bar chart depicts the distributions along six dimensions: upper cloth texture, lower cloth texture, upper cloth length, lower cloth length, gender, and ethnicity. In particular, the number of females is approximately 44 times the number of males in filtered DeepFashion , while our dataset features a more balanced female-to-male ratio of 1.491.49. With the help of DeepFace API , it is shown that SHHQ is more diverse in terms of ethnicity. Advantages are also shown in the other five attributes. In terms of garment-related attributes, images with specific labels in filtered DeepFashion are too scarce to be used as a training set. The Stylish-Humans-HQ dataset boosts the number of each category by an average of 24.424.4 times.

Systematic Investigation

Our investigations are built on the official StyleGAN2 codebase https://github.com/NVlabs/stylegan2-ada-pytorch and StyleGAN2 architecture. The detailed training settings can be found in Appendix C.

We conduct extensive experiments to study three factors concerning the quality of generated images: 1) data size (Section 4.1), 2) data distribution (Section 4.2), and 3) data alignment (Section 4.3).

Motivation. Data size is an essential factor that determines the quality of generated images. Previous literature always takes different strategies to improve the generation performance according to different dataset sizes: regularization techniques are employed to train a large dataset, while augmentation and conditional feature transferring are proposed to tackle the limited data of faces and non-rigid objects. Here, we design sets of experiments to examine the relationship between training data size and the image quality of generated humans.

Experimental Settings. To determine the relationship between data size and image quality for the unconditional human GAN, we construct 66 sub-datasets and denoted these subsets as S0S0 (10K10K), S1S1 (20K20K), S2S2 (40K40K), S3S3 (80K80K), S4S4 (160K160K) and S5S5 (230K230K). Here, S0S0 is the pruned DeepFashion dataset. We perform the training on two resolution settings for each set: 1024×5121024\times 512 and 512×256512\times 256. Considering the case of limited data, we also conduct additional training experiments with adaptive discriminator augmentation (ADA) for small datasets S0S0, S1S1, and S2S2. Fréchet Inception Distance (FID) and Inception Score (IS) are the indicators for evaluating the model performance.

Results. As shown in Figure 4 (a), the FID scores (solid lines) decrease as the size of the training dataset increases for both resolution settings. The declining trend is gradually flattening and tends to converge. S0S0 generates the least satisfactory results, with FID of 7.807.80 and 7.237.23 for low- and high-resolution, respectively, while S1S1 achieves corresponding improvements of 42%42\% and 40%40\% on FID with only an additional 10K10K training images. When the training size reaches 40K40K for both resolutions, the FID curves start to converge to a certain extent.

The dotted lines indicate the results of ADA experiments with subsets S0S0 - S2S2. The employed data augmentation strategy helps to reduce FID when training data is less than 40K40K. Table 2 and Figure 12 in Appendix B show the detailed quantitative results of FID and IS scores, from where the IS increases with the size of training data and slows down after the amount of data reaches 40K40K.

Discussion. The experiments confirm that ADA can improve the generation quality for datasets smaller than 40K40K images, in terms of FID and IS. However, ADA still cannot fully compensate for the impact of insufficient data. Besides, when the amount of data is less than 40K40K, the relationship between image quality and data size is close to linear. As the amount of data increases to 40K40K and more, the improvement in the quality of the resulting images slows down and is less significant.

2 Data Distribution

Motivation. The nature of GAN makes the model inherits the distribution of the training dataset and introduces generation bias due to dataset imbalance . This bias severely affects the performance of GAN models. To address this issue, studies for unfairness mitigation have attracted substantial research interest. In this work, we explore the question of data distribution in human generation and conduct experiments to verify whether a uniform data distribution can improve the performance of a human generation model.

Experimental Settings. This study decomposes the distribution of the human body into Face Orientation and Clothing Texture, since face fidelity has a significant impact on visual perception and clothing occupies a large portion of the full-body image. Figure 5 depicts the distribution of face orientation angle in SHHQ. The general features of human faces are relatively symmetrical; thus, we fold yaw distribution vertically along 0∘0^{\circ} and get the long-tailed distribution. For the face and clothing experiments, we collect an equal number of long-tailed and uniformly distributed datasets from SHHQ for face rotation angle and upper-body clothing texture, respectively.

Results. To evaluate the image quality in terms of different distributions, the cropped faces and clothing regions are used to calculate FID, and FID is calculated separately for each bin. Result can be found in Figure 4 (b) and (c).

1) Face Orientation: As for the long-tailed experiment (blue curve in Figure 4 (b)), the FID progressively grows as the face yaw angle increases and remains high when the facial rotation angle is too large. By contrast, the upward trend for the face FID in the uniform experiment (red) is more gradual. In addition, the amount of the training data of the first two bins in the uniform set is greatly reduced compared to the long-tail experiment, but the damage to FID is slight.

Figure 5 presents the random samples belonging to different bins in face experiments. It can be visually observed that the right-most samples in the uniform experiment have better image quality. More cropped faces in different bins from both experiments can be found in Appendix B.

2) Clothing Texture: From Figure 4 (c), except for the first bin (“plain” pattern), the FID curve climbs steadily as the amount of training data for the long-tailed experiment decreases, and the FID curve for the uniform experiment also shows a near-uniform pattern. In particular, FID of the last bin for the uniform experiment is lower than that in the long-tailed setting. We infer that the training samples for “plaid” clothing texture in the long-tailed experiment are too few to be learned by the model.

As for the “plain” bin results, the long-tailed distribution has a lower FID score in this bin. The reason may lie in that the number of plain textures in the long-tailed distribution is considerably higher than that in the uniform distribution. Also, it can be observed that the training patches in this bin are mainly textureless color blocks (see Appendix B), where such patterns may be easier to capture by models.

Discussion. Based on the above analysis, we conclude that the uniform distribution of face rotation angles can effectively reduce the FID of rare training faces while maintaining acceptable image quality for the dominant faces. However, simply balancing the distribution of texture patterns does not always reduce the corresponding FID effectively. This phenomenon raises an interesting question that can be further explored: is the relation between image quality and data distribution also entangled with other factors, e.g., image pattern and data size? Additionally, due to the nature of GAN-based structures, a GAN model memorizes the entire dataset, and usually, the discriminator tends to overfit those poorly sampled images at the tail of the distribution. Consequently, the long-tailed situation accumulated as “tail” images is barely generated. From this perspective, it also can be seen that the uniform distribution preserves the diversity of faces/textures and partially alleviates this problem.

3 Data Alignment

Motivation. Recently, researchers have drawn attention to spatial bias in generation tasks. Several works align face images with keypoints for face generation, and other studies propose different alignment schemes to preprocess non-rigid objects . In this paper, we study the relationship between the spatial deviation of the entire human body and the generated image quality.

Experimental Settings. We randomly sample a set of 50K50K images from the SHHQ dataset and align every image separately using three different alignment strategies: aligning the image based on the face center, pelvis, and the midpoint of the whole body, as shown in Figure 6.

Following are the reasons for selecting these three positions as alignment centers. 1) For the face center, we hypothesize that faces contain rich semantic information that is valuable for learning and may account for a heavy proportion in human generation. 2) For the pelvis, studies related to human pose estimation conventionally predict the body joint coordinates relative to the pelvis. Thus we employ the pelvis as the alignment anchor. 3) For the body’s midpoint, the leg-to-body ratio (the proportion of upper and lower body length) may vary among different people; therefore, we try to find the mean coordinates of the full body with the help of the segmentation mask.

Results. Human images are complex and easily affected by various extrinsic factors such as body poses and camera viewpoints. The FID scores for the face-aligned, pelvis-aligned, and mid-body-aligned experiments are 3.53.5, 2.82.8, and 2.42.4, respectively. Figure 6 further interprets this perspective as the human bodies in (b) and (c) are tilted, and the overall image quality is degraded. The example shown in Figure 6 (c) also presents the inconsistent human positions caused by different leg-to-body ratios.

Discussion. Both FID scores and visualizations suggest that the human generative models gain more stable spatial semantic information through the mid-body alignment method than face- and pelvis-centered methods. We believe this observation could benefit later studies on human generation.

4 Experimental Insights

Now the questions can be answered based on the above investigations:

For Question-1 (Data Size): A large dataset with more than 40K40K images helps to train a high-fidelity unconditional human generation model, for both 512×256512\times 256 and 1024×5121024\times 512 resolution.

For Question-2 (Data Distribution): The uniform distribution of face rotation angles helps reduce the FID of rare faces while maintaining a reasonable quality of dominant faces. But simply balancing the clothing texture distribution does not effectively improve the generation quality.

For Question-3 (Data Alignment): Aligning the human by the center of the full body presents a quality improvement over aligning the human by face or pelvis centers.

Model Zoo and Editing Benchmark

In the field of face generation, a pre-trained StyleGAN model has shown remarkable potential and success in various downstream tasks, including editing , neural rendering , and super-resolution , which have spawned a series of compelling research. Nevertheless, a publicly available pre-trained model is still lacking for the human generation task. To fill this gap, we train our baseline model on the collected 230K230K images (SHHQ) using the StyleGAN2 framework. The training takes about 55 days on 88 Tesla V100 GPUs, and the model provides the best FID of 1.571.57. As seen in Figures 1 and 13, our model has the ability to generate full-body images with diverse poses and clothing textures under satisfactory image quality. The other two models are fully trained with StyleGAN and StyleGAN3 , both in a resolution of 1024×5121024\times 512. To adapt various application scenarios, we train models with different StyleGAN architectures in a lower image resolution (512×256512\times 256) as well. In total, all the 66 models will be released for future research. We believe they can contribute to the exploration of various tasks related to human generation and continuously benefit the community.

Furthermore, the style mixing results of the baseline model show the interpretability of the corresponding latent space, which can be seen in Figure 7. As seen in the figure, source and reference images are sampled from the baseline model, and the rest images are the style-mixing results. We see that copying low layers from reference images to source images brings changes in geometry features (pose) from reference to the source. In contrast, other features such as skin colors, garment colors, and personal identities in source images are preserved. When copying middle styles, clothing type and identical appearance are copied from reference to the source. Finally, we observe that fine styles from high-resolution layers control the clothing color. More examples are displayed in Appendix E.1. These style mixing results suggest that the provided model’s geometry and appearance information are well disentangled.

2 Editing Benchmark

StyleGAN has presented remarkable editing capabilities over faces. In this section, we extend it to the full-scale human by using off-the-shelf inversion and editing methods, in which we validate the potential of our proposed model zoo. In addition, we re-implement the concurrent human generation method, InsetGAN , to further demonstrate another practical usage with the provided model zoo.

First, we leverage several SOTA StyleGAN-based facial editing techniques, such as InterFaceGAN , StyleSpace , and SeFa , with multiple editing directions: garment length for tops and bottoms, and global pose orientation. To examine the ability of editing real images with the provided model, PTI is adopted to invert images before editing.

As illustrated in Figure 8, PTI presents the ability to invert real full-body human images. For attributes manipulation, StyleSpace expresses better disentanglement compared to InterFaceGAN and SeFa , as only the attribute-targeted region has been changed. However, as for the regions to be edited, the results of InterFaceGAN are more natural and photo-realistic. It turns out that the latent space of the human body is more complicated than other domains such as faces, objects, and scenes, and more attention should be paid to disentangle human attributes. More editing results are shown in Appendix E.2.

Moreover, InsetGAN proposes a multi-GAN optimization method to fuse face and body generated from separate GAN models. We re-implement this process by iteratively optimizing the latent codes for random faces and bodies generated by the FFHQ and our baseline model, respectively. In Figure 9, we show the fused full-body images of six human postures with different male and female faces. The optimization procedure blends diverse faces and bodies in a graceful manner. Our baseline model can be adapted to different faces and generate more complicated and diverse full-body images.

In this section, we adopt a representative facial editing method on humans generated by our pre-trained model and obtain impressive results. We also show that images generated from the released human model can be further locally optimized using existing GAN models. All these works demonstrate the effectiveness and convenience of our provided model zoo and verify its potential in human-centric tasks.

Future Work

In this study, we take a preliminary step towards the exploration of the human generation/editing problem. We believe many future works can be further explored based on the SHHQ dataset and the provided model zoo. In the following, we discuss three interesting directions, i.e., Human Generation/Editing, Neural Rendering, and Multi-modal Generation.

Human Generation / Editing. Studies in unconditional human generation , human editing , virtual try-on , and motion transfer heavily rely on large datasets to train or use existing pre-trained models as the first step of transfer learning. Furthermore, editing benchmarks show that disentangled editing of the human body remains challenging for existing methods . In this context, the released model zoo could expedite such research progress. Additionally, we further analyze failure cases generated by the provided model and discuss corresponding potential efforts that could be made to human generation tasks in Appendix D.

Neural Rendering. Another future research direction is to improve 3D consistency and mitigate artifacts in full-body human generation through neural rendering . Similar to work such as EG3D , StyleNeRF , and StyleSDF , we encourage researchers to use our human models to facilitate human generation with multi-view consistency.

Multi-modal Generation. Cross-modal representation is an emerging research trend, such as CLIP and Imagebert . Hundreds of studies are made on text-driven image generation and manipulation , e.g., DALLE and AttnGAN . In the meantime, several studies show interest in probing the transfer learning benefits of large-scale pre-trained models . Most of these works focus on faces and objects, whereas research fields related to full-scale humans could be explored more, for example, text-to-human generation and text-driven human attributes manipulation, with the help of the provided full-body human models.

Conclusion

This work mainly probes how to train unconditional human-based GAN models to generate photo-realistic images from a data-centric perspective. By leveraging the 230K230K SHHQ dataset, we analyze three fundamental yet critical issues that the community cares most about: data size, data distribution, and data alignment. While experimenting with StyleGAN and large-scale data, we obtain several empirical insights. Apart from these, we create a model zoo, consisting of six human-GAN models, and the effectiveness of the model zoo is demonstrated by employing several state-of-the-art face editing methods.

Acknowledgements. We thank Hao Zhu, Zhaoyang Liu and Zhuoqian Yang for their feedback and discussions. This study is partly supported by NTU NAP, MOE AcRF Tier 1 (2021-T1-001-088), and under the RIE2020 Industry Alignment Fund Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s).

References

Appendix A SHHQ: StyleGAN-Human Datasets

The dataset we collected consists of 230K230K high-quality images of humans that vary in clothing appearance, ethnicity, and pose. Several training samples are shown in Figure 10. Please note that all these images are unprocessed. As shown in Figure 11, we also conduct qualitative comparisons with other human datasets to demonstrate the superiority of our clean, high-quality data. Besides, we display more generated human images from the baseline model trained with our SHHQ in Figure 13.

Appendix B Experiment Results

Table 2 and Figure 12 display the results of data size experiment elaborated in Section 4.1. The results align with our expectation that increment training data will improve IS scores and reduce FID scores. Figure 15 and 16 depict the comparison between cropped faces and textures generated by the long-tail and uniform experiments. Due to privacy concerns, cropped training faces are not shown.

Appendix C Training Scheme

We adopt the official NVIDIA Pytorch version of StyleGAN2-ADA as our codebase, and use the architecture of StyleGAN2. Here are several settings we use to accommodate this human generation task: (a) The input human-image has a width-to-height ratio of 1:21:2, and the input resolution in the script is changed accordingly. (b) We adopt the same eight mapping layers as the original StyleGAN . (c) There is no such a pretrained model for human images, so all the experiments are trained from scratch with the corresponding subset. (d) All other training hyper-parameters adopt the default values.

Appendix D Limitations

Compared to face generation, training an unconditional human GAN is an arduous task because the semantic features of the full-body are much more complicated than a single face. Figure 14 shows some failure cases generated by the baseline model, suggesting several directions that can be strengthened in future human generation work. Artifacts caused by entangled features of faces/hands and clothing accessories are revealed in Figures 14 (a) - (c). Case (d) exhibits three hands on a person, which indicates that the global perception of the model needs to be improved . We observe inferior hand quality in rare poses such as (e) and (f). To address this, the potential work could be augmenting training with such extreme poses, changing the data distribution, or implementing independent networks (i.e. fine-grained discriminators) to enhance local details . The face and texture quality in cases (b) and (g) could be enhanced by local refinement as well.

Appendix E Visualization of the Applications

Here we provide more examples of images generated by style-mixing on our baseline model. Figure 17, 18, and 19 represent the results of style-mixing on coarse, middle, and high resolution respectively. It shows that the latent at different scales control different high-level attributes of the clothed human, which is similar to face images.

E.2 Human Editing

Figure 20 displays the rotation of the human from the front view to the back view. The editing is done in WW space. Figure 21 demonstrates the editing results in the length of sleeves and bottoms, based on StyleSpace .