Text2Human: Text-Driven Controllable Human Image Generation
Yuming Jiang, Shuai Yang, Haonan Qiu, Wayne Wu, Chen Change Loy, Ziwei Liu
Introduction
Recent years have witnessed the rapid progress of image generation since the emergence of Generative Adversarial Networks (GANs) (Goodfellow et al., 2014). Nowadays, we can easily generate diverse faces of high fidelity using a pretrained StyleGAN (Karras et al., 2020), which further supports several downstream tasks, such as facial attribute editing (Abdal et al., 2021; Patashnik et al., 2021; Jiang et al., 2021) and face stylization (Song et al., 2021; Pinkney and Adler, 2020; Yang et al., 2022).
Human full-body images, another type of human-related media, are more diverse, richer, and fine-grained in content. Furthermore, human image generation (Fu et al., 2022; Frühstück et al., 2022; Grigorev et al., 2021) has wide applications, including human pose transfer (Albahar et al., 2021; Sarkar et al., 2021a), virtual try-on (Lewis et al., 2021; Cui et al., 2021), and animations (Yoon et al., 2021; Chan et al., 2019; Hong et al., 2022). From the perspective of applications and interactions, apart from generating high-fidelity human images, it is even desirable to intuitively control the synthesized human images for layman users. For example, they may want to generate a person wearing a floral T-shirt and jeans without expert software knowledge. Human image generation with explicit textual controls makes it possible for users to create 2D avatars more easily.
Despite the great potential, controllable human body image generation with high fidelity and diversity is less explored due to the following challenges: 1) Compared to faces, human body images are more complex with multiple factors, including the diversity of human poses, the complicated silhouettes of clothing, and sundry textures of clothing; 2) Existing human body image generation methods (Sarkar et al., 2021b; Yildirim et al., 2019; Weng et al., 2020) fail to generate diverse styles of clothes since they tend to generate clothes with simple patterns like pure color, let alone fine-grained controls on the textures of clothes in the generated images. 3) The generation of clothes with textual controls relies on additional fine-grained annotations. However, currently, there is a lack of human image generation datasets containing fine-grained labels on clothes shapes and textures (Liu et al., 2016a, b; Cai et al., 2022). To bridge the gap, in this work, we propose the Text2Human framework for the text-driven controllable human image generation. As shown in Fig. 1, given a human pose, users can specify the clothes shapes and textures using solely natural language descriptions. Human images are then synthesized in accordance with the textual requests.
Due to the complexity of human body images, it is challenging to handle all involving factors in a single generative model. We decompose the human generation task into two stages. Stage I generates a human parsing mask with diverse clothes shapes based on the given human pose and user-specified texts describing the clothes shapes. Then Stage II enriches the human parsing mask with diverse textures of clothes based on texts describing the clothes textures.
Considering the high diversity of clothes textures, we introduce the concept of codebook, which is widely used in VQVAE-based methods (Van Den Oord et al., 2017; Esser et al., 2021a), into our framework. The codebook learns discrete neural representations of images. To adaptively characterize textures, we propose a hierarchical VQVAE with texture-aware codebook designs. Specifically, the codebooks are constructed in multiple scales. The codebook in the coarser scale contains more structural information about textures of clothes, while the codebook in finer scales includes more detailed textures. Due to the different natures of different textures, we also build codebooks separately for each texture.
In order to conditionally generate human images consistent with the texts describing the textures, we need a sampler to select appropriate texture representations (i.e., codebook indices) from the codebook, and then re-arrange them in a reasonable order in the spatial domain. In this manner, with rich texture representations stored in codebooks, the human generation task is formulated as to sample an intermediate feature map from the learned codebooks. We adopt the diffusion model based transformer (Bond-Taylor et al., 2021; Gu et al., 2022; Esser et al., 2021b) as the sampler. With the texture-aware codebook design, we incorporate mixture-of-experts (Shazeer et al., 2017) into the sampler. The sampler has multiple index prediction expert heads to predict indices for different textures.
With the hierarchical codebooks, we need to sample intermediate feature maps from the coarse level to the fine level, i.e., sampling indices for both the coarse-level and fine-level codebook is required for the image synthesis. Thanks to the implicit relationship between codebooks at different levels learned by our proposed hierarchical VQVAE, the indices of codebook at the coarse level can provide hints for the sampling of the fine level features. A similar idea is also adopted VQVAE2 (Razavi et al., 2019). However, in VQVAE2, the pixel-wise sampling by auto-regressive models is time-consuming. By comparison, we propose a feed-forward codebook index prediction network, which predicts the desired fine-level codebook indices directly from the coarse-level features. The proposed index prediction network speeds up the sampling process and ensures the generation quality.
To facilitate the controllable human generation, we construct a large-scale full-body human image dataset dubbed DeepFashion-MultiModal dataset, which contains rich clothes shape and texture annotations, human parsing masks with diverse fashion attribute classes, and human poses. Both the textual attribute annotations and human parsing masks are manually labeled. The human poses are extracted using (Güler et al., 2018). All images are collected from the high-resolution version of DeepFashion dataset. These images are further cleaned and selected to ensure they are full-body and of good quality The dataset is available at https://github.com/yumingj/DeepFashion-MultiModal..
To summarize, our main contributions are as follows: 1) We propose the Text2Human framework for the task of text-driven controllable human generation. Our proposed framework is able to generate photo-realistic human images from natural language descriptions. 2) We build a hierarchical VQVAE with the texture-aware codebook design. We propose a transformer-based sampler with the concept of mixture-of-experts. The features are routed to different expert heads according to the required attributes. The hierarchical design and mixture-of-experts sampler enable the synthesis and control of complicated textures. 3) We propose a feed-forward index prediction network to predict codebook indices of fine-level codebook based on the features sampled at the coarse level, which overcomes the limitation of the time-consuming sampling process in classical hierarchical VQVAE methods. 4) We contribute a large-scale and high-quality human image dataset with rich clothes shape and texture annotations as well as human parsing masks to facilitate the task of controllable human synthesis.
Related Work
Generative Adversarial Network (GAN) has demonstrated its powerful capabilities in generating high-fidelity images. Since (Goodfellow et al., 2014) proposed the first generative model in 2014, different variants of GAN (Brock et al., 2019; Karras et al., 2020, 2021, 2019; Chai et al., 2022) have been proposed. In addition to unconditional generation, conditional GANs (Mirza and Osindero, 2014) were proposed to generate images based on conditions like segmentation mask (Isola et al., 2017; Wang et al., 2018; Park et al., 2019) and natural language (Xu et al., 2018; Surya et al., 2020). Our proposed Text2Human is a conditional image generation framework by taking human poses and texts as inputs. In parallel to GAN, VAE (Kingma and Welling, 2013) is another paradigm for image generation. It embeds input images into a latent distribution and synthesizes images by sampling vectors from the prior distribution. Several VAE-based works (Larsen et al., 2016; Esser et al., 2018; Van Den Oord et al., 2017; Esser et al., 2021a) have been proposed to improve the visual quality of the generated images. Our proposed method shares some similarities with existing VAE-based methods but differs in the texture-aware codebook, sampler with mixture-of-experts, and feed-forward index prediction network for the hierarchical sampling.
The goal of pose transfer (Ma et al., 2017, 2018; Liu et al., 2019, 2020; Balakrishnan et al., 2018; Tao et al., 2022) is to transfer the appearance of the same person from one pose to another. (Albahar et al., 2021) proposed a pose-conditioned StyleGAN framework. The details of the source image are warped to the target pose and then are used to spatially modulate the features for synthesis. (Zhou et al., 2019) proposed a method for the text-guided pose transfer task. (Men et al., 2020) proposed ADGAN for controllable person image synthesis. The person image is synthesized by providing a pose and several example images. All of these tasks require a source person image to synthesize the target person. Recently, TryOnGAN (Lewis et al., 2021) and HumanGAN (Sarkar et al., 2021b) are proposed to support the human image generation conditioned on human pose only. TryOnGAN trained a pose conditioned StyleGAN2 network and can generate human images under the given pose condition. HumanGAN proposed a VAE-based human image generation framework. Human images are generated by sampling from the learned distribution. However, these methods do not offer fine-grained controls on human generation. Our proposed framework allows for controllable human generation by giving texts describing the desired attributes.
Text2Human
where is the attribute embedder for and fuses attribute embeddings from attribute embedders. denotes the concatenation operation.
Together with , the is then fed into the Pose-to-Parsing Module, which is composed of an encoder and a decoder . The operation at layer of is defined as follows:
where is the spatial broadcast operation so that is broadcasted to have the same spatial size with , and .
The operation of at layer can be expressed as . The final decoded feature is fed into fully convolutional layers to make the final parsing prediction. We use the cross-entropy loss to train the whole Pose-to-Parsing Module.
2. Stage II: Parsing to Human
Then the image is reconstructed using the quantized representation . The encoder, decoder and codebook are end-to-end trained through the following loss function:
where denotes the stop-gradient operation.
To sample images from learned codebooks, autoregressive models (Salimans et al., 2017; Chen et al., 2018) are employed to predict the orderings of codebook indices. Autoregressive models predict indices in a fixed unidirectional manner and the prediction of the incoming index only relies on already sampled top-left parts. In VQVAE, PixelCNN (Van den Oord et al., 2016) is adopted as the autoregressive model. In recently proposed VQGAN (Esser et al., 2021a), transformer (Vaswani et al., 2017) is adopted for its capability to capture long-term dependencies among codebook indices (In transformer, codebook indices are referred to as ‘tokens’). Recently, some works (Bond-Taylor et al., 2021; Gu et al., 2022; Esser et al., 2021b; Chang et al., 2022) proposed to use the diffusion model to replace the autoregressive model motivated by two advantages: 1) Indices are predicted based on global and bidirectional context, resulting in more coherent sampled images; 2) Indices are predicted in parallel, leading to much faster sampling speed. Specifically, in diffusion-based transformer, starting from fully-masked indices , the final prediction of indices are sampled steps by transformers. The indices at the step are sampled following the distributions:
where is the parameters of transformers. At each time step, the indices are randomly replaced with newly sampled ones.
2.2. Hierarchical VQVAE with Texture-Aware Codebook
Then the image is reconstructed using the quantized feature through the decoder : . Here we view as two consecutive parts . The spatial sizes of the inputs to and are and , respectively.
Once the top-level codebook is trained, we move to build the bottom-level codebook . The image features represented by the codes of already recover the coarse information. Therefore, just needs to learn residual information to . We introduce a residual encoder to extract fine-level feature , which is quantized into with . The image is then constructed as follows:
During the training of bottom-level codebook and , and are fixed. The network is optimized by Eq. (4) combined with the perceptual loss and discriminator loss.
To make the codes in contain richer texture information as well as keep the well-learned structure information in , the code shape is set to rather than the conventional . It is implemented by dividing into non-overlapping patches with spatial size of . Once the features are divided into patches, the quantization process is the same as Eq. (3).
Our hierarchical VQVAE shares some similarities with VQVAE2 (Razavi et al., 2019) in the hierarchical design, but differs in the following aspects: 1) Codes in our fine-level codebook have a spatial size of , while the codes in codebooks of VQVAE2 has no spatial size; 2) Our hierarchical design is motivated by representing textures at multiple scales while VQVAE2 is motivated to learn more powerful priors over the latent codes. 3) VQVAE2 trains the whole network end-to-end, which leads to poor representation ability of coarse-level features. Our stage-wise training strategy ensures meaningful representations at all levels.
Apart from multi-level codebooks, we further design a texture-aware codebook. The motivation behind the texture-awareness of the codebook lies in that the textures with different appearances at the original scale may appear to be similar at downsampled scales, leading to an ambiguity problem if we build a single coarse-level codebook for all textures. Therefore, we build different codebooks for different texture attributes separately. We will divide features extracted by the encoders according to their texture attributes at the image level and feed them into different codebooks to get the quantized features.
2.3. Sampler with Mixture-of-Experts
To incorporate texture-aware codebooks, we adapt the diffusion-based transformer into a texture-aware one as well. A straightforward idea is to train multiple samplers for different textures. However, this naive idea has two shortcomings: 1) Contextual information in the whole image is vital for the sampling of codebook indices, while training sampler for one single texture makes such information blind to the network. 2) Training multiple samplers are not ideal if we adopt the transformer as the sampler, since multiple transformers are too heavy for modern GPU devices.
Therefore, we introduce the idea of mixture-of-experts (Shazeer et al., 2017) into the diffusion-based transformer. The inputs to the mixture-of-experts sampler consist of three parts: 1) codebook index , 2) tokenized human segmentation masks , and 3) tokenized texture masks . The texture mask is obtained by filling the texture attribute labels of clothes in the corresponding regions of the segmentation mask. The multi-head attention of the transformer is computed among all of the tokens:
where , and are learnable embeddings.
The feature extracted by the multi-head attention is routed to different experts heads. The router routes the specific textures based on the texture attribute information provided by . Each expert head is in charge of the prediction of tokens for a single texture. The prediction of tokens is formulated as a classification task, where the class number is the size of the codebook. The final codebook indices are composed of outputs from all expert heads.
During training, the codebook index is the coarse-level codebook index obtained by the hierarchical VQVAE. When it comes to sampling, is initialized with masked tokens and it is iteratively filled with newly sampled ones until fully filled.
2.4. Feed-forward Codebook Index Prediction
To sample an image from the hierarchical VQVAE, multiple feature maps composed of the hierarchical codebooks need to be fed into the decoders. The traditional paradigm (Razavi et al., 2019) is to sample multiple features at different scales. However, token-wisely sampling at larger feature scales is time-consuming. Besides, when sampling at a large feature scale, long-term dependencies are hard to capture, and thus the generated images are of poor quality.
Motivated by these, we propose a feed-forward codebook index prediction network by harnessing the implicit relationship between codebooks at different levels learned by our proposed hierarchical VQVAE. Specifically, features, which are token-wisely sampled at the coarse level, are fed into the codebook index prediction network to predict the fine-level codebook indices. The codebook index prediction network is defined as:
The encoder-decoder network is adopted for the index prediction network. It should be noted that the codebook index prediction network is texture-aware as well. Shared features are extracted by the encoder and decoder, but fed into different classifier heads according to the attributes.
The use of the codebook index prediction network and the hierarchical codebooks improves the quality of generated images compared to images generated with only one level codebook. Thanks to the feed-forward index prediction network, the sampling process at larger scales under the hierarchical VQVAE design can be achieved within only one single forward pass. It speeds up the sampling process compared to the token-wisely autoregressive sampling used in (Razavi et al., 2019).
3. Text-driven Synthesis
Our framework is a text-driven one. To transform the texts requested by users into attributes, we have some predefined text descriptions for each attribute. We use the pretrained Sentence-BERT model (Reimers and Gurevych, 2019) to extract the word embeddings of our predefined texts and the text requested by users and then calculate their cosine similarities. According to the cosine similarities of word embeddings, we then classify the texts into their corresponding attributes.
4. Interactive User Interface
We present an interactive user interface for our Text2Human as shown in Fig. 1(a). Users can upload a human pose map and then type a text describing the clothing shapes. A human parsing map will be generated accordingly. Then users provide another text describing the clothing textures, and Text2Human generates the corresponding final human image. On the right side of the interface, we provide a parsing palette, which enables users to edit the human parsing. For example, as shown in Fig. 4, users can draw some holes on jeans and make the right pant leg longer using the palette to make the generated images more customized.
DeepFashion-MultiModal Dataset
Currently, most human generation methods are developed on the low-resolution version of the DeepFashion dataset and the datasets lack fine-grained annotations. Therefore, a publicly available and well-annotated high-quality human image dataset is important for the research on the human generation task. Motivated by this, we set up a large-scale high-quality human dataset with rich attribute annotations named DeepFashion-MultiModal Dataset. In a nutshell, our dataset has the following properties: 1) It contains 11,484 high-quality images at resolution. 2) For each image, we manually annotate the human parsing labels with 24 classes. 3) Each image is annotated with attributes for both clothes shapes and textures. 4) We provide densepose for each human image.
DeepFashion dataset is a large-scale clothes database that contains over 800,000 fashion images, ranging from in-shop images to unconstrained photos uploaded by customers on e-commerce websites with varying quality. Since images from the in-shop clothes retrieval benchmark are mostly of high quality with pure color background, we filter full-body images from this benchmark. There are 11,484 full-body images in total. Similar to the data alignment method used in FFHQ (Karras et al., 2019), we align the full-body images based on their poses.
1) Human Pose Representations: We extract densepose for each image using the off-the-shelf method (Güler et al., 2018). 2) Human Parsing Annotations: Human parsing serves as an effective intermedium in pose-to-photo synthesis. For each image, we provide human parsing annotations including 24 semantic labels of body components (face, hair, skin), clothes (top, outer, skirt, dress, pants, rompers) and accessories (headwear, eyeglasses, neckwear, etc.). The human parsing is manually annotated from scratch by annotators using Photoshop. 3) Clothes Shape Annotations: We manually label the clothes shape attributes for each image. The annotations include the length of upper clothes and lower clothes, the presence of fashion accessories (e.g., hat, glasses, neckwear), and the shapes of the upper clothes’ necklines. The length of upper clothes falls into four classes: sleeveless, short-sleeve, medium-sleeve, and long-sleeve. The categories for lower clothes are three-point shorts, shorts, cropped pants, and trousers. The shapes of necklines are roughly divided into V-shape, square-shape, crew neck, turtleneck, and lapel. The presence of fashion accessories has two states, i.e., presence or absence. When we annotate clothes shapes for jumpsuits (e.g., dress and rompers), the upper part and the lower part of garments are treated separately. 4) Clothes Texture Annotations: We manually label the clothes textures by two orthogonal dimensions: clothes colors and clothes fabrics. Clothes colors consist of floral, patterned, stripes, solid color, lattice, color blocks, and hybrid colors. Clothes fabrics are divided into denim, cotton, leather, furry, knitted, tulle, and other materials.
Experiments
We split the dataset into a training set and a testing set. The training set contains images and the testing set contains images. We downsample the images to resolution. The texture attribute labels are the combinations of clothes colors and fabrics annotations. The modules in the whole pipeline are trained stage by stage. All of our models are trained on one NVIDIA Tesla V100 GPU. We adopt the Adam optimizer. The learning rate is set as . For the training of Stage I (i.e., Pose to Parsing), we use the (human pose, clothes shape labels) pairs as inputs and the labeled human parsing masks as ground truths. We use the instance channel of densepose (three-channel IUV maps in original) as the human pose . Each shape attribute is represented as one-hot embeddings. We train the Stage I module for epochs. The batch size is set as . For the training of hierarchical VQVAE in Stage II, we first train the top-level codebook, , and decoder for 110 epochs, and then train the bottom-level codebook, , and for 60 epochs with top-level related parameters fixed. The batch size is set as . The sampler with mixture-of-experts in Stage II requires and . is obtained by a human parsing tokenizer, which is trained by reconstructing the human parsing maps for epochs with batch size . is obtained by directly downsampling the texture instance maps to the same size of codebook indices maps using nearest interpolation. The cross-entropy loss is employed for training. The sampler is trained for epochs with the batch size of . For the feed-forward index prediction network, we use the top-level features and bottom-level codebook indices as the input and ground-truth pairs. The feed-forward index prediction network is optimized using the cross-entropy loss. The index prediction network is trained for epochs and the batch size is set as .
2. Comparison Methods
(Wang et al., 2018) is a conditional GAN for semantic map guided image synthesis. Here, we use the human parsing map and the texture map obtained by filling texture attribute labels in the human parsing map as inputs.
(Park et al., 2019) is a conditional GAN for semantic map guided synthesis. It is adapted in a similar way to Pix2PixHD.
(Weng et al., 2020) synthesizes human images based on a human parsing map and some attributes about the clothes.
(Sarkar et al., 2021b) is a pose-conditioned VAE-based human generation method, which generates diverse human appearances by sampling from a fixed distribution (e.g., Gaussian distribution).
(Lewis et al., 2021) is a pose-conditioned StyleGAN method. The constant noise is replaced with the pose features. We train the model with the same human pose representation as our method for fair comparisons.
(Esser et al., 2021a) is a VQVAE-based method and also shows an application to conditional human image generation. For a fair comparison, we use human parsing as the input condition.
3. Evaluation Metrics
For image generation tasks, Fréchet Inception Distance (FID) is a metric evaluating the similarities between generated images and training images. A lower FID indicates a higher quality.
We use a pretrained predictor to predict the texture attributes of generated images. The prediction accuracy is reported to measure the realism of the generated texture. We also use the pretrained predictor to calculate the ratios of complicated textures (floral, stripe, lattice) to evaluate the diversity.
A user study is performed to evaluate the quality of the generated images. Users are presented with 20 groups of results. Each group has five images generated by baselines and our method. A total of 16 users are asked to 1) rank images according to photorealism (rank 5 is the best) and 2) score texture consistency with the given three attribute labels for upper clothes, lower clothes and outer clothes. The full score is 3. If the outer clothing is not required, the score for the outer clothing is 1.
4. Quantitative Comparisons
We report the quantitative results under two different settings: human image generation 1) from a human parsing, and 2) from a given human pose. Table 1 shows the comparisons with state-of-the-art conditional image generation methods. A well-annotated human parsing map and labels for clothes texture annotations are provided to synthesize the human images. As shown in Table 1, our method achieves the lowest FID, which demonstrates the fidelity and diversity of our generated human images. In addition, the best texture attribute prediction accuracy shows that our proposed Text2Human framework can accurately generate human images conditioned on provided textures. In Table 2, we show the quantitative comparisons on pose-guided human image synthesis. Since it is non-trivial to add clothes shape and texture controls for HumanGAN and TryOnGAN, under this setting, we report the ratio of complicated textures among all generated images. The highest ratio demonstrates that our methods can synthesize diverse textures for clothes. The user study results are shown in Fig. 5. Our proposed Text2Human gets the highest rank in terms of the photorealism of the generated images. As for the clothes textures, the images synthesized by our framework are more consistent with the required texture attributes. The user study results are consistent with other quantitative results.
5. Qualitative Comparisons
Figure 6 shows visual comparisons on synthesized human images given human parsing maps and clothes textures. Our proposed method can generate complicated textures with finer details and high-fidelity faces. Figure 7 shows visual comparisons with state-of-the-art pose-guided TryOnGAN (Lewis et al., 2021) and HumanGAN (Sarkar et al., 2021b). The compared baselines do not offer any controls on clothes shapes and textures, while our method can explicitly control these attributes. We also compare our proposed Text2Human with another VQVAE-based method, Taming Transformer (Esser et al., 2021a). As shown in Fig. 8, given the same human parsing map, our method can generate more plausible human images.
6. Ablation Study
Fig. 9(a) shows the improvement brought by the proposed hierarchical VQVAE on the recovery of plaid and stripe patterns. With the hierarchical design, the reconstructed images contain more high-frequency details, verifying better texture representations are learned. The hierarchical design reduces the reconstruction loss (i.e., loss + perceptual loss) from to on the whole testing set.
To evaluate the effectiveness of our texture-aware and mixture-of-experts design, we train a diffusion-based sampler with only one codebook for all textures. As shown in Fig. 9(b), the sampler without mixture-of-experts and texture-aware codebook cannot generate requested floral textures, demonstrating our design makes the sampler better conditioned on the textual inputs. We report attribute prediction accuracies on complicated textures (i.e., floral, stripe, and denim). The results are shown in Table 3. Without mixture-of-experts, the attribute prediction accuracy drops by , , and on floral, stripe, and denim textures, respectively. There are more denim textures (3449 images) than floral (325 images) and stripe (361 images) textures in the training set. It is easier for models to capture the patterns of denim textures even without the mixture-of-expert design. As a result, we can observe a smaller performance gap for denim textures, compared to those for floral and stripe textures. It indicates that the mixture-of-experts design is more effective in generating uncommon textures with fewer training samples.
To overcome the limitations of the hierarchical sampling paradigm of VQVAE2, we propose a feed-forward index prediction network to speed up the sampling speed as well as refine the textures. In terms of running time, our feed-forward network predicts fine-level codebook indices within 0.6s while VQVAE2 takes 25mins. In terms of quality, we conduct a comparative experiment with VQVAE2. For a fair comparison, we use the “ground-truth” coarse-level codebook indices obtained when reconstructing a given human image as input to predict fine-level indices by the autoregressive model of VQVAE2 or our feed-forward network. As shown in Fig. 9(c), our method reconstructs more clear and high-fidelity clothes textures than VQVAE2. We report LPIPS distance (Zhang et al., 2018) and ArcFace distance (Deng et al., 2019) between the reconstructed images and the original images in Table 4. It further verifies the effectiveness of our proposed feed-forward index prediction network in terms of reconstruction performance. Fig. 9(d) further provides a visualization of the refinement of our feed-forward network. Our network effectively refines the synthesized lattice patterns sampled from the coarse-level codebook.
7. Limitations
In this section, we discuss three common limitations of our proposed Text2Human.
1) Uncommon poses. The performance would degrade with human poses which are uncommon in the DeepFashion-MultiModal dataset. Two examples of uncommon poses are shown in Fig. 10(a). The first pose is with two legs crossed, artifacts would appear in the cross-region. The second person stands facing the side rather than the front. In this case, artifacts would appear in the face region, as the model is prone to generate faces heading the front. And thus, the generated image looks unnatural. Our framework is data-driven and can benefit from more diverse human datasets in future work. 2) Plaid textures are blurry as shown in as shown in Fig. 10(b). This is attributed to the imbalanced textures in DeepFashion. Only 162 out of 10335 training images have plaid patterns in upper clothes. This is a common problem for all baselines, and our performance is superior. In future work, the performance could be boosted by adding more data with such complicated patterns. For newly added data, the labels for clothes attributes could be provided by the attribute predictor trained on our dataset. Some techniques dealing with imbalanced data could also be employed to mitigate the problem. 3) Potential error in word embeddings. Translating text descriptions to one-hot embeddings inevitably introduces errors. For example, for the length of sleeves, we only define four classes, i.e., sleeveless, short sleeves, medium sleeves, and long sleeves. If the user wants to generate a sweater with sleeves covering the elbow but not reaching the wrist, the synthesized human parsing cannot be perfectly aligned with the text inputs as the predefined texts cannot handle sleeves with arbitrary lengths. In future work, continuous word embeddings could be employed to provide richer and more robust information.
Conclusions
In this work, we proposed the Text2Human framework for text-driven controllable human generation in two stages: pose-to-parsing and parsing-to-human. The first stage synthesizes the human parsing masks based on required clothes shapes. In the second stage, we propose a hierarchical VQVAE with texture-aware codebooks to capture the rich multi-scale representations for diverse clothes textures, and then propose a sampler with mixture-of-experts to sample desired human images conditioned on the texts describing the textures. To speed up the sampling process of hierarchical VQVAE and further refine the sampled images from the coarse level, a feed-forward codebook index prediction network is employed. Our proposed Text2Human is able to generate human images with high diversity and fidelity in clothes textures and shapes. We also contribute a large-scale dataset, named DeepFashion-MultiModal dataset, for the controllable human image generation task.