PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, Zhenguo Li

Introduction

Recently, the advancement of text-to-image (T2I) generative models, such as DALL·E 2 (OpenAI, 2023), Imagen (Saharia et al., 2022), and Stable Diffusion (Rombach et al., 2022) has started a new era of photorealistic image synthesis, profoundly impacting numerous downstream applications, such as image editing (Kim et al., 2022), video generation (Wu et al., 2022), 3D assets creation (Poole et al., 2022), etc.

However, the training of these advanced models demands immense computational resources. For instance, training SDv1.5 (Podell et al., 2023) necessitates 6K A100 GPU days, approximately costing 320,000,andtherecentlargermodel,RAPHAEL(Xueetal.,2023b),evencosts60KA100GPUdays–requiringaround320,000, and the recent larger model, RAPHAEL (Xue et al., 2023b), even costs 60K A100 GPU days – requiring around3,080,000, as detailed in Table 2. Additionally, the training contributes substantial CO2 emissions, posing environmental stress; e.g. RAPHAEL’s (Xue et al., 2023b) training results in 35 tons of CO2 emissions, equivalent to the amount one person emits over 7 years, as shown in Figure 4. Such a huge cost imposes significant barriers for both the research community and entrepreneurs in accessing those models, causing a significant hindrance to the crucial advancement of the AIGC community. Given these challenges, a pivotal question arises: Can we develop a high-quality image generator with affordable resource consumption?

In this paper, we introduce PixArt-α\alpha, which significantly reduces computational demands of training while maintaining competitive image generation quality to the current state-of-the-art image generators, as illustrated in Figure 1. To achieve this, we propose three core designs:

We decompose the intricate text-to-image generation task into three streamlined subtasks: (1) learning the pixel distribution of natural images, (2) learning text-image alignment, and (3) enhancing the aesthetic quality of images. For the first subtask, we propose initializing the T2I model with a low-cost class-condition model, significantly reducing the learning cost. For the second and third subtasks, we formulate a training paradigm consisting of pretraining and fine-tuning: pretraining on text-image pair data rich in information density, followed by fine-tuning on data with superior aesthetic quality, boosting the training efficiency.

Based on the Diffusion Transformer (DiT) (Peebles & Xie, 2023), we incorporate cross-attention modules to inject text conditions and streamline the computation-intensive class-condition branch to improve efficiency. Furthermore, we introduce a re-parameterization technique that allows the adjusted text-to-image model to load the original class-condition model’s parameters directly. Consequently, we can leverage prior knowledge learned from ImageNet (Deng et al., 2009) about natural image distribution to give a reasonable initialization for the T2I Transformer and accelerate its training.

Our investigation reveals notable shortcomings in existing text-image pair datasets, exemplified by LAION (Schuhmann et al., 2021), where textual captions often suffer from a lack of informative content (i.e., typically describing only a partial of objects in the images) and a severe long-tail effect (i.e., with a large number of nouns appearing with extremely low frequencies). These deficiencies significantly hamper the training efficiency for T2I models and lead to millions of iterations to learn stable text-image alignments. To address them, we propose an auto-labeling pipeline utilizing the state-of-the-art vision-language model (LLaVA (Liu et al., 2023)) to generate captions on the SAM (Kirillov et al., 2023). Referencing in Section 2.4, the SAM dataset is advantageous due to its rich and diverse collection of objects, making it an ideal resource for creating high-information-density text-image pairs, more suitable for text-image alignment learning.

Our effective designs result in remarkable training efficiency for our model, costing only 753 A100 GPU days and 28,400.AsdemonstratedinFigure4,ourmethodconsumeslessthan1.2528,400. As demonstrated in Figure 4, our method consumes less than 1.25% training data volume compared to SDv1.5 and costs less than 2% training time compared to RAPHAEL. Compared to RAPHAEL, our training costs are only 1%, saving approximately3,000,000 (PixArt-α\alpha’s 28,400vs.RAPHAEL’s28,400 vs. RAPHAEL’s3,080,000). Regarding generation quality, our user study experiments indicate that PixArt-α\alpha offers superior image quality and semantic alignment compared to existing SOTA T2I models (e.g., DALL·E 2 (OpenAI, 2023), Stable Diffusion (Rombach et al., 2022), etc.), and its performance on T2I-CompBench (Huang et al., 2023) also evidences our advantage in semantic control. We hope our attempts to train T2I models efficiently can offer valuable insights for the AIGC community and help more individual researchers or startups create their own high-quality T2I models at lower costs.

Method

The reasons for slow T2I training lie in two aspects: the training pipeline and the data.

The T2I generation task can be decomposed into three aspects: Capturing Pixel Dependency: Generating realistic images involves understanding intricate pixel-level dependencies within images and capturing their distribution; Alignment between Text and Image: Precise alignment learning is required for understanding how to generate images that accurately match the text description; High Aesthetic Quality: Besides faithful textual descriptions, being aesthetically pleasing is another vital attribute of generated images. Current methods entangle these three problems together and directly train from scratch using vast amount of data, resulting in inefficient training. To solve this issue, we disentangle these aspects into three stages, as will be described in Section 2.2.

Another problem, depicted in Figure 3, is with the quality of captions of the current dataset. The current text-image pairs often suffer from text-image misalignment, deficient descriptions, infrequent diverse vocabulary usage, and inclusion of low-quality data. These problems introduce difficulties in training, resulting in unnecessarily millions of iterations to achieve stable alignment between text and images. To address this challenge, we introduce an innovative auto-labeling pipeline to generate precise image captions, as will be described in Section 2.4.

2 Training strategy Decomposition

The model’s generative capabilities can be gradually optimized by partitioning the training into three stages with different data types.

Stage1: Pixel dependency learning. The current class-guided approach (Peebles & Xie, 2023) has shown exemplary performance in generating semantically coherent and reasonable pixels in individual images. Training a class-conditional image generation model (Peebles & Xie, 2023) for natural images is relatively easy and inexpensive, as explained in Appendix A.5. Additionally, we find that a suitable initialization can significantly boost training efficiency. Therefore, we boost our model from an ImageNet-pretrained model, and the architecture of our model is designed to be compatible with the pretrained weights.

Stage2: Text-image alignment learning. The primary challenge in transitioning from pretrained class-guided image generation to text-to-image generation is on how to achieve accurate alignment between significantly increased text concepts and images.

This alignment process is not only time-consuming but also inherently challenging. To efficiently facilitate this process, we construct a dataset consisting of precise text-image pairs with high concept density. The data creation pipeline will be described in Section 2.4. By employing accurate and information-rich data, our training process can efficiently handle a larger number of nouns in each iteration while encountering considerably less ambiguity compared to previous datasets. This strategic approach empowers our network to align textual descriptions with images effectively.

Stage3: High-resolution and aesthetic image generation. In the third stage, we fine-tune our model using high-quality aesthetic data for high-resolution image generation. Remarkably, we observe that the adaptation process in this stage converges significantly faster, primarily owing to the strong prior knowledge established in the preceding stages.

Decoupling the training process into different stages significantly alleviates the training difficulties and achieves highly efficient training.

3 Efficient T2I Transformer

PixArt-α\alpha adopts the Diffusion Transformer (DiT) (Peebles & Xie, 2023) as the base architecture and innovatively tailors the Transformer blocks to handle the unique challenges of T2I tasks, as depicted in Figure 4. Several dedicated designs are proposed as follows:

Cross-Attention layer. We incorporate a multi-head cross-attention layer to the DiT block. It is positioned between the self-attention layer and feed-forward layer so that the model can flexibly interact with the text embedding extracted from the language model. To facilitate the pretrained weights, we initialize the output projection layer in the cross-attention layer to zero, effectively acting as an identity mapping and preserving the input for the subsequent layers.

AdaLN-single. We find that the linear projections in the adaptive normalization layers (Perez et al., 2018) (adaLN) module of the DiT account for a substantial proportion (27%) of the parameters. Such a large number of parameters is not useful since the class condition is not employed for our T2I model. Thus, we propose adaLN-single, which only uses time embedding as input in the first block for independent control (shown on the right side of Figure 4). Specifically, in the iith block, let S(i)=[β1(i),β2(i),γ1(i),γ2(i),α1(i),α2(i)]S^{(i)}=[\beta^{(i)}_{1},\beta^{(i)}_{2},\gamma^{(i)}_{1},\gamma^{(i)}_{2},\alpha^{(i)}_{1},\alpha^{(i)}_{2}] be a tuple of all the scales and shift parameters in adaLN. In the DiT, S(i)S^{(i)} is obtained through a block-specific MLP S(i)=f(i)(c+t)S^{(i)}=f^{(i)}(c+t), where cc and tt denotes the class condition and time embedding, respectively. However, in adaLN-single, one global set of shifts and scales are computed as S‾=f(t)\overline{S}=f(t) only at the first block which is shared across all the blocks. Then, S(i)S^{(i)} is obtained as S(i)=g(S‾,E(i))S^{(i)}=g(\overline{S},E^{(i)}), where gg is a summation function, and E(i)E^{(i)} is a layer-specific trainable embedding with the same shape as S‾\overline{S}, which adaptively adjusts the scale and shift parameters in different blocks.

Re-parameterization. To utilize the aforementioned pretrained weights, all E(i)E^{(i)}’s are initialized to values that yield the same S(i)S^{(i)} as the DiT without cc for a selected tt (empirically, we use t=500t=500). This design effectively replaces the layer-specific MLPs with a global MLP and layer-specific trainable embeddings while preserving compatibility with the pretrained weights.

Experiments demonstrate that incorporating a global MLP and layer-wise embeddings for time-step information, as well as cross-attention layers for handling textual information, persists the model’s generative abilities while effectively reducing its size.

4 Dataset construction

The captions of the LAION dataset exhibit various issues, such as text-image misalignment, deficient descriptions, and infrequent vocabulary as shown in Figure 3. To generate captions with high information density, we leverage the state-of-the-art vision-language model LLaVA (Liu et al., 2023). Employing the prompt, “Describe this image and its style in a very detailed manner”, we have significantly improved the quality of captions, as shown in Figure 3.

However, it is worth noting that the LAION dataset predominantly comprises of simplistic product previews from shopping websites, which are not ideal for training text-to-image generation that seeks diversity in object combinations. Consequently, we have opted to utilize the SAM dataset (Kirillov et al., 2023), which is originally used for segmentation tasks but features imagery rich in diverse objects. By applying LLaVA to SAM, we have successfully acquired high-quality text-image pairs characterized by a high concept density, as shown in Figure 10 and Figure 11 in the Appendix.

In the third stage, we construct our training dataset by incorporating JourneyDB (Pan et al., 2023) and a 10M internal dataset to enhance the aesthetic quality of generated images beyond realistic photographs. Refer to Appendix A.5 for details.

As a result, we show the vocabulary analysis (NLTK, 2023) in Table 1, and we define the valid distinct nouns as those appearing more than 10 times in the dataset. We apply LLaVA on LAION to generate LAION-LLaVA. The LAION dataset has 2.46 M distinct nouns, but only 8.5% are valid. This valid noun proportion significantly increases from 8.5% to 13.3% with LLaVA-labeled captions. Despite LAION’s original captions containing a staggering 210K distinct nouns, its total noun number is a mere 72M. However, LAION-LLaVA contains 234M noun numbers with 85K distinct nouns, and the average number of nouns per image increases from 6.4 to 21, indicating the incompleteness of the original LAION captions. Additionally, SAM-LLaVA outperforms LAION-LLaVA with a total noun number of 328M and 30 nouns per image, demonstrating SAM contains richer objectives and superior informative density per image. Lastly, the internal data also ensures sufficient valid nouns and average information density for fine-tuning. LLaVA-labeled captions significantly increase the valid ratio and average noun count per image, improving concept density.

Experiment

This section begins by outlining the detailed training and evaluation protocols. Subsequently, we provide comprehensive comparisons across three main metrics. We then delve into the critical designs implemented in PixArt-α\alpha to achieve superior efficiency and effectiveness through ablation studies. Finally, we demonstrate the versatility of our PixArt-α\alpha through application extensions.

Training Details. We follow Imagen (Saharia et al., 2022) and DeepFloyd (DeepFloyd, 2023) to employ the T5 large language model (i.e., 4.3B Flan-T5-XXL) as the text encoder for conditional feature extraction, and use DiT-XL/2 (Peebles & Xie, 2023) as our base network architecture. Unlike previous works that extract a standard and fixed 77 text tokens, we adjust the length of extracted text tokens to 120, as the caption curated in PixArt-α\alpha is much denser to provide more fine-grained details. To capture the latent features of input images, we employ a pre-trained and frozen VAE from LDM (Rombach et al., 2022). Before feeding the images into the VAE, we resize and center-crop them to have the same size. We also employ multi-aspect augmentation introduced in SDXL (Podell et al., 2023) to enable arbitrary aspect image generation. The AdamW optimizer (Loshchilov & Hutter, 2017) is utilized with a weight decay of 0.03 and a constant 2e-5 learning rate. Our final model is trained on 64 V100 for approximately 26 days. See more details in Appendix A.5.

Evaluation Metrics. We comprehensively evaluate PixArt-α\alpha via three primary metrics, i.e., Fréchet Inception Distance (FID) (Heusel et al., 2017) on MSCOCO dataset (Lin et al., 2014), compositionality on T2I-CompBench (Huang et al., 2023), and human-preference rate on user study.

2 Performance Comparisons and Analysis

Fidelity Assessment. The FID is a metric to evaluate the quality of generated images. The comparison between our method and other methods in terms of FID and their training time is summarized in Table 2. When tested for zero-shot performance on the COCO dataset, PixArt-α\alpha achieves a FID score of 7.32. It is particularly notable as it is accomplished in merely 12% of the training time (753 vs. 6250 A100 GPU days) and merely 1.25% of the training samples (25M vs. 2B images) relative to the second most efficient method. Compared to state-of-the-art methods typically trained using substantial resources, PixArt-α\alpha remarkably consumes approximately 2% of the training resources while achieving a comparable FID performance. Although the best-performing model (RAPHEAL) exhibits a lower FID, it relies on unaffordable resources (i.e., 200×200\times more training samples, 80×80\times longer training time, and 5×5\times more network parameters than PixArt-α\alpha). We argue that FID may not be an appropriate metric for image quality evaluation, and it is more appropriate to use the evaluation of human users, as stated in Appendix A.8. We leave scaling of PixArt-α\alpha for future exploration for performance enhancement.

Alignment Assessment. Beyond the above evaluation, we also assess the alignment between the generated images and text condition using T2I-Compbench (Huang et al., 2023), a comprehensive benchmark for evaluating the compositional text-to-image generation capability. As depicted in Table 3, we evaluate several crucial aspects, including attribute binding, object relationships, and complex compositions. PixArt-α\alpha exhibited outstanding performance across nearly all (5/6) evaluation metrics. This remarkable performance is primarily attributed to the text-image alignment learning in Stage 2 training described in Section 2.2, where high-quality text-image pairs were leveraged to achieve superior alignment capabilities.

User Study. While quantitative evaluation metrics measure the overall distribution of two image sets, they may not comprehensively evaluate the visual quality of the images. Consequently, we conducted a user study to supplement our evaluation and provide a more intuitive assessment of PixArt-α\alpha’s performance. Since user study involves human evaluators and can be time-consuming, we selected the top-performing models, namely DALLE-2, SDv2, SDXL, and DeepFloyd, which are accessible through APIs and capable of generating images.

For each model, we employ a consistent set of 300 prompts from Feng et al. (2023) to generate images. These images are then distributed among 50 individuals for evaluation. Participants are asked to rank each model based on the perceptual quality of the generated images and the precision of alignments between the text prompts and the corresponding images. The results presented in Figure 5 clearly indicate that PixArt-α\alpha excels in both higher fidelity and superior alignment. For example, compared to SDv2, a current top-tier T2I model, PixArt-α\alpha exhibits a 7.2% improvement in image quality and a substantial 42.4% enhancement in alignment.

3 Ablation Study

We then conduct ablation studies on the crucial modifications discussed in Section 2.3, including structure modifications and re-parameterization design. In Figure 6, we provide visual results and perform a FID analysis. We randomly choose 8 prompts from the SAM test set for visualization and compute the zero-shot FID-5K score on the SAM dataset. Details are described below.

“w/o re-param” results are generated from the model trained from scratch without re-parameterization design. We supplemented with an additional 200K iterations to compensate for the missing iterations from the pretraining stage for a fair comparison. “adaLN” results are from the model following the DiT structure to use the sum of time and text feature as input to the MLP layer for the scale and shift parameters within each block. “adaLN-single” results are obtained from the model using Transformer blocks with the adaLN-single module in Section 2.3. In both “adaLN” and “adaLN-single”, we employ the re-parameterization design and training for 200K iterations.

As depicted in Figure 6, despite “adaLN” performing lower FID, its visual results are on par with our “adaLN-single” design. The GPU memory consumption of “adaLN” is 29GB, whereas “adaLN-single” achieves a reduction to 23GB, saving 21% in GPU memory consumption. Furthermore, considering the model parameters, the “adaLN” method consumes 833M, whereas our approach reduces to a mere 611M, resulting in an impressive 26% reduction. “adaLN-single-L (Ours)” results are generated from the model with same setting as “adaLN-single”, but training for a Longer training period of 1500K iterations. Considering memory and parameter efficiency, we incorporate the “adaLN-single-L” into our final design.

The visual results clearly indicate that, although the differences in FID scores between the “adaLN” and “adaLN-single” models are relatively small, a significant discrepancy exists in their visual outcomes. The “w/o re-param” model consistently displays distorted target images and lacks crucial details across the entire test set.

Related work

We review related works in three aspects: Denoising diffusion probabilistic models (DDPM), Latent Diffusion Model, and Diffusion Transformer. More related works can be found in Appendix A.1. DDPMs (Ho et al., 2020; Sohl-Dickstein et al., 2015) have emerged as highly successful approaches for image generation, which employs an iterative denoising process to transform Gaussian noise into an image. Latent Diffusion Model (Rombach et al., 2022) enhances the traditional DDPMs by employing score-matching on the image latent space and introducing cross-attention-based controlling. Witnessed the success of Transformer architecture on many computer vision tasks, Diffusion Transformer (DiT) (Peebles & Xie, 2023) and its variant (Bao et al., 2023; Zheng et al., 2023) further replace the Convolutional-based U-Net (Ronneberger et al., 2015) backbone with Transformers for increased scalability.

Conclusion

In this paper, we introduced PixArt-α\alpha, a Transformer-based text-to-image (T2I) diffusion model, which achieves superior image generation quality while significantly reducing training costs and CO2 emissions. Our three core designs, including the training strategy decomposition, efficient T2I Transformer and high-informative data, contribute to the success of PixArt-α\alpha. Through extensive experiments, we have demonstrated that PixArt-α\alpha achieves near-commercial application standards in image generation quality. With the above designs, PixArt-α\alpha provides new insights to the AIGC community and startups, enabling them to build their own high-quality yet low-cost T2I models. We hope that our work inspires further innovation and advancements in this field.

We would like to express our gratitude to Shuchen Xue for identifying and correcting the FID score in the paper.

Appendix A Appendix

Diffusion models (Ho et al., 2020; Sohl-Dickstein et al., 2015) and score-based generative models (Song & Ermon, 2019; Song et al., 2021) have emerged as highly successful approaches for image generation, surpassing previous generative models such as GANs (Goodfellow et al., 2014), VAEs (Kingma & Welling, 2013), and Flow (Rezende & Mohamed, 2015). Unlike traditional models that directly map from a Gaussian distribution to the data distribution, diffusion models employ an iterative denoising process to transform Gaussian noise into an image that follows the data distribution. This process can be reversely learned from an untrainable forward process, where a small amount of Gaussian noise is iteratively added to the original image.

A.1.2 Latent Diffusion Model

Latent Diffusion Model (a.k.a. Stable diffusion) (Rombach et al., 2022) is a recent advancement in diffusion models. This approach enhances the traditional diffusion model by employing score-matching on the image latent space and introducing cross-attention-based controlling. The results obtained with this approach have been impressive, particularly in tasks involving high-density image generation, such as text-to-image synthesis. This has served as a source of inspiration for numerous subsequent works aimed at improving text-to-image synthesis, including those by Saharia et al. (2022); Balaji et al. (2022); Feng et al. (2023); Xue et al. (2023b); Podell et al. (2023), and others. Additionally, Stable diffusion and its variants have been effectively combined with various low-cost fine-tuning (Hu et al., 2021; Xie et al., 2023) and customization (Zhang et al., 2023; Mou et al., 2023) technologies.

A.1.3 Diffusion Transformer

Transformer architecture (Vaswani et al., 2017) have achieved great success in language models (Radford et al., 2018; 2019), and many recent works (Dosovitskiy et al., 2020a; He et al., 2022) show it is also a promising architecture on many computer vision tasks like image classification (Touvron et al., 2021; Zhou et al., 2021; Yuan et al., 2021; Han et al., 2021), object detection (Liu et al., 2021; Wang et al., 2021; 2022; Ge et al., 2023; Carion et al., 2020), semantic segmentation (Zheng et al., 2021; Xie et al., 2021; Strudel et al., 2021) and so on (Sun et al., 2020; Li et al., 2022b; Zhao et al., 2021; Liu et al., 2022; He et al., 2022; Li et al., 2022a). The Diffusion Transformer (DiT) (Peebles & Xie, 2023) and its variant (Bao et al., 2023; Zheng et al., 2023) follow the step to further replace the Convolutional-based U-Net (Ronneberger et al., 2015) backbone with Transformers. This architectural choice brings about increased scalability compared to U-Net-based diffusion models, allowing for the straightforward expansion of its parameters. In our paper, we leverage DiT as a scalable foundational model and adapt it for text-to-image generation tasks.

A.2 PixArt-α𝛼\alpha vs. Midjourney

In Figure 7, we present the images generated using PixArt-α\alpha and the current SOTA product-level method Midjourney (Midjourney, 2023) with randomly sampled prompts online. Here, we conceal the annotations of images belonging to which method. Readers are encouraged to make assessments based on the prompts provided. The answers will be disclosed at the end of the appendix.

A.3 PixArt-α𝛼\alpha vs. Prestigious Diffusion Models

In Figure 8 and 9, we present the comparison results using a test prompt selected by RAPHAEL. The instances depicted here exhibit performance that is on par with, or even surpasses, that of existing powerful generative models.

A.4 Auto-labeling Techniques

To generate captions with high information density, we leverage state-of-the-art vision-language models LLaVA (Liu et al., 2023). Employing the prompt, “Describe this image and its style in a very detailed manner”, we have significantly improved the quality of captions. We show the prompt design and process of auto-labeling in Figure 10. More image-text pair samples on the SAM dataset are shown in Figure 11.

A.5 Additional Implementation Details

We include detailed information about all of our PixArt-α\alpha models in this section. As shown in Table 4, among the 256×\times256 phases, our model primarily focuses on the text-to-image alignment stage, with less time on fine-tuning and only 1/8 of that time spent on ImageNet pixel dependency.

For the embedding of input timesteps, we employ a 256-dimensional frequency embedding (Dhariwal & Nichol, 2021). This is followed by a two-layer MLP that features a dimensionality matching the transformer’s hidden size, coupled with SiLU activations. We adopt the DiT-XL model, which has 28 Transformer blocks in total for better performance, and the patch size of the PatchEmbed layer in ViT (Dosovitskiy et al., 2020b) is 2×\times.

Inspired by Podell et al. (2023), we incorporate the multi-scale training strategy into our pipeline. Specifically, We divide the image size into 40 buckets with different aspect ratios, each with varying aspect ratios ranging from 0.25 to 4, mirroring the method used in SDXL. During optimization, a training batch is composed using images from a single bucket, and we alternate the bucket sizes for each training step. In practice, we only apply multi-scale training in the high-aesthetics stage after pretraining the model at a fixed aspect ratio and resolution (i.e. 256px). We adopt the positional encoding trick in DiffFit (Xie et al., 2023) since the image resolution and aspect change during different training stages.

Beside the training time discussed in Table 4, data labeling and VAE training may need additional time. We treat the pre-trained VAE as a ready-made component of a model zoo, the same as pre-trained CLIP/T5-XXL text encoder, and our total training process does not include the training of VAE. However, our attempt to train a VAE resulted in an approximate training duration of 25 hours, utilizing 64 V100 GPUs on the OpenImage dataset. As for auto-labeling, we use LLAVA-7B to generate captions. LLaVA’s annotation time on the SAM dataset is approximately 24 hours with 64 V100 GPUs. To ensure a fair comparison, we have temporarily excluded the training time and data quantity of VAE training, T5 training time, and LLaVA auto-labeling time.

In this study, we incorporated three sampling algorithms, namely iDDPM (Nichol & Dhariwal, 2021), DPM-Solver (Lu et al., 2022), and SA-Solver (Xue et al., 2023a). We observe these three algorithms perform similarly in terms of semantic control, albeit with minor differences in sampling frequency and color representation. To optimize computational efficiency, we ultimately chose to employ the DPM-Solver with 20 inference steps.

A.6 Hyper-parameters analysis

In Figure 20, we illustrate the variations in the model’s metrics under different configurations across various datasets. we first investigate FID for the model and plot FID-vs-CLIP curves in Figure 20(a) for 10k text-image paed from MSCOCO. The results show a marginal enhancement over SDv1.5. In Figure 20(b) and 20(c), we demonstrate the corresponding T2ICompBench scores across a range of classifier-free guidance (cfg) (Ho & Salimans, 2022) scales. The outcomes reveal a consistent and commendable model performance under these varying scales.

A.7 More Images generated by PixArt-α𝛼\alpha

More visual results generated by PixArt-α\alpha are shown in Figure 12, 13, and 14. The samples generated by PixArt-α\alpha demonstrate outstanding quality, marked by their exceptional fidelity and precision in faithfully adhering to the given textual descriptions. As depicted in Figure 15, PixArt-α\alpha demonstrates the ability to synthesize high-resolution images up to 1024×10241024\times 1024 pixels and contains rich details, and is capable of generating images with arbitrary aspect ratios, enhancing its versatility for real-world applications. Figure 16 illustrates PixArt-α\alpha’s remarkable capacity to manipulate image styles through text prompts directly, demonstrating its versatility and creativity.

A.8 Disccusion of FID metric for evaluating image quality

During our experiments, we observed that the FID (Fréchet Inception Distance) score may not accurately reflect the visual quality of generated images. Recent studies such as SDXL (Podell et al., 2023) and Pick-a-pic (Kirstain et al., 2023) have presented evidence suggesting that the COCO zero-shot FID is negatively correlated with visual aesthetics.

Furthermore, it has been stated by Betzalel et al. (Betzalel et al., 2022) that the feature extraction network used in FID is pretrained on the ImageNet dataset, which exhibits limited overlap with the current text-to-image generation data. Consequently, FID may not be an appropriate metric for evaluating the generative performance of such models, and (Betzalel et al., 2022) recommended employing human evaluators for more suitable assessments.

Thus, we conducted a user study to validate the effectiveness of our method.

A.9 Customized Extension

In text-to-image generation, the ability to customize generated outputs to a specific style or condition is a crucial application. We extend the capabilities of PixArt-α\alpha by incorporating two commonly used customization methods: DreamBooth (Ruiz et al., 2022) and ControlNet (Zhang et al., 2023).

DreamBooth can be seamlessly applied to PixArt-α\alpha without further modifications. The process entails fine-tuning PixArt-α\alpha using a learning rate of 5e-6 for 300 steps, without the incorporation of a class-preservation loss.

As depicted in Figure 17(a), given a few images and text prompts, PixArt-α\alpha demonstrates the capacity to generate high-fidelity images. These images present natural interactions with the environment under various lighting conditions. Additionally, PixArt-α\alpha is also capable of precisely modifying the attribute of a specific object such as color, as shown in 17(b). Our appealing visual results demonstrate PixArt-α\alpha can generate images of exceptional quality and its strong capability for customized extension.

Following the general design of ControlNet (Zhang et al., 2023), we freeze each DiT Block and create a trainable copy, augmenting with two zero linear layers before and after it. The control signal cc is obtained by applying the same VAE to the control image and is shared among all blocks. For each block, we process the control signal cc by first passing it through the first zero linear layer, adding it to the layer input xx, and then feeding it into the trainable copy and the second zero linear layer. The processed control signal is then added to the output yy of the frozen block, which is obtained from input xx. We trained the ControlNet on HED (Xie & Tu, 2015) signals using a learning rate of 5e-6 for 20,000 steps.

As depicted in Figure 18, when provided with a reference image and control signals, such as edge maps, we leverage various text prompts to generate a wide range of high-fidelity and diverse images. Our results demonstrate the capacity of PixArt-α\alpha to yield personalized extensions of exceptional quality.

A.10 Discussion on Transformer vs. U-Net

The Transformer-based network’s superiority over convolutional networks has been widely established in various studies, showcasing attributes such as robustness (Zhou et al., 2022; Xie et al., 2021), effective modality fusion (Girdhar et al., 2023), and scalability (Peebles & Xie, 2023). Similarly, the findings on multi-modality fusion are consistent with our observations in this study compared to the CNN-based generator (U-Net). For instance, Table 3 illustrates that our model, PixArt-α\alpha, significantly outperforms prevalent U-Net generators in terms of compositionality. This advantage is not solely due to the high-quality alignment achieved in the second training stage but also to the multi-head attention-based fusion mechanism, which excels at modeling long dependencies. This mechanism effectively integrates compositional semantic information, guiding the generation of vision latent vectors more efficiently and producing images that closely align with the input texts. These findings underscore the unique advantages of Transformer architectures in effectively fusing multi-modal information.

A.11 Limitations & Failure cases

In Figure 19, we highlight the model’s failure cases in red text and yellow circle. Our analysis reveals the model’s weaknesses in accurately controlling the number of targets and handling specific details, such as features of human hands. Additionally, the model’s text generation capability is somewhat weak due to our data’s limited number of font and letter-related images. We aim to explore these unresolved issues in the generation field, enhancing the model’s abilities in text generation, detail control, and quantity control in the future.

A.12 Unveil the answer

In Figure 7, we present a comparison between PixArt-α\alpha and Midjourney and conceal the correspondence between images and their respective methods, inviting the readers to guess. Finally, in Figure 21, we unveil the answer to this question. It is difficult to distinguish between PixArt-α\alpha and Midjourney, which demonstrates PixArt-α\alpha’s exceptional performance.

References