DiffiT: Diffusion Vision Transformers for Image Generation
Ali Hatamizadeh, Jiaming Song, Guilin Liu, Jan Kautz, Arash Vahdat
Introduction
Diffusion models have revolutionized the domain of generative learning, with successful frameworks in the front line such as DALLE 3 , Imagen , Stable diffusion , and eDiff-I . They have enabled generating diverse complex scenes in high fidelity which were once considered out of reach for prior models. Specifically, synthesis in diffusion models is formulated as an iterative process in which random image-shaped Gaussian noise is denoised gradually towards realistic samples . The core building block in this process is a denoising autoencoder network that takes a noisy image and predicts the denoising direction, equivalent to the score function . This network, which is shared across different time steps of the denoising process, is often a variant of U-Net that consists of convolutional residual blocks as well as self-attention layers in several resolutions of the network. Although the self-attention layers have shown to be important for capturing long-range spatial dependencies, yet there exists a lack of standard design patterns on how to incorporate them. In fact, most denoising networks often leverage self-attention layers only in their low-resolution feature maps to avoid their expensive computational complexity. Recently, several works have observed that diffusion models exhibit a unique temporal dynamic during generation. At the beginning of the denoising process, when the image contains strong Gaussian noise, the high-frequency content of the image is completely perturbed, and the denoising network primarily focuses on predicting the low-frequency content. However, towards the end of denoising, in which most of the image structure is generated, the network tends to focus on predicting high-frequency details. The time dependency of the denoising network is often implemented via simple temporal positional embeddings that are fed to different residual blocks via arithmetic operations such as spatial addition. In fact, the convolutional filters in the denoising network are not time-dependent and the time embedding only applies a channel-wise shift and scaling. Hence, such a simple mechanism may not be able to optimally capture the time dependency of the network during the entire denoising process.
In this work, we aim to address the issue of lacking fine-grained control over capturing the time-dependent component in self-attention modules for denoising diffusion models. We introduce a novel Vision Transformer-based model for image generation, called DiffiT (pronounced di-feet) which achieves state-of-the-art performance in terms of FID score of image generation on CIFAR10 and FFHQ-64 (image space) as well as ImageNet-256 and ImageNet-512 (latent space) datasets. Specifically, DiffiT proposes a new paradigm in which temporal dependency is only integrated into the self-attention layers where the key, query, and value weights are adapted per time step. This allows the denoising model to dynamically change its attention mechanism for different denoising stages. In an effort to unify the architecture design patterns, we also propose a hierarchical transformer-based architecture for latent space synthesis tasks.
The following summarizes our contributions in this work:
We introduce a novel time-dependent self-attention module that is specifically tailored to capture both short- and long-range spatial dependencies. Our proposed time-dependent self-attention dynamically adapts its behavior over sampling time steps.
We propose a novel transformer-based architecture, denoted as DiffiT, which unifies the design patterns of denoising networks.
We show that DiffiT can achieve state-of-the-art performance on a variety of datasets for both image and latent space generation tasks.
Related Work
Transformer-based models have achieved competitive performance in different generative learning models in the visual domain . A number of transformer-based architectures have emerged for GANs . TransGAN proposed to use a pure transformer-based generator and discriminator architecture for pixel-wise image generation. Gansformer introduced a bipartite transformer that encourages the similarity between latent and image features. Styleformer uses Linformers to scale the synthesis to higher resolution images. Recently, a number of efforts have leveraged Transformer-based architectures for diffusion models and achieved competitive performance. In particular, Diffusion Transformer (DiT) proposed a latent diffusion model in which the regular U-Net backbone is replaced with a Transformer. In DiT, the conditioning on input noise is done by using Adaptive LayerNorm (AdaLN) blocks. Using the DiT architecture, Masked Diffusion Transformer (MDT) introduced a masked latent modeling approach to effectively capture contextual information. In comparison to DiT, although MDT achieves faster learning speed and better FID scores on ImageNet-256 dataset , it has a more complex training pipeline. Unlike DiT and MDT, the proposed DiffiT does not use shift and scale, as in AdaLN formulation, for conditioning. Instread, DiffiT proposes a time-dependent self-attention (i.e. TMSA) to jointly learn the spatial and temporal dependencies. In addition, DiffiT proposes both image and latent space models for different image generation tasks with different resolutions with SOTA performance.
Diffusion models have driven significant advances in various domains, such as text-to-image generation , natural language processing , text-to-speech synthesis , 3D point cloud generation , time series modeling , molecular conformal generation , and machine learning security . These models synthesize samples via an iterative denoising process and thus are also known in the community as noise-conditioned score networks. Since its initial success on small-scale datasets like CIFAR-10 , diffusion models have been gaining popularity compared to other existing families of generative models. Compared with variational autoencoders , diffusion models divide the synthesis procedure into small parts that are easier to optimize, and have better coverage of the latent space ; compared with generative adversarial networks , diffusion models have better training stability and are much easier to invert . Diffusion models are also well-suited for image restoration, editing and re-synthesis tasks with minimal modifications to the existing architecture , making it well-suited for various downstream applications.
Methodology
The distributions of these random variables are the marginal distributions of forward diffusion processes (Markovian or not ) that gradually reduces the “signal-to-noise” ratio between the data and noise. As a generative model, diffusion models are trained to approximate the reverse diffusion process, that is, to transform from the initial noisy distribution (that is approximately Gaussian) to a distribution that is close to the data one.
Despite being derived from different perspectives, diffusion models can generally be written as learning the following denoising autoencoder objective
Intuitively, given a noisy sample from (generated via ), a neural network is trained to predict the amount of noise added (i.e., ). Equivalently, the neural network can also be trained to predict instead . The above objective is also known as denoising score matching , where the goal is to try to fit the data score (i.e., ) with a neural network, also known as the score network . The score network can be related to via the relationship .
Samples from the diffusion model can be simulated by the following family of stochastic differential equations that solve from to :
where is the reverse standard Wiener process, and is a function that describes the amount of stochastic noise during the sampling process. If for all , then the process becomes a probabilistic ordinary differential equation (ODE), and can be solved by ODE integrators such as denoising diffusion implicit models (DDIM ). Otherwise, solvers for stochastic differential equations (SDE) can be used, including the one for the original denoising diffusion probabilistic models (DDPM ). Typically, ODE solvers can converge to high-quality samples in fewer steps and SDE solvers are more robust to inaccurate score models .
2 DiffiT Model
At every layer, our transformer block receives , a set of tokens arranged spatially on a 2D grid in its input. It also receives , a time token representing the time step. Similar to Ho et al. , we obtain the time token by feeding positional time embeddings to a small MLP with swish activation . This time token is passed to all layers in our denoising network. We introduce our time-dependent multi-head self-attention, which captures both long-range spatial and temporal dependencies by projecting feature and time token embeddings in a shared space. Specifically, time-dependent queries , keys and values in the shared space are computed by a linear projection of spatial and time embeddings and via
where , , , , , denote spatial and temporal linear projection weights for their corresponding queries, keys, and values respectively.
We note that the operations listed in Eq. 3 to 5 are equivalent to a linear projection of each spatial token, concatenated with the time token. As a result, key, query, and value are all linear functions of both time and spatial tokens and they can adaptively modify the behavior of attention for different time steps. We define , , and which are stacked form of query, key, and values in rows of a matrix. The self-attention is then computed as follows
In which, is a scaling factor for keys , and B corresponds to a relative position bias . For computing the attention, the relative position bias allows for the encoding of information across each attention head. Note that although the relative position bias is implicitly affected by the input time embedding, directly integrating it with this component may result in sub-optimal performance as it needs to capture both spatial and temporal information. Please see Sec. 5.4 for more analysis.
The DiffiT transformer block (see Fig. 3) is a core building block of the proposed architecture and is defined as
where TMSA denotes time-dependent multi-head self-attention, as described in the above, is the time-embedding token, is a spatial token, and LN and MLP denote Layer Norm and multi-layer perceptron (MLP) respectively.
2.1 Image Space
DiffiT uses a symmetrical U-Shaped encoder-decoder architecture in which the contracting and expanding paths are connected to each other via skip connections at every resolution. Specifically, each resolution of the encoder or decoder paths consists of consecutive DiffiT blocks, containing our proposed time-dependent self-attention modules. In the beginning of each path, for both the encoder and decoder, a convolutional layer is employed to match the number of feature maps. In addition, a convolutional upsampling or downsampling layer is also used for transitioning between each resolution. We speculate that the use of these convolutional layers embeds inductive image bias that can further improve the performance. In the remainder of this section, we discuss the DiffiT Transformer block and our proposed time-dependent self-attention mechanism. We use our proposed Transformer block as the residual cells when constructing the U-shaped denoising architecture.
We define our final residual cell by combining our proposed DiffiT Transformer block with an additional convolutional layer in the form:
where GN denotes the group normalization operation and DiffiT-Transformer is defined in Eq. 7 and Eq. 8 (shown in Fig. 3). Our residual cell for image space diffusion models is a hybrid cell combining both a convolutional layer and our Transformer block.
2.2 Latent Space
Recently, latent diffusion models have been shown effective in generating high-quality large-resolution images . In Fig. 4, we show the architecture of latent DiffiT model. We first encode the images using a pre-trained variational auto-encoder network . The feature maps are then converted into non-overlapping patches and projected into a new embedding space. Similar to the DiT model , we use a vision transformer, without upsampling or downsampling layers, as the denoising network in the latent space. In addition, we also utilize a three-channel classifier-free guidance to improve the quality of generated samples. The final layer of the architecture is a simple linear layer to decode the output.
Results
We have trained the proposed DiffiT model on CIFAR-10, FFHQ-64 datasets respectively. In Table. 1, we compare the performance of our model against a variety of different generative models including other score-based diffusion models as well as GANs, and VAEs. DiffiT achieves a state-of-the-art image generation FID score of 1.95 on the CIFAR-10 dataset, outperforming state-of-the-art diffusion models such as EDM and LSGM . In comparison to two recent ViT-based diffusion models, our proposed DiffiT significantly outperforms U-ViT and GenViT models in terms of FID score in CIFAR-10 dataset. Additionally, DiffiT significantly outperforms EDM and DDPM++ models, both on VP and VE training configurations, in terms of FID score. In Fig. 5, we illustrate the generated images on FFHQ-64 dataset. Please see supplementary materials for CIFAR-10 generated images.
2 Latent Space
We have also trained the latent DiffiT model on ImageNet-512 and ImageNet-256 dataset respectively. In Table. 2, we present a comparison against other approaches using various image quality metrics. For this comparison, we select the best performance metrics from each model which may include techniques such classifier-free guidance. In ImageNet-256 dataset, the latent DiffiT model outperforms competing approaches, such as MDT-G , DiT-XL/2-G and StyleGAN-XL , in terms of FID score and sets a new SOTA FID score of 1.73. In terms of other metrics such as IS and sFID, the latent DiffiT model shows a competitive performance, hence indicating the effectiveness of the proposed time-dependant self-attention. In ImageNet-512 dataset, the latent DiffiT model significantly outperforms DiT-XL/2-G in terms of both FID and Inception Score (IS). Although StyleGAN-XL shows better performance in terms of FID and IS, GAN-based models are known to suffer from issues such as low diversity that are not captured by the FID score. These issues are reflected in sub-optimal performance of StyleGAN-XL in terms of both Precision and Recall. In addition, in Fig. 6, we show a visualization of uncurated images that are generated on ImageNet-256 and ImageNet-512 dataset. We observe that latent DiffiT model is capable of generating diverse high quality images across different classes.
Ablation
In this section, we provide additional ablation studies to provide insights into DiffiT. We address four main questions: (1) What strikes the right balance between time and feature token dimensions ? (2) How do different components of DiffiT contribute to the final generative performance, (3) What is the optimal way of introducing time dependency in our Transformer block? and (4) How does our time-dependent attention behave as a function of time?
We conduct experiments to study the effect of the size of time and feature token dimensions on the overall performance. As shown below, we observe degradation of performance when the token dimension is increased from 256 to 512. Furthermore, decreasing the time embedding dimension from 512 to 256 impacts the performance negatively.
2 Effect of Architecture Design
As presented in Table 4, we study the effect of various components of both encoder and decoder in the architecture design on the image generation performance in terms of FID score on CIFAR-10. For these experiments, the projected temporal component is adaptively scaled and simply added to the spatial component in each stage. We start from the original ViT base model with 12 layers and employ it as the encoder (config A). For the decoder, we use the Multi-Level Feature Aggregation variant of SETR (SETR-MLA) to generate images in the input resolution. Our experiments show this architecture is sub-optimal as it yields a final FID score of 5.34. We hypothesize this could be due to the isotropic architecture of ViT which does not allow learning representations at multiple scales.
We then extend the encoder ViT into 4 different multi-resolution stages with a convolutional layer in between each stage for downsampling (config B). We denote this setup as Multi-Resolution and observe that these changes and learning multi-scale feature representations in the encoder substantially improve the FID score to 4.64.
In addition, instead of SETR-MLA decoder, we construct a symmetric U-like architecture by using the same Multi-Resolution setup except for using convolutional layer between stage for upsampling (config C). These changes further improve the FID score to 3.71. Furthermore, we first add the DiffiT Transformer blocks and construct a DiffiT Encoder and observe that FID scores substantially improve to 2.27 (config D). As a result, this validates the effectiveness of the proposed TMSA in which the self-attention models both spatial and temporal dependencies. Using the DiffiT decoder further improves the FID score to 1.95 (config E), hence demonstrating the importance of DiffiT Transformer blocks for decoding.
3 Time-Dependent Self-Attention
We evaluate the effectiveness of our proposed TMSA layers in a generic denoising network. Specifically, using the DDPM++ model, we replace the original self-attention layers with TMSA layers for both VE and VP settings for image generation on the CIFAR10 dataset. Note that we did not change the original hyper-parameters for this study. As shown in Table 5 employing TMSA decreases the FID scores by 0.28 and 0.25 for VE and VP settings respectively. These results demonstrate the effectiveness of the proposed TMSA to dynamically adapt to different sampling steps and capture temporal information.
4 Impact of Self-Attention Components
In Table 6, we study different design choices for introducing time-dependency in self-attention layers. In the first baseline, we remove the temporal component from our proposed TMSA and we only add the temporal tokens to relative positional bias (config F). We observe a significant increase in the FID score to 3.97 from 1.95. In the second baseline, instead of using relative positional bias, we add temporal tokens to the MLP layer of DiffiT Transformer block (config G). We observe that the FID score slightly improves to 3.81, but it is still sub-optimal compared to our proposed TMSA (config H). Hence, this experiment validates the effectiveness of our proposed TMSA that integrates time tokens directly with spatial tokens when forming queries, keys, and values in self-attention layers.
5 Visualization of Self-Attention Maps
One of our key motivations in proposing TMSA is to allow the self-attention module to adapt its behavior dynamically for different stages of the denoising process. In Fig. 7, we demonstrate a qualitative comparison of self-attention maps. Although the attention maps without TMSA change in accordance to noise information, they lack fine-grained object details that are perfectly captured by TMSA.
6 Effect of Classifier-Free Guidance
We investigate the effect of classifier-free guidance scale on the quality of generated samples in terms of FID score. For ImageNet-256 experiment, we used the improved classifier-free guidance which uses a power-cosine schedule to increase the diversity of generated images in early sampling stages. This scheme was not used for ImageNet-512 experiment, since it did not result in any significant improvements. As shown in Fig. 5.6, the guidance scales of 4.6 and 1.49 correspond to best FID scores of 1.73 and 2.67 for ImageNet-256 and ImageNet-512 experiments, respectively. Increasing the guidance scale beyond these values result in degradation of FID score.
Conclusion
In this work, we presented Diffusion Vision Transformers (DiffiT) which is a novel transformer-based model for diffusion-based image generation. The proposed DiffiT model unifies the design pattern of denoising diffusion architectures. We proposed a novel time-dependent self-attention layer that jointly learns both spatial and temporal dependencies. Our proposed self-attention allows for capturing short and long-range information in different time steps. Analysis of time-dependent self-attention maps reveals strong localization and dynamic temporal behavior over sampling steps. We introduced the latent DiffiT for high-resolution image generation. We have evaluated the effectiveness of DiffiT using both image and latent space experiments.
References
G Ablation
We investigate if treating time embedding as a seperate token in TMSA maybe a beneficial choice. Specifically, we apply self-attention to spatial and time tokens separately to understand the impact of decoupling them. As shown in Table S.1, we observe the degradation of performance for CIFAR10, FFHQ64 datasets, in terms of FID score. Hence, the decoupling of spatial and temporal information in TMSA leads to sub-optimal performance.
G.2 Sensitivity to Time Embedding
We study the sensitivity of DiffiT model to different time embeddings representations such as Fourier and positional time embeddings. As shown in Table S.2, using a Fourier time embedding leads to degradation of performance in terms of FID score for both CIFAR10 and FFHQ-64 datasets.
G.3 Comparison to DiT and LDM
On contrary to LDM and DiT , the latent DiffiT does not rely on shift and scale, as in AdaLN , or concatenation to incorporate time embedding into the denoising networks. However, DiffiT uses a time-dependent self-attention (i.e. TMSA) to jointly learn the spatial and temporal dependencies. In addition, DiffiT proposes both image and latent space models for different image generation tasks with different resolutions with SOTA performance. Specifically, as shown in Table S.3, DiffiT significantly outperforms LDM and DiT by 31.26% and 51.94% in terms of FID score on ImageNet-256 dataset. In addition, DiffiT outperforms DiT by 13.85% on ImageNet-512 dataset. Hence, these benchmarks validate the effectiveness of the proposes architecture and TMSA design in DiffiT model as opposed to previous SOTA for both CNN and Transformer-based diffusion models.
H Architecture
We provide the details of blocks and their corresponding output sizes for both the encoder and decoder of the DiffiT model in Table S.4 and Table S.5, respectively. The presented architecture details denote models that are trained with 6464 resolution. Without loss of generality, the architecture can be extended for 3232 resolution. For FFHQ-64 dataset, the values of , , and are 4, 4, 4, and 4 respectively. For CIFAR-10 dataset, the architecture spans across three different resolution levels (i.e. 32, 16, 8), and the values of , , are 4, 4, 4 respectively. Please refer to the paper for more information regarding the architecture details.
H.2 Latent Space
In Fig S.1, we illustrate the architecture of the latent DiffiT model. Our model is comparable to DiT-XL/2-G variant which 032 uses a patch size of 2. Specifically, we use a depth of 30 layers with hidden size dimension of 1152, number of heads dimension of 16 and MLP ratio of 4. In addition, for the classifier-free guidance implementation, we only apply the guidance to the first three input channels with a scale of where is the input latent.
I Implementation Details
We strictly followed the training configurations and data augmentation strategies of the EDM model for the experiments on CIFAR10 , and FFHQ-64 datasets, all in an unconditional setting. All the experiments were trained for 200000 iterations with Adam optimizer and used PyTorch framework and 8 NVIDIA A100 GPUs. We used batch sizes of 512 and 256, learning rates of and and training images of sizes and on experiments for CIFAR10 and FFHQ-64 datasets, respectively.
We use the deterministic sampler of EDM model with 18, 40 and 40 steps for CIFAR-10 and FFHQ-64 datasets, respectively. For FFHQ-64 dataset, our DiffiT network spans across 4 different stages with 1, 2, 2, 2 blocks at each stage. We also use window-based attention with local window size of 8 at each stage. For CIFAR-10 dataset, the DiffiT network has 3 stages with 2 blocks at each stage. Similarly, we compute attentions on local windows with size 4 at each stage. Note that for all networks, the resolution is decreased by a factor of 2 in between stages. However, except for when transitioning from the first to second stage, we keep the number of channels constant in the rest of the stages to maintain both the number of parameters and latency in our network. Furthermore, we employ traditional convolutional-based downsampling and upsampling layers for transitioning into lower or higher resolutions. We achieved similar image generation performance by using bilinear interpolation for feature resizing instead of convolution. For fair comparison, in all of our experiments, we used the FID score which is computed on 50K samples and using the training set as the reference set.
I.2 Latent Space
We employ learning rates of and and batch sizes of 256 and 512 for ImageNet-256 and ImageNet-512 experiments, respectively. We also use the exponential moving average (EMA) of weights using a decay of 0.9999 for both experiments. We also use the same diffusion hyper-parameters as in the ADM model. For a fair comparison, we use the DDPM sampler with 250 steps and report FID-50K for both ImageNet-256 and ImageNet-512 experiments.
J Qualitative Results
We illustrate visualization of generated images for CIFAR-10 and FFHQ-64 datasets in Figures S.2 and S.3, respectively. In addition, in Figures S.4, S.5, S.6 and S.7, we visualize the the generated images by the latent DiffiT model for ImageNet-512 dataset. Similarly, the generated images for ImageNet-256 are shown in Figures S.8, S.9 and S.10. We observe that the proposed DiffiT model is capable of capturing fine-grained details and produce high fidelity images across these datasets.