Perception Prioritized Training of Diffusion Models

Jooyoung Choi, Jungbeom Lee, Chaehun Shin, Sungwon Kim, Hyunwoo Kim, Sungroh Yoon

Introduction

Diffusion models , a recent family of generative models, have achieved remarkable image generation performance. Diffusion models have been rapidly studied, as they offer several desirable properties for image synthesis, including stable training, easy model scaling, and good distribution coverage . Starting from Ho et al. , recent works have shown that the diffusion models can render high-fidelity images comparable to those generated by generative adversarial networks (GANs) , especially in class-conditional settings, by relying on additional efforts such as classifier guidance and cascaded models . However, the unconditional generation of single models still has considerable room for improvement, and performance has not been explored for various high-resolution datasets (e.g., FFHQ , MetFaces ) where other families of generative models mainly compete.

Starting from tractable noise distribution, a diffusion model generates images by progressively removing noise. To achieve this, a model learns the reverse of the predefined diffusion process, which sequentially corrupts the contents of an image with various levels of noise. A model is trained by optimizing the sum of denoising score matching losses for various noise levels , which aims to learn the recovery of clean images from corrupted images. Instead of using a simple sum of losses, Ho et al. observed that their empirically obtained weighted sum of losses was more beneficial to sample quality. Their weighted objective is the current de facto standard objective for training diffusion models . However, surprisingly, it remains unknown why this performs well or whether it is optimal for sample quality. To the best of our knowledge, the design of a better weighting scheme to achieve better sample quality has not yet been explored.

Given the success of diffusion models with the standard weighted objective, we aim to amplify this benefit by exploring a more appropriate weighting scheme for the objective function. However, designing a weighting scheme is difficult owing to two factors. First, there are thousands of noise levels; therefore, an exhaustive grid search is impossible. Second, it is not clear what information the model learns at each noise level during training, therefore hard to determine the priority of each level.

In this paper, we first investigate what a diffusion model learns at each noise level. Our key intuition is that the diffusion model learns rich visual concepts by solving pretext tasks for each level, which is to recover the image from corrupted images. At the noise level where the images are slightly corrupted, images are already available for perceptually rich content and thus, recovering images does not require prior knowledge of image contexts. For example, the model can recover noisy pixels from neighboring clean pixels. Therefore, the model learns imperceptible details, rather than high-level contexts. In contrast, when images are highly corrupted so that the contents are unrecognizable, the model learns perceptually recognizable contents to solve the given pretext task. Our observations motivate us to propose P2 (perception prioritized) weighting, which aims to prioritize solving the pretext task of more important noise levels. We assign higher weights to the loss at levels where the model learns perceptually rich contents while minimal weights to which the model learns imperceptible details.

To validate the effectiveness of the proposed P2 weighting, we first compare diffusion models trained with previous standard weighting scheme and P2 weighting on various datasets. Models trained with our objective are consistently superior to the previous standard objective by large margins. Moreover, we show that diffusion models trained with our objective achieve state-of-the-art performance on CelebA-HQ and Oxford-flowers datasets, and comparable performance on FFHQ among various types of generative models, including generative adversarial networks (GANs) . We further analyze whether P2 weighting is effective to various model configurations and sampling steps. Our main contributions are as follows:

We introduce a simple and effective weighting scheme of training objectives to encourage the model to learn rich visual concepts.

We investigate how the diffusion models learn visual concepts from each noise level.

We show consistent improvement of diffusion models across various datasets, model configurations, and sampling steps.

Background

Diffusion models transform complex data distribution pdata(x)p_{data}(x) into simple noise distribution N(0,I)\mathcal{N}(0,\mathbf{I}) and learn to recover data from noise. The diffusion process of diffusion models gradually corrupts data x0x_{0} with predefined noise scales 0<β1,β2,...,βT<10<\beta_{1},\beta_{2},...,\beta_{T}<1, indexed by time step tt. Corrupted data x1,...,xTx_{1},...,x_{T} are sampled from data x0∼pdata(x)x_{0}\sim p_{data}(x), with a diffusion process, which is defined as Gaussian transition:

Noisy data xtx_{t} can be sampled from x0x_{0} directly:

where ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,\mathbf{I}) and αt:=∏s=1t(1−βs)\alpha_{t}:=\prod_{s=1}^{t}(1-\beta_{s}). We note that data x0x_{0}, noisy data x1,...,xTx_{1},...,x_{T}, and noise ϵ\epsilon are of the same dimensionality. To ensure p(xT)∼N(0,I)p(x_{T})\sim\mathcal{N}(0,\mathbf{I}) and the reversibility of the diffusion process , one should set βt\beta_{t} to be small and αT\alpha_{T} to be near zero. To this end, Ho et al. and Dhariwal et al. use a linear noise schedule where βt\beta_{t} increases linearly from β1\beta_{1} to βT\beta_{T}. Nichol et al. use a cosine schedule where αt\alpha_{t} resembles the cosine function.

Diffusion models generate data x0x_{0} with the learned denoising process pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) which reverses the diffusion process of Eq. 1. Starting from noise xT∼N(0,I)x_{T}\sim\mathcal{N}(0,\mathbf{I}), we iteratively subtract the noise predicted by noise predictor ϵθ\epsilon_{\theta}:

where σt2\sigma_{t}^{2} is a variance of the denoising process and z∼N(0,I)z\sim\mathcal{N}(0,\mathbf{I}). Ho et al. used βt\beta_{t} as σt2\sigma_{t}^{2}.

Recent work Kingma et al. simplified the noise schedules of diffusion models in terms of signal-to-noise ratio (SNR). SNR of corrupted data xtx_{t} is a ratio of squares of mean and variance from Eq. 2, which can be written as:

and thus the variance of noisy data xtx_{t} can be written in terms of SNR: αt=1−1/(1+SNR(t))\alpha_{t}=1-1/(1+\text{SNR}(t)). We would like to note that SNR(tt) is a monotonically decreasing function.

2 Training Objectives

The diffusion model is a type of variational auto-encoder (VAE); where the encoder is defined as a fixed diffusion process rather than a learnable neural network, and the decoder is defined as a learnable denoising process that generates data. Similar to VAE, we can train diffusion models by optimizing a variational lower bound (VLB), which is a sum of denoising score matching losses : Lvlb=∑tLtL_{vlb}=\sum_{t}L_{t}, where weights for each loss term are uniform. For each step tt, denoising score matching loss LtL_{t} is a distance between two Gaussian distributions, which can be rewritten in terms of noise predictor ϵθ\epsilon_{\theta} as:

Intuitively, we train a neural network ϵθ\epsilon_{\theta} to predict the noise ϵ\epsilon added in noisy image xtx_{t} for given time step tt.

Ho et al. empirically observed that the following simplified objective is more beneficial to sample quality:

In terms of VLB, their objective is Lsimple=∑tλtLtL_{simple}=\sum_{t}\lambda_{t}L_{t} with weighting scheme λt=(1−βt)(1−αt)/βt\lambda_{t}=(1-\beta_{t})(1-\alpha_{t})/\beta_{t}. In a continuous-time setting, this scheme can be expressed in terms of SNR:

where SNR′(t)=dSNR(t)dt\text{SNR}^{\prime}(t)=\frac{d\text{SNR}(t)}{dt}. See appendix for derivations.

While Ho et al. use fixed values for the variance σt\sigma_{t}, Nichol et al. propose to learn it with hybrid objective Lhybird=Lsimple+cLvlbL_{hybird}=L_{simple}+cL_{vlb}, where c=1e−3c=1e^{-3}. They observed that learning σt\sigma_{t} enables reducing sampling steps while maintaining the generation performance. We inherit their hybrid objective for efficient sampling and modify LsimpleL_{simple} to improve performance.

3 Evaluation Metrics

We use FID and KID for quantitative evaluations. FID is well-known to be analogous to human perception and well-used as a default metric for measuring generation performances. KID is a well-used metric to measure performance on small datasets . However, since both metrics are sensitive to the preprocessing , we use a correctly implemented library . We compute FID and KID between the generated samples and the entire training set. We measured final scores with 50k samples and conducted ablation studies with 10k samples for efficiency, following . We denote them as FID-50k and FID-10k respectively.

Method

We first investigate what the model learns at each diffusion step in Sec. 3.1. Then, we propose our weighting scheme in Sec. 3.2. We provide discussions on how our weighting scheme is effective in Sec. 3.3.

Diffusion models learn visual concepts by solving pretext task at each noise level, which is to recover signals from corrupted signals. More specifically, the model predicts the noise component ϵ\epsilon of a noisy image xtx_{t}, where the time step tt is an index of the noise level. While the output of diffusion models is noise, other generative models (VAE, GAN) directly output images. Because noise does not contain any content or signals, it is difficult to understand how the noise predictions contribute to learning rich visual concepts. Such nature of diffusion models arises the following question: what information does the model learn at each step during training?

Investigating diffusion process. We first investigate the predefined diffusion process to explore what the model can learn from each noise level. Let say we have two different clean images x0x_{0}, x0′x^{\prime}_{0} and three noisy images xtA, xtB∼q(xt∣x0)x_{tA},~{}x_{tB}\sim q(x_{t}|x_{0}), xt′∼q(xt∣x0′)x^{\prime}_{t}\sim q(x_{t}|x^{\prime}_{0}), where qq is the diffusion process. In Fig. 1 (left), we measure perceptual distances (LPIPS ) in two cases: the distance between xtAx_{tA} and xtBx_{tB} (blue line), which share the same x0x_{0}, and the distance between xtAx_{tA} and xt′x^{\prime}_{t} (orange line), which were synthesized from different images x0x_{0} and x0′x^{\prime}_{0}. We present the distances of the two cases as functions of the signal-to-noise ratio (SNR) introduced in Eq. 4, which characterizes the noise level at each step. To briefly review, SNR decreases through the diffusion process, as shown in Fig. 1 (right), and increases through the denoising process.

The early steps of the diffusion process have large SNRs, which indicates invisibly small noise; thus, noisy images xtx_{t} retain a large amount of contents from the clean image x0x_{0}. Therefore, in the early steps, xtAx_{tA} and xtBx_{tB} are perceptually similar, while xtAx_{tA} and xt′x^{\prime}_{t} are perceptually different, as shown by the large SNR side in Fig. 1 (left). A model can recover signals without understanding holistic contexts, as perceptually rich signals are already prepared in the image. Thus the model will learn only imperceptible details by solving recovery tasks when SNR is large.

In contrast, the late steps have small SNRs, indicating a sufficiently large noise to remove the contents of x0x_{0}. Therefore, distances of both cases start to converge to a constant value, as the noisy images become difficult to recognize the high-level contents. It is shown in the small SNR side in Fig. 1 (left). Here, a model needs prior knowledge to recover signals because the noisy images lack recognizable content. We argue that the model will learn perceptually rich contents by solving recovery tasks when SNR is small.

Investigating a trained model. We would like to verify the aforementioned discussions with a trained model. Given an input image x0x_{0}, we first perturb it to xtx_{t} using a diffusion process q(xt∣x0)q(x_{t}|x_{0}) and reconstruct it with the learned denoising process pθ(x^0∣xt)p_{\theta}(\hat{x}_{0}|x_{t}), as illustrated in Fig. 2 (left). When tt is small, the reconstruction x^0\hat{x}_{0} will be highly similar to the input x0x_{0} as the diffusion process removes a small amount of signals, while x^0\hat{x}_{0} will share less content with x0x_{0} when tt is large. In Fig. 2 (right), we compare x0x_{0} and x^0\hat{x}_{0} among various tt to show how each step contributes to the sample. Samples in the first two columns share only coarse features (e.g., global color scheme) with the input on the rightmost column, whereas samples in the third and fourth columns share perceptually discriminative contents. This suggests that the model learns coarse features when the SNR of step tt is smaller than 10−210^{-2} and the model learns the content when SNR is between 10−210^{-2} and 10010^{0}. When the SNR is larger than 10010^{0} (fifth column), reconstructions are perceptually identical to the inputs, suggesting that the model learns imperceptible details that do not contribute to perceptually recognizable contents.

Based on the above observations, we hypothesize that diffusion models learn coarse features (e.g., global color structure) at steps of small SNRs (0–10−20\text{--}10^{-2}), perceptually rich contents at medium SNRs (10−2–10010^{-2}\text{--}10^{0}), and remove remaining noise at large SNRs (100–10410^{0}\text{--}10^{4}). According to our hypothesis, we group noise levels into three stages, which we term coarse, content, and clean-up stages.

2 Perception Prioritized Weighting

In the previous section, we explored what the diffusion model learns from each step in terms of SNR. We discussed that the model learns coarse features (e.g., global color structure), perceptually rich contents, and to clean up the remaining noise at three groups of noise levels. We pointed out that the model learns imperceptible details at the clean-up stage. In this section, we introduce Perception Prioritized (P2) weighting, a new weighting scheme for the training objective, which aims to prioritize learning from more important noise levels.

We opt to assign minimal weights to the unnecessary clean-up stage thereby assigning relatively higher weights to the rest. In particular, we aim to emphasize training on the content stage to encourage the model to learn perceptually rich contexts. To this end, we construct the following weighting scheme:

where λt\lambda_{t} is the previous standard weighting scheme (Eq. 7) and γ\gamma is a hyperparameter that controls the strength of down-weighting focus on learning imperceptible details. kk is a hyperparameter that prevents exploding weights for extremely small SNRs and determines sharpness of the weighting scheme. While multiple designs are possible, we show that even the simplest choice (P2) outperforms the standard scheme λt\lambda_{t}. Our method is applicable to existing diffusion models by replacing ∑tλtLt\sum_{t}\lambda_{t}L_{t} with ∑tλt′Lt\sum_{t}\lambda^{\prime}_{t}L_{t}.

In fact, our weighting scheme λt′\lambda^{\prime}_{t} is a generalization of the popularly used weighting scheme λt\lambda_{t} of Ho et al. (Eq. 7), where λt′\lambda^{\prime}_{t} arrives at λt\lambda_{t} when γ=0\gamma=0. We refer to λt\lambda_{t} as the baseline herein.

3 Effectiveness of P2 Weighting

Prior works empirically suggest that the baseline objective ∑tλtLt\sum_{t}\lambda_{t}L_{t} offers a better inductive bias for sample quality than the VLB objective ∑tLt\sum_{t}L_{t}, which does not impose any inductive bias during training. Fig. 3 exhibits λt′\lambda^{\prime}_{t} and λt\lambda_{t} for both linear and cosine noise schedules, which are explained in Sec. 2.1, indicating that both weighting schemes focus training on the content stage the most and the cleaning stage the least. The success of the baseline weighting is in line with our previous hypothesis that models learn perceptually rich content by solving pretext tasks at the content stage.

However, despite the success of the baseline objective, we argue that the baseline objective still imposes an undeserved focus on learning imperceptible details and prevents from learning perceptually rich content. Fig. 3 shows that our λt′\lambda^{\prime}_{t} further suppresses the weights for the cleaning stage, which relatively uplifts the weights for the coarse and the content stages. To visualize relative changes of weights, we exhibit normalized weighting schemes. Fig. 4 supports our method in that FID of the diffusion model trained with our weighting scheme (γ=1\gamma=1) beats the baseline for both linear and cosine schedules throughout the training.

Another notable result from Fig. 4 is that the cosine schedule is inferior to the linear schedule by a large margin, although our weighting scheme improves the FID by a large gap. Sec. 2.2 indicates that the weighting scheme is closely related to the noise schedule. As shown in Fig. 3, the cosine schedule assigns smaller weights to the content stage compared with the linear schedule. We would like to note that designing weighting schemes and noise schedules are correlated but not equivalent, as the noise schedules affects both weights and MSE terms.

To summarize, our P2 weighting provides a good inductive bias for learning rich visual concepts, by uplifting weights at the coarse and the content stages, and suppressing weights at the clean-up stage.

4 Implementation

We set kk as 1 for easy deployment, because 1/(1+SNR(t))=1−αt1/(1+\text{SNR}(t))=1-\alpha_{t}, as discussed in Sec. 2.1. We set γ\gamma as either 0.5 and 1. We empirically observed that γ\gamma over 2 suffers noise artifacts in the sample because it assigns almost zero weight to the clean-up stage. We set T=1000T=1000 for all experiments. We implemented the proposed approach on top of ADM , which offers well-designed architecture and efficient sampling. We use lighter version of ADM through our experiments. Our code and models are availablehttps://github.com/jychoi118/P2-weighting.

Experiment

We start by exhibiting the effectiveness of our new training objective over the baseline objective in Sec. 4.1. Then, we compare with prior literature of various types of generative models in Sec. 4.2. Finally, we conduct analysis studies to further support our method in Sec. 4.3. Samples generated with our models are shown in Fig. 5.

Quantitative comparison. We trained diffusion models by optimizing training objectives with both baseline and our weighting scheme on FFHQ , AFHQ-dog , MetFaces , and CUB datasets. These datasets contain approximately 70k, 50k, 1k, and 12k images respectively. We resized and center-cropped data to 256×\times256 pixels, following the pre-processing performed by ADM .

Tab. 1 shows the results. Our method consistently exhibits superior performance to the baseline in terms of FID and KID. The results suggest that our weighting scheme imposes a good inductive bias for training diffusion models, regardless of the dataset. Our method outperforms the baseline by a large margin especially on MetFaces, which contains only 1k images. Hence, we assume that wasting the model capacity on learning imperceptible details is very harmful when training with limited data.

Qualitative comparison. We observe that diffusion models trained with the baseline objective are likely to suffer color shift artifacts, as shown in Fig. 6. We assume that the baseline training objective unnecessarily focuses on the imperceptible details; therefore, it fails to learn global color schemes properly. In contrast, our objective encourages the model to learn global and holistic concepts in a dataset.

2 Comparison to the Prior Literature

We compare diffusion models trained with our method to existing models on FFHQ , Oxford flowers , and CelebA-HQ datasets, as shown in Tab. 2. We use 256×\times256 resolutions for all datasets. We achieve state-of-the-art FIDs on the Oxford Flowers and CelebA-HQ datasets. While our models are trained with T=1000T=1000, we already achieve state-of-the-art with reduced sampling steps; 250 and 500 steps respectively. On FFHQ, we achieve a superior result to most models except StyleGAN2 , whose architecture was carefully designed for FFHQ. We note that our method brought the diffusion model closer to the state-of-the-art, and scaling model architectures and sampling steps will further improve the performance.

3 Analysis

In this section, we analyze whether our weighting scheme is robust to the model configurations, number of sampling steps, and sampling schedules.

Model configuration matters? Previous experiments are conducted using our default model for fair comparisons. Here, we show that P2 weighting is effective regardless of the model configurations. Tab. 3 shows that our method achieves consistently superior performance to the baseline for various configurations. We investigated for following variations: replacing the BigGAN residual block with the residual block of Ho et al. , removing self-attention at 16×\times16, using two BigGAN residual blocks, and training our default model with a learning rate of 2.5×10−52.5\times 10^{-5}. Our default model contains a single BigGAN residual block and is trained with a learning rate 2×10−52\times 10^{-5}. Our weighting scheme consistently improves FID and KID by a large margin, across various model configurations. Our method is especially effective when the self-attention is removed ((c)), indicating that P2 encourages learning global dependency.

Sampling step matters? We trained our models on 1000 diffusion steps following the convention of previous studies. However, it requires more than 10 min to generate a high-resolution image with a modern GPU. Nichol et al. have shown that their sampling strategy maintains performance even when reducing the sampling steps. They also observed that using the DDIM sampler was effective when using 50 or fewer sampling steps.

Fig. 7 shows the FID scores of various sampling steps with models trained on the FFHQ. A model trained with our weighting scheme consistently outperforms the baseline by considerable margins. It should be noted that our weighting scheme consistently achieves better performance with half the number of sampling steps required by the baseline.

Why not schedule sampling steps? In addition to the consistent improvement across various sampling steps, we sweep over the sampling steps in Tab. 4. Sweeping sampling schedules slightly improves FID and KID but does not reach our improvement. Our method is more effective compared to scheduling sampling steps, as we improve the model training, which benefits predictions at all steps.

Related Work

Diffusion models and score-based models are two recent families of generative models that generate data using a learned denoising process. Song et al. showed that both families can be expressed with stochastic differentiable equations (SDE) with different noise schedules. We note that score-based models may enjoy different hyperparamters of P2 (γ\gamma and kk) as noise schedules are correlated with weighting schemes (Sec. 2.2). Recent studies have achieved remarkable improvements in sample quality. However, they rely on heavy architectures, long training and sampling steps , classifier guidance , and a cascade of multiple models . In contrast, we improved the performance by simply redesigning the training objective without requiring heavy computations and additional models. Along with the success in the image domain, diffusion models have also shown effectiveness in speech synthesis .

2 Advantages of Diffusion Models

Diffusion models have several advantages over other generative models. First, their sample quality is superior to likelihood-based methods such as autoregressive models , flow models , and variational autoencoders (VAEs) . Second, because of the stable training, scaling and applying diffusion models to new domains and datasets is much easier than generative adversarial networks (GANs) , which rely on unstable adversarial training.

Moreover, pre-trained diffusion models are surprisingly easy to apply to downstream image synthesis tasks. Recent works have demonstrated that pre-trained diffusion models can easily adapt to image translation and image editing. Compared to GAN-based methods , they adapt a single diffusion model to various tasks without task-specific training and loss functions. They also show that diffusion models allow stochastic (one-to-many) generation in those tasks, while GAN-based methods suffer deterministic (one-to-one) generations .

3 Redesigning Training Objectives

Recent works introduced new training objectives to achieve state-of-the-art likelihood. However, their objectives suffer from degradation of sample quality and training instability, therefore rely on importance sampling or sophisticated parameterization . Because the likelihood focuses on fine-scale details, their objectives impede understanding global consistency and high-level concepts of images. For this reason, use different weighting schemes for likelihood training and FID training. Our P2 weighting provides a good inductive bias for perceptually rich contents, allowing the model to achieve improved sample quality with stable training.

Conclusion

We proposed perception prioritized weighting, a new weighting scheme for the training objective of the diffusion models. We investigated how the model learns visual concepts at each noise level during training, and divided diffusion steps into three groups. We showed that even the simplest choice (P2) improves diffusion models across datasets, model configurations, and sampling steps. Designing a more sophisticated weighting scheme may further improve the performance, which we leave as future work. We believe that our method will open new opportunities to boost the performance of diffusion models.

Acknowledgements: This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) [NO.2021-0-01343, Artificial Intelligence Graduate School Program (Seoul National University)], LG AI Research, Samsung SDS, AIRS Company in Hyundai Motor and Kia through HMC/KIA-SNU AI Consortium Fund, and the BK21 FOUR program of the Education and Research Program for Future ICT Pioneers, Seoul National University in 2022.

References

A Weighting Schemes

In the main text, we showed weights of both our new weighting scheme and the baseline, as functions of signal-to-noise ratio (SNR). In Fig. A (left), we show weights as functions of time steps (tt). To exhibit relative changes of weights, we show normalized weights as functions of both time steps (Fig. A (middle)) and SNR (Fig. A (right)). We normalized so that the sum of weights for all time steps become 1. Normalized weights suggest that larger γ\gamma suppresses weights at steps near t=0t=0 and uplifts weights at larger steps. Note that weights of VLB objective are equal to a constant, as such objective does not impose any inductive bias for training. In contrast, as discussed in the main text, our method encourages the model to learn rich content rather than imperceptible details.

A.2 Derivations

In the main text, we wrote the baseline weighting scheme λt\lambda_{t} as a funtion of SNR, which characterizes the noise level at each step tt. Below is the derivation:

which is a differential of log-SNR(tt) regarding time-step tt.

B Discussions

Despite the promising performances achieved by our method, diffusion models still need multiple sampling steps. Diffusion models require at least 25 feed-forwards with DDIM sampler, which makes it difficult to use diffusion models in real-time applications. Yet, they are faster than autoregressive models which generate a pixel at each step. In addition, we have observed in section 4.3 that our method enables better FID with half the number of steps required by the baseline. Along with our method, optimizing sampling schedules with dynamic programming or distilling DDIM sampling into a single step model might be promising future directions for faster sampling.

B.2 Broader Impacts

The proposed method in this work allows high-fidelity image generation with diffusion-based generative models. Improving the performance of generative models can enable multiple creative applications . However, such improvements have the potential to be exploited for deception. Works in deepfake detection or watermarking can alleviate the problems. Investigating invisible frequency artifacts in samples of diffusion models might be promising approach to detect fake images.

C Implementation Details

For a given time-step tt, the input noisy image xtx_{t} and output noise prediction ϵ\epsilon and variance σt\sigma_{t} are images of the same resolution. Therefore, ϵθ\epsilon_{\theta} is parameterized with the U-Net -style architecture of three input and six output channel dimensions. We inherit the architecture of ADM , which is a U-Net with large channel dimension, BigGAN residual blocks, multi-resolution attention, and multi-head attention with fixed channels per head. Time-step tt is provided to the model by adaptive group normalization (AdaGN), which transforms tt embeddings to scales and biases of group normalizations . However, for efficiency, we use fewer base channels, fewer residual blocks, and a self-attention at a single resolution (16×\times16).

Hyperparameters for training models are in Tab. A. We use γ=0.5\gamma=0.5 for FFHQ and CelebA-HQ as it achieve slightly better FIDs than γ=1.0\gamma=1.0 on those datasets. Models consist of one or two residual blocks per resolution and self-attention blocks at 16×\times16 resolution or at bottleneck layers of 8×\times8 resolution. Our default model has only 94M parameters, while recent works rely on large models (larger than 500M) . While recent works use 2 or 4 blocks per resolution, we use only one block, which leads to speed-up of training and inference. We use dropout when training on limited data. We trained models using EMA rate of 0.9999, 32-bit precision, and AdamW optimizer .

D Additional Results

Qualitative. Additional samples for all datasets mentioned in the paper are in Fig. D.

Quantitative. In Fig. 1 of the main text, we measured perceptual distances to investigate how the diffusion process corrupts perceptual contents. In Fig. 2, we qualitatively explored what a trained model learned at each step (Fig. 2). Here, we reproduce Fig. 1 at various datasets and resolutions in Fig. B and show the quantitative result of Fig. 2 in Fig. C. These results indicate that our investigation in Sec. 3.1 holds for various datasets and resolutions.