Efficient Diffusion Training via Min-SNR Weighting Strategy

Tiankai Hang, Shuyang Gu, Chen Li, Jianmin Bao, Dong Chen, Han Hu, Xin Geng, Baining Guo

Introduction

In recent years, denoising diffusion models have emerged as a promising new class of deep generative models due to their remarkable ability to model complicated distributions. Compared to prior Generative Adversarial Networks (GANs), diffusion models have demonstrated superior performance across a range of generation tasks in various modalities, including text-to-image generation , image manipulation , video synthesis , text generation , 3D avatar synthesis , etc. A key limitation of present denoising diffusion models is their slow convergence rate, requiring substantial amounts of GPU hours for training . This constitutes a considerable challenge for researchers seeking to effectively experiment with these models.

In this paper, we first conducted a thorough examination of this issue, revealing that the slow convergence rate likely arises from conflicting optimization directions for different timesteps during training. In fact, we find that by dedicatedly optimizing the denoising function for a specific noise level can even harm the reconstruction performance for other noise levels, as shown in Figure 2. This indicates that the optimal weight gradients for different noise levels are in conflict with one another. Given that current denoising diffusion models employ shared model weights for various noise levels, the conflicting weight gradients will impede the overall convergence rate, if without careful consideration on the balance of these noise timesteps.

To tackle this problem, we propose the Min-SNR-γ\gamma loss weighting strategy. This strategy treats the denoising process of each timestep as an individual task, thus diffusion training can be considered as a multi-task learning problem. To balance various tasks, we assign loss weights for each task according to their difficulty. Specifically, we adopt a clamped signal-to-noise ratio (SNR) as loss weight to alleviate the conflicting gradients issue. By organizing various timesteps using this new weighting strategy, the diffusion training process can converge much faster than previous approaches, as illustrated in Figure 1.

Generic multi-task learning methods usually seek to mitigate conflicts between tasks by adjusting the loss weight of each task based on their gradients. One classical approach , Pareto optimization, aims to seek a gradient descent direction to improve all the tasks. However, these approaches differ from our Min-SNR-γ\gamma weighting strategy in three aspects: 1) Sparsity. Most previous studies in the generic multi-task learning field have focused on scenarios with a small number of tasks, which differs from the diffusion training where the number of tasks can be up to thousands. As in our experiments, Pareto optimal solutions in diffusion training tend to set loss weights of most timesteps as 0. In this way, many timesteps will be left without any learning, and thus harm the entire denoising process. 2) Instability. The gradients computed for each timestep in each iteration are often noisy, owing to a limited number of samples for each timestep. This hampers the accurate computation of Pareto optimal solutions. 3) Inefficiency. The calculation of Pareto optimal solutions is time-consuming, significantly slowing down the overall training.

Our proposed Min-SNR-γ\gamma strategy is a predefined global step-wise loss weighting setting, instead of run-time adaptive loss weights for each iteration as in the original Pareto optimization, thus avoiding the sparsity issue. Moreover, the global loss weighting strategy eliminates the need for noisy computation of gradients and the time-consuming Pareto optimization process, making it more efficient and stable. Though suboptimal, the global strategy can be also almost as effective: Firstly, the optimization dynamics of each denoising task are largely shaped by the task’s noise level, without the need to account for individual samples too much. Secondly, after a moderate number of iterations, the gradients of the majority subsequent training process become more stable, thus it can be approximated by a stationery weighting strategy.

To validate the effectiveness of the Min-SNR-γ\gamma weighting strategy, we first compute its Pareto objective value and compare it with the optimal step-wise loss weights obtained by directly solving the Pareto problem. Together, we also compare it with several conventional loss weighting strategies, including constant weighting, SNR weighting, and SNR with an lower bound. Figure 4 shows that our Min-SNR-γ\gamma weighting strategy produces Pareto objective values almost as low as the optimal one, significantly better than other existing works, indicating a significant alleviation of the gradient conflicting issue. As a result, the proposed weighting strategy not only converges much faster than previous approaches, but is also effective and general for various generation scenarios. It achieves a new record of FID score 2.06 on the ImageNet 256×\times256 benchmark, and proves to also improve models using other prediction targets and network architectures.

Our contributions are summarized as follows:

We have uncovered a compelling explanation for the slow convergence issue in diffusion training: a conflict in gradients across various timesteps.

We have proposed a new loss weighting strategy for diffusion model training, which greatly mitigates the conflicting gradients across timesteps and results in a marked acceleration of convergence speed.

We have established a new FID score record on the ImageNet 256×256256\times 256 image generation benchmark.

Related Works

Denoising Diffusion Models. Diffusion models are strong generative models, particularly in the field of image generation, due to their ability to model complex distributions. This advantage has led to superiority over previous GAN models in terms of both high-fidelity and diversity of generated images . Besides, diffusion models also show great success in text-to-video generation , 3D Avatar generation , image to image translation , image manipulation , music generation , and even drug discovery . The most widely used network structure for diffusion models in the field of image generation is UNet . Recently, researchers have also explored the use of Vision Transformers as an alternative, with U-ViT borrowing the skip connection design from UNet and DiT leveraging Adaptive LayerNorm and discovering that the zero initialization strategy is critical for achieving state-of-the-art class-conditional ImageNet generation results.

Improved Diffusion Models. Recent studies have tried to improve the diffusion models from different perspectives. Some works aim to improve the quality of generated images by guiding the sampling process . Other studies propose fast sampling methods that require only a dozen steps to generating high-quality images. Some works have further distilled the diffusion models for even fewer steps in the sampling process . Meanwhile, some researchers have noticed that the noise schedule is important for diffusion models. Other works have found that different predicting targets from denoising networks affect the training stability and final performance. Finally, some works have proposed using the Mixture of Experts (MoE) approach to handle noise from different levels, which can boost the performance of diffusion models, but require a larger number of parameters and longer training time.

Multi-task Learning. The goal of Multi-task learning (MTL) is to learn multiple related tasks jointly so that the knowledge contained in a task can be leveraged by other tasks. One of the main challenges in MTL is negative transfer , means the joint training of tasks hurts learning instead of helping it. From an optimization perspective, it manifests as the presence of conflicting task gradients. To address this issue, some previous works try to modulate the gradient to prevent conflicts. Meanwhile, other works attempt to balance different tasks through carefully design the loss weights . GradNorm considers loss weight as learnable parameters and updates them through gradient descent. Another approach MTO regards the multi-task learning problem as a multi-objective optimization problem and obtains the loss weights by solving a quadratic programming problem.

Method

Diffusion models consist of two processes: a forward noising process and a reverse denoising process. We denote the distribution of training data as p(x0)p(\mathbf{x}_{0}). The forward process is a Gaussian transition, gradually adds noise with different scales to a real data point x0∼p(x0)\mathbf{x}_{0}\sim p(\mathbf{x}_{0}) to obtain a series of noisy latent variables {x1,x2,…,xT}\{\mathbf{x}_{1},\mathbf{x}_{2},\ldots,\mathbf{x}_{T}\}:

where ϵ\boldsymbol{\epsilon} is the noise sampled from Gaussian distribution N(0,I)\mathcal{N}(0,\mathbf{I}). The noise schedule σt\sigma_{t} denotes the magnitude of noise added to the clean data at tt timestep. It increases monotonically with tt. In this paper, we adopt the standard variance-preserving diffusion process, where αt=1−σt2\alpha_{t}=\sqrt{1-\sigma_{t}^{2}}.

The reverse process is parameterized by another Gaussian transition, gradually denoises the latent variables and restores the real data x0\mathbf{x}_{0} from a Gaussian noise:

μ^θ\mathbf{\hat{\mu}}_{\theta} and Σ^θ\hat{\Sigma}_{\theta} are predicted statistics. Ho et al. set Σ^θ(xt)\hat{\Sigma}_{\theta}(\mathbf{x}_{t}) to the constant σt2I\sigma_{t}^{2}\mathbf{I}, and μ^θ\hat{\mu}_{\theta} can be decomposed into the linear combination of xt\mathbf{x}_{t} and a noise approximation model ϵ^θ\hat{\epsilon}_{\theta}. They find using a network to predict noise ϵ\mathbf{\epsilon} works well, especially when combined with a simple re-weighted loss function:

Most previous works follow this strategy and predict the noise. Later works use another re-parameterization that predicts the noiseless state x0x_{0}:

And some other works even employ the network to directly predict velocity vv. Despite their prediction targets being different, we can derive that they are mathematically equivalent by modifying their loss weights.

2 Diffusion Training as Multi-Task Learning

To reduce the number of parameters, previous studies often share the parameters of the denoising models across all steps. However, it’s important to keep in mind that different steps may have vastly different requirements. At each step of a diffusion model, the strength of the denoising varies. For example, easier denoising tasks (when t→0t\to 0) may require simple reconstructions of the input in order to achieve lower denoising loss. This strategy, unfortunately, does not work as well for noisier tasks (when t→Tt\to T). Thus, it’s extremely important to analyze the correlation between different timesteps.

In this regard, we conduct a simple experiment. We begin by clustering the denoising process into several separate bins. Then we finetune the diffusion model by sampling timesteps in each bin. Lastly, we evaluate its effectiveness by looking at how it impacted the loss of other bins. As shown in Figure 2, we can observe that finetuning specific steps benefited those surrounding steps. However, it’s often detrimental for other steps that are far away. This inspires us to consider whether we can find a more efficient solution that benefits all timesteps simultaneously.

We re-organized our goal from the perspective of multitask learning. The training process of denoising diffusion models contains TT different tasks, each task represents an individual timestep. We denote the model parameters as θ\theta and the corresponding training loss is Lt(θ),t∈{1,2,…,T}\mathcal{L}^{t}(\theta),t\in\{1,2,\ldots,T\}. Our goal is to find a update direction δ≠0\delta\neq 0, that satisfies:

We consider the first-order Taylor expansion:

Thus, the ideal update direction is equivalent to satisfy:

3 Pareto optimality of diffusion models

Consider a update direction δ∗\delta^{*}:

of which wtw_{t} is the solution to the optimization problem:

If the optimal solution to the Equation 8 exists, then δ∗\delta^{*} should satisfy it. Otherwise, it means that we must sacrifice a certain task in exchange for the loss decrease of other tasks. In other words, we have reached the Pareto Stationary and the training has converged.

A more general form of this theorem was first proposed in and we leave a succinct proof in the appendix. Since diffusion models are required to go through all the timesteps when generating images. So any timestep should not be ignored during training. Consequently, a regularization term is included to prevent the loss weights from becoming excessively small. The optimization goal in Equation 10 becomes:

where λ\lambda controls the regularization strength.

To solve Equation 11, leverages the Frank-Wolfe algorithm to obtain the weight {wt}\{w_{t}\} through iterative optimization. Another approach is to adopt Unconstrained Gradient Descent(UGD). Specifically, we re-parameterize wtw_{t} through βt\beta_{t}:

Combined with Equation 11, we can use gradient descent to optimize each term independently:

However, whether leveraging the Frank-Wolfe or the UGD algorithm, there are two disadvantages: 1) Inefficiency. Both of these two methods need additional optimization at each training iteration, it greatly increases the training cost. 2) Instability. In practice, by using a limited number of samples to calculate the gradient term ∇θLt(θ)\nabla_{\theta}\mathcal{L}^{t}(\theta), the optimization results are unstable(as shown in Figure 3). In other words, the loss weights for each denoising task vary greatly during training, making the entire diffusion training inefficient.

4 Min-SNR-γ𝛾\gamma Loss Weight Strategy

In order to avoid the inefficiency and instability caused by the iterative optimization in each iteration, one possible attempt is to adopt a stationery loss weight strategy.

To simplify the discussion, we assume that the network is reparametered to predict the noiseless state x0\mathbf{x}_{0}. However, it’s worth noting that different prediction objectives can be transformed into one another, we will delve into it in Section 4.2. Now, we consider the following alternative training loss weights:

Constant weighting. wt=1w_{t}=1. Which treats different tasks as equally weighted and has been used in both discrete diffusion models and continuous diffusion models .

SNR weighting. wt=SNR(t)w_{t}=\text{SNR}(t), where SNR(t)=αt2/σt2\text{SNR}(t)=\alpha_{t}^{2}/\sigma_{t}^{2}. It’s the most widely used weighting strategy . By combining with Equation 2, we can find it’s numerically equivalent to the constant weighting strategy when the predicting target is noise.

Max-SNR-γ\gamma weighting. wt=max⁡{SNR(t),γ}w_{t}=\max\{\text{SNR}(t),\gamma\}. This modification of SNR weighting is first proposed in to avoid a weight of zero with zero SNR steps. They set γ=1\gamma=1 as their default setting. However, the weights still concentrate on small noise levels.

Min-SNR-γ\gamma weighting. wt=min⁡{SNR(t),γ}w_{t}=\min\{\text{SNR}(t),\gamma\}. We propose this weighting strategy to avoid the model focusing too much on small noise levels.

UGD optimization weighting. wtw_{t} is optimized from Equation 13 in each timestep. Compared with the previous setting, this strategy changes during training.

First, we combine these weighting strategies into Equation 11 to validate whether they are approach to the Pareto optimality state. As shown in Figure 4, the UGD optimization weighting strategy can achieve the lowest score on our optimization target. In addition, the Min-SNR-γ\gamma weighting strategy is the closest to the optimum, demonstrating it has the property to optimize different timesteps simultaneously.

In the following section, we present experimental results to demonstrate the effectiveness of our Min-SNR-γ\gamma weighting strategy in balancing diverse noise levels. Our approach aims to achieve faster convergence and strong performance.

Experiments

In this section, we first provide an overview of the experimental setup. Subsequently, we conduct comprehensive ablation studies to show that our method is versatile and suitable for various prediction targets and network architectures. Finally, we compare our approach to the state-of-the-art methods across multiple image generation benchmarks, demonstrating not only its accelerated convergence but also its superior capability in generating high-quality images.

Datasets. We perform experiments on both unconditional and conditional image generation using the CelebA dataset and the ImageNet dataset . The CelebA dataset, which comprises 162,770 human faces, is a widely-used resource for unconditional image generation studies. We follow ScoreSDE for data pre-processing, which involves center cropping each image to a resolution of 140×140140\times 140 and then resizing it to 64×6464\times 64. For the class conditional image generation, we adopt the ImageNet dataset with a total of 1.3 million images from 1000 different classes. We test the performance on both 64×6464\times 64 and 256×256256\times 256 resolutions.

Training Details. For low resolution (64×6464\times 64) image generation, we follow ADM and directly train the diffusion model on the pixel-level. For high-resolution image generation, we utilize LDM approach by first compressing the images into latent space, then training a diffusion model to model the latent distributions. To obtain the latent for images, we employ VQ-VAE from Stable Diffusionhttps://huggingface.co/stabilityai/sd-vae-ft-mse-original, which encodes a high-resolution image (256×256×3256\times 256\times 3) into 32×32×432\times 32\times 4 latent codes.

In our experiments, we employ both ViT and UNet as our diffusion model backbones. We adopt a vanilla ViT structure without any modifications as our default setting. we incorporate the timestep tt and class condition c\mathbf{c} as learnable input tokens to the model. Although further customization of the network structure may improve performance, our focus in this paper is to analyze the general properties of diffusion models. For the UNet structure, we follow ADM and keep the FLOPs similar to the ViT-B model, which has 1.5×1.5\times parameters. Additional details can be found in the appendix.

For the diffusion settings, we use a cosine noise scheduler following the approach in . The total number of timesteps is standardized to T=1000T=1000 across all datasets. We adopt AdamW as our optimizer. For the CelebA dataset, we train our model for 500K iterations with a batch size of 128. During the first 5,000 iterations, we implement a linear warm-up and keep the learning rate at 1×10−41\times 10^{-4} for the remaining training. For the ImageNet dataset, the default learning rate is fixed at 1×10−41\times 10^{-4}. The batch size is set to 10241024 for 64264^{2} resolution and 256256 for 2562256^{2} resolution.

Evaluation Settings. To evaluate the performance of our models, we utilize an Exponential Moving Average (EMA) model with a rate of 0.9999. During the evaluation phase, we generate images with the Heun sampler from EDM . For conditional image generation, we also implement the classifier-free sampling strategy to achieve better results. Finally, we measure the quality of the generated images using the FID score calculated on 50K images.

2 Analysis of the Proposed Min-SNR-γ𝛾\gamma

Comparison of Different Weighting Strategies. To demonstrate the significance of the loss weighting strategy, we conduct experiments with different loss weight settings for predicting x0\mathbf{x}_{0}. These settings include: 1) constant weighting, where wt=1w_{t}=1, 2) SNR weighting, with wt=SNR(t)w_{t}=\text{SNR}(t), 3) truncated SNR weighting, with wt=max⁡{SNR(t),γ}w_{t}=\max\{\text{SNR}(t),\gamma\} (following with a set value of γ=1\gamma=1), and 4) our proposed Min-SNR-γ\gamma weighting strategy, with wt=min⁡{SNR(t),γ}w_{t}=\min\{\text{SNR}(t),\gamma\}, we set γ=5\gamma=5 as the default value.

The ViT-B serves as our default backbone and experiments are performed on ImageNet 256×256256\times 256. As illustrated in Figure 5, we observe that all results improve as the number of training iterations increases. However, our method demonstrates a significantly faster convergence compared to other methods. Specifically, it exhibits a 3.4×3.4\times speedup in reaching an FID score of 1010. It is worth mentioning that the SNR weighting strategy performed the worst, which could be due to its disproportionate focus on less noisy stages.

For a deeper understanding of the reasons behind the varying convergence rates, we analyzed their training loss at different noise levels. For a fair comparison, we exclude the loss weight term by only calculating ∥x0−x^θ∥22\lVert\mathbf{x}_{0}-\mathbf{\hat{x}_{\theta}}\rVert_{2}^{2}. Considering that the loss of different noise levels varies greatly, we calculate the loss in different bins and present the results in Figure 6. The results show that while the constant weighting strategy is effective for high noise intensities, it performs poorly at low noise intensities. Conversely, the SNR weighting strategy exhibits the opposite behavior. In contrast, our proposed Min-SNR-γ\gamma strategy achieves a lower training loss across all cases, and indicates quicker convergence through the FID metric.

Furthermore, we present visual results in Figure 7 to demonstrate the fast convergence of the Min-SNR-γ\gamma strategy. We apply the same random seed for noise to sample images from training iteration 50K, 200K, 400K, and 1M with different loss weight settings. Our results show that the Min-SNR-γ\gamma strategy generates a clear object with only 200K iterations, which is significantly better in quality than the results obtained by other methods.

Min-SNR-γ\gamma for Different Prediction Targets. Instead of predicting the original signal x0\mathbf{x}_{0} from the network, some recent works have employed alternative re-parameterizations, such as predicting noise ϵ\epsilon, or velocity v\mathbf{v} . To verify the applicability of our weighting strategy to these prediction targets, we conduct experiments comparing the four aforementioned weighting strategies across these different re-parameterizations.

As we discussed in Section 3.4, predicting noise ϵ\epsilon is mathematically equivalent to predicting x0\mathbf{x}_{0} by intrinsically involving Signal-to-Noise Ratio as a weight factor, thus we divide the SNR term in practice. For example, the Min-SNR-γ\gamma strategy in predicting noise can be expressed as wt=min⁡{SNR(t),γ}SNR(t)=min⁡{γSNR(t),1}w_{t}=\frac{\min\{\text{SNR}(t),\gamma\}}{\text{SNR}(t)}=\min\{\frac{\gamma}{\text{SNR}(t)},1\}. And the SNR strategy in predicting noise is equivalent to a “constant strategy”. For simplicity and consistency, we still refer to them as Min-SNR-γ\gamma and SNR strategies. Similarly, we can derive that when predicting velocity v\mathbf{v}, the loss weight factor must be divided by (SNR+1)(\text{SNR}+1). These strategies are still referred to by their original names for ease of reference.

We conduct experiments on these two variants and present the results in Figure 5. Taking the neural network output as noise with const or Max-SNR-γ\gamma setting leads to divergence. Meanwhile, our proposed Min-SNR-γ\gamma strategy converges faster than other loss weighting strategies for both prediction noise and predicting velocity. These demonstrate that balancing the loss weights for different timesteps is intrinsic, independent of any re-parameterization.

Min-SNR-γ\gamma on Different Network Architectures. The Min-SNR-γ\gamma strategy is versatile and robust for different prediction targets and network structures. We conduct experiments on the widely used UNet and keep the number of parameters close to the ViT-B model. For each experiment, models were trained for 1 million iterations and their FID scores were calculated at multiple intervals. The results in Table 1 indicate that the Min-SNR-γ\gamma strategy converges significantly faster than the baseline and provides better performance for both predicting x0\mathbf{x}_{0} and predicting noise.

Robustness Analysis. Our approach utilizes a single hyperparameter, γ\gamma, as the truncate value. To assess its robustness, we conducted a thorough robustness analysis in various settings. Our experiments were performed on the ImageNet-256 dataset using the ViT-B model and the prediction target of the network is x0\mathbf{x}_{0}. We varied the truncate value γ\gamma by setting it to 1, 5, 10, and 20 and evaluated their performance. The results are shown in Table 2. We find there are only minor variations in the FID score when γ\gamma is smaller than 20. Additionally, we conducted more experiments by modifying the predicting target to the noise ϵ\epsilon, and modifying the network structure to UNet. We find that the results were also consistently stable. Our results indicate that good performance can usually be achieved when γ\gamma is set to 5, making it the established default setting.

3 Comparison with state-of-the-art Methods

CelebA-64. We conduct experiments on the CelebA 64×6464\times 64 dataset for unconditional image generation. Both UNet and ViT are used as our backbones and are trained for 500K iterations. During the evaluation, we use the EDM sampler to generate 50K samples and calculate the FID score. The results are summarized in Table 3. Our ViT-Small model outperforms previous ViT-based models with an FID score of 2.14. It is worth mentioning that no modifications are made to the naive network structure, demonstrating that the results could still be improved further. Meanwhile, our method using the UNet structure achieves an even better FID score of 1.60, outperforming previous UNet methods.

ImageNet-64. We also validate our method on class-conditional image generation on the ImageNet 64×6464\times 64 dataset. During training, the class label is dropped with the probability 0.150.15 for classifier-free inference . The model is trained for 800K iterations and images are synthesized using classifier-free guidance with a scale of cfg=1.5\text{cfg}=1.5 and the EDM sampler for image generation. For a fair comparison, we adopt a 21-layer ViT-Large model without additional architecture designs, which has a similar number of parameters to U-ViT-Large . The results presented in Table 4 show that our method achieves an FID score of 2.28, significantly improving upon the U-ViT-Large model.

ImageNet-256. We also apply diffusion models for higher-resolution image generation on the ImageNet 256×256256\times 256 benchmark. To enhance training efficiency, we first compress 256×256×3256\times 256\times 3 images into 32×32×432\times 32\times 4 latent codes using the encoder from LDM . During the sampling process, we employ the EDM sampler and the classifier-free guidance to generate images. The FID comparison is presented in Table 5. Under the setting of predicting ϵ\epsilon with Min-SNR-5, our ViT-XL model achieves the FID of 2.082.08 for only 2.1M iterations, which is 3.3×3.3\times faster than DiT and outperforms the previous state-of-the-art FID record of 2.272.27. Moreover, with longer training (about 7M iterations as in ), we are able to achieve the FID score of 2.06 by predicting x0\mathbf{x}_{0} with Min-SNR-5. Our UNet-based model with 395M parameters is trained for about 1.4M iterations and achieves FID score of 2.81.

Conclusion

In this paper, we point out that the conflicting optimization directions between different timesteps may cause slow convergence in diffusion training. To address it, we regard the diffusion training process as a multi-task learning problem and introduce a novel weighting strategy, named Min-SNR-γ\gamma, to effectively balance different timesteps. Experiments demonstrate our method can boost diffusion training several times faster, and achieves the state-of-the-art FID score on ImageNet-256 dataset.

References

Appendix A Proof for Theorem 1

First, we introduce the Pareto Optimality mentioned in the paper. Assume the loss for each task is Lt(θ),t∈{1,2,…,T}\mathcal{L}^{t}(\theta),t\in\{1,2,\ldots,T\} and the respective gradient to θ\theta is ∇θLt(θ)\nabla_{\theta}\mathcal{L}^{t}(\theta). For simplicity, we denote Lt(θ)\mathcal{L}^{t}(\theta) as Lt\mathcal{L}^{t}. If we treat each task with equal importance, we assume each loss item L1,L2,…,LT\mathcal{L}^{1},\mathcal{L}^{2},\ldots,\mathcal{L}^{T} is decreasing or kept the same. There exists one point θ∗\theta^{*} where any change of the point will leads to the increase of one loss item. We call the point θ∗\theta^{*} “Pareto Optimality”. In other words, we cannot sacrifice one task for another task’s improvement. To reach Pareto Optimality, we need to find an update direction δ\delta which meet:

⟨⋅,⋅⟩\left\langle\cdot,\cdot\right\rangle denotes the inner product of two vectors. It is worth noting that δ=0\delta=0 satisfies all the above inequalities. We care more about the non-zero solution and adopt it for updating the network parameter θ\theta. If the non-zero point does not exist, it may already achieve the “Pareto Optimality”, which is referred as “Pareto Stationary”.

For simplicity, we denote the gradient for each loss item ∇θLt\nabla_{\theta}\mathcal{L}^{t} as gt\mathbf{g}_{t}. Suppose we have a gradient vector u\mathbf{u} to satisfy that all ⟨gt,u⟩≥0,t∈{1,2,…,T}\left\langle\mathbf{g}_{t},\mathbf{u}\right\rangle\geq 0,t\in\{1,2,\ldots,T\}. Then −u-\mathbf{u} is the updating direction ensuring a lower loss for each task.

As proposed in , ⟨gt,u⟩≥0,∀t∈{1,2,…,T}\left\langle\mathbf{g}_{t},\mathbf{u}\right\rangle\geq 0,\forall t\in\{1,2,\ldots,T\} is equivalent to min⁡t⟨gt,u⟩≥0\min_{t}\left\langle\mathbf{g}_{t},\mathbf{u}\right\rangle\geq 0. And it could be achieved when the minimal value of ⟨gt,u⟩\left\langle\mathbf{g}_{t},\mathbf{u}\right\rangle is maximized. Thus the problem is further converted to:

There is no constraint for the vector u\mathbf{u}, so it may become infinity and make the updating unstable. To avoid it, we add a regularization term to it

And notice that the max⁡\max function ensures the value is always greater than or equal to a specific value u=0\mathbf{u}=0.

which also means max⁡umin⁡t⟨gt,u⟩≥12∥u∥22≥0\max_{\mathbf{u}}\min_{t}\left\langle\mathbf{g}_{t},\mathbf{u}\right\rangle\geq\frac{1}{2}\lVert\mathbf{u}\rVert_{2}^{2}\geq 0. Therefore, the solution of Equation 19 satisfies our optimization goal of ⟨gt,u⟩≥0,∀t∈{1,2,…,T}\left\langle\mathbf{g}_{t},\mathbf{u}\right\rangle\geq 0,\forall t\in\{1,2,\ldots,T\}.

We define CT\mathcal{C}^{T} as a set of nn-dimensional variables

We can also verify the above function is concave with respect to u\mathbf{u} and α\alpha. According to Von Neumann’s Minmax theorem , the objective with regularization in Equation 19 is equivalent to

Finally, we achieved Theorem 1 in the main paper.

Appendix B Relationship between Different Targets

The most common predicting target is in ϵ\epsilon-space. Loss for prediction in x0\mathbf{x}_{0}-space and ϵ\epsilon-space can be transformed by the SNR loss weight.

where ϵ^θ\hat{\epsilon}_{\theta} is the network to predict the noise and x^θ\hat{\mathbf{x}}_{\theta} is to predict the clean data.

Prediction target v=αtϵ−σtx0\mathbf{v}=\alpha_{t}\epsilon-\sigma_{t}\mathbf{x}_{0} is proposed in , we can derive the related loss

Appendix C Hyper-parameter

Here we list more details about the architecture, training and evaluation setting.

The ViT setting adopted in the paper are as follows,

We use ViT-Small for face generation on CelebA 64×6464\times 64. Besides, we adopt ViT-Base as the default backbone for the ablation study. To make relative fair comparison with U-ViT, we use a 21-layer ViT-Large for ImageNet 64×6464\times 64 benchmark. To compare with former state-of-the-art method DiT on ImageNet 256×256256\times 256, we adopt the similar setting ViT-XL with the same depth, hidden size, and patch size.

In the paper, we also evaluate our method’s robustness to model architectures using the UNet backbone. For ablation study, we adjust the setting based on ADM to make the parameters and FLOPs close to ViT-B. The setting is

We also conduct experiments with the same architecture (296M) in ADM on ImageNet 64×6464\times 64. After 900K training iterations with batch size 1024, it could achieve an FID score of 2.11.

For high resolution generation on ImageNet 256×256256\times 256. We use the 395M setting from LDM , which operates on the 32×32×432\times 32\times 4 latent space.

C.2 Training Settings

The training iterations and learning rate have been reported in the paper. We use AdamW as our default optimizer. (β1,β2)(\beta_{1},\beta_{2}) is set to (0.9,0.999)(0.9,0.999) for UNet backbone. Following , we set (β1,β2)(\beta_{1},\beta_{2}) to (0.99,0.99)(0.99,0.99) for ViT backbone.

C.3 Sampling Settings

If not otherwise specified, we only use EDM’s Heun sampler. We only adjust the sampling steps for better results. For ablation study with ViT-B and UNet, we set the number of steps to 30. For ImageNet 64×6464\times 64 in Table 4, the number of steps is set to 20. For ImageNet 256×256256\times 256 in Table 5, the number of sampling steps is set to 50.

Appendix D Additional Results

In the paper, most of the ablation study is conducted on ImageNet 256×256256\times 256’s latent space. Here, we present the results on ImageNet 64×6464\times 64 pixel space. We adopt a ViT-B model as our backbone and train the diffusion model for 800K iterations with batch size 512. Our predicting targets are x0\mathbf{x}_{0} and ϵ\epsilon and they are equipped with our proposed simple Min-SNR-γ\gamma loss weight (γ=5\gamma=5). We adopt the pre-trained noisy classifier at 64×6464\times 64 from ADM as conditional guidance. We can see that the loss weighting strategy contributes to the faster convergence for both x0\mathbf{x}_{0} and ϵ\epsilon.

D.2 Visual Results on Different Datasets

We provide additional generated results in Figure 9-12. Figure 9 shows the generated samples with UNet backbone on CelebA 64×6464\times 64. Figure 10 and Figure 11 demonstrate the generated samples on conditional ImageNet 64×6464\times 64 benchmark with ViT-Large and UNet backbone respectively. The visual results on CelebA 64×6464\times 64 and ImageNet 64×6464\times 64 are randomly synthesized without cherry-pick.

We also present some visual results on ImageNet 256×256256\times 256 with our model which can achieve the FID 2.06 in Figure 12.