On the Importance of Noise Scheduling for Diffusion Models
Ting Chen
Why is noise scheduling important for diffusion models?
Diffusion models define a noising process of data by where is an input example (e.g., an image), is a sample from a isotropic Gaussian distribution, and is a continuous number between 0 and 1. The training of diffusion models is simple: we first sample to diffuse the input example to , and then train a denoising network to predict either noise or clean data . As is uniformly distributed, the noise schedule determines the distribution of noise levels that the neural network is trained on.
The importance of noise schedule can be demonstrated by the example in Figure 2. As we increase the image size, the denoising task at the same noise level (i.e. the same ) becomes simpler. This is due to the redundancy of information in data (e.g., correlation among nearby pixels) typically increases with the image size. Furthermore, the noises are independently added to each pixels, making it easier to recover the original signal when image size increases. Therefore, the optimal schedule at a smaller resolution may not be optimal at a higher resolution. And if we do not adjust the scheduling accordingly, it may lead to under training of certain noise levels. Similar observations are made in concurrent work .
Strategies to adjust noise scheduling
Built on top of existing work related to noise scheduling , we systematically study two different noise scheduling strategies for diffusion models.
The first strategy is to parameterized noise schedule with a one-dimensional function . Here we present ones based on part of cosine or sigmoid functions, with temperature scaling. Note that the original Cosine schedule is proposed in , with a fixed part of cosine curve that cannot be adjusted, and the simgoid schedule is proposed in . Other than these two types of functions, we further propose a simple linear noise schedule function, which is just (note that this is not the linear schedule proposed in ). Algorithm 1 presents the code for these instantiations of the continuous time noise schedule function .
Figure 3 visualizes the noise schedule functions under different choice of hyper-parameters, and their corresponding logSNR (signal-to-noise ratio). We can see that both cosine and sigmoid functions can parameterize a rich set of noise distributions. Please note that here we choose the hyper-parameters so that the noise distribution is skewed towards noisier levels, which we find to be more helpful.
2 Strategy 2: adjusting input scaling factor
Another way to indirectly adjust noise scheduling, proposed in , is to scale the input by a constant factor , which results in the following noising processing.
When , the variance of can change even has the same mean and variance as , which could lead to decreased performance . In this case, to ensure the variance keep fixed, one can scale by a factor of . However, in practice, we find that it works well by simply normalize the by its variance to make sure it has unit variance before feeding it to the denoising network . This variance normalization operation can also be seen as the first layer of the denoising network.
While this input scaling strategy is similar to changing the noise scheduling function above, it achieves slightly different effect in the logSNR when compared to cosine and sigmoid schedules, particularly when is closer to 0, as shown in Figure 5. In fact, the input scaling shifts the logSNR along y-axis while keeping its shape unchanged, which is different from all the noise schedule functions considered above. Although, one may also equivalently parameterize function in other ways to avoid scaling the inputs, as nicely demonstrated by the concurrent work .
3 Putting it together: a simple compound noise scheduling strategy
Here we propose to combine these two strategies by having a single noise schedule function, such as , and scale the input by a factor of . The training and inference strategies are given in the following.
Algorithm 2 shows how to incorporate the combined noising scheduling strategy into the training of diffusion models, with main changes highlighted in blue.
Inference/sampling strategy
If the variance normalization is used during the training, it should also be used during the sampling (i.e., the normalization can be seen as the first layer of the denoising network). Note that since we use a continuous time steps , so the inference schedule does not need to be the same as training schedule. During the inference we use a uniform discretization of the time between 0 and 1 into a given number of steps, and then we can chose a desired function to determine the level of noises at inference time. In practice, we find that standard cosine schedule works well for sampling.
Experiments
We mainly conduct experiments on class-conditional ImageNet image generation, and we follow common practice of evaluation, using FID and Inception Score as metrics computed on 50K samples, generated by 1000 steps of DDPM.
We follow for model specification but use smaller models as well as shorter overall training steps (except for >256 resolutions) to conserve compute. This results in worse performance in general but due to the improvement of noise scheduling, we can still achieve similar performance at lower resolutions (6464 and 128128), but significantly better results at higher resolutions (256256 or higher).
For hyper-parameters, we use LAMB optimizer with and weight decay of 0.01, self-conditioning rate of 0.9, and EMA decay of 0.9999. Table 1 and 2 summarize major hyper-parameters.
2 The effect of strategy 1 (noise schedule functions)
We first keep the input scaling fixed to 1, and evaluate the effect of noise schedules based on cosine, sigmoid and linear functions. As shown in Table 3, different image resolutions require different noise schedule functions to obtain the best performance, and it is difficult to find the optimal schedule due to several hyper-parameters involved.
3 The effect of strategy 2 (input scaling)
Here we keep the noise schedule functions fixed, and adjust the input scaling factor. The results are shown in Table 4. We find that 1) as image resolution increases, the optimal input scaling factor becomes smaller, 2) compared to the best result from Table 3 where we only change the noise schedule function while keeping input scaling fixed, adjusting input scaling is better (drop FID from 4.28 to 3.52 for ), and it is also easier to find as we can just tune a single scaling factor. Finally, seems to be a slightly better noise schedule than cosine (s=0.2,e=1,).
Table 5 demonstrates that the simple compound noise scheduling strategy, combined with RIN , enables state-of-the-art generation of high resolution images based on pure pixels. We forgo latent diffusion models where “pixels” are replaced with learned latent codes, since our scheduling technique is only tested on pixel-based diffusion models, but note these are orthogonal techniques and can potentially be combined. We note that state-of-the-art GANs can achieve similar or better performance but with multi-stage generation, as well as classifier-guidance , which we do not use for quantitative evaluation.
5 Visualization of generated samples
Even though we do not use label dropout for images at resolution of , we still find the classifier-free guidance during sampling improves the fidelity of generated samples. Therefore, we generate all the visualization samples with a guidance weight of 3. Figure 6, 7 and 8 show image samples generated from our trained model. Note that these are random samples, without cherry picking, generated conditioned on the given classes. Overall, we do see the global structure is well preserved across various resolutions, though object parts at smaller scale may be imperfect. We believe it can be improved with scaling the model and/or dataset (e.g., with more detailed text descriptions instead of just the class labels), and also the hyper-parameters tuning (as we do not thoroughly tune them for high resolutions).
Conclusion
In this work, we empirically study noise scheduling strategies for diffusion models and show their importance. The noise scheduling not only plays an important role in image generation but also for other tasks such as panoptic segmentation . A simple strategy of adjusting input scaling factor works well across different image resolutions. When combined with recently proposed RIN architecture , our noise scheduling strategy enables single-stage generation of high resolution images. For practitioners, our work suggests that it is important to select a proper noise scheduling scheme when training diffusion models for a new task or a new dataset.
Acknowledgements
We thank David Fleet and Allan Jabri for helpful discussions.