Rolling Diffusion Models

David Ruhe, Jonathan Heek, Tim Salimans, Emiel Hoogeboom

Introduction

Diffusion models (Sohl-Dickstein et al., 2015; Song & Ermon, 2019; Ho et al., 2020) have significantly boosted the field of generative modeling. They provided the fundaments for large-scale text-to-image systems like DALL-E 2 (Ramesh et al., 2022), Imagen (Saharia et al., 2022), Parti (Yu et al., 2022), and Stable Diffusion (Rombach et al., 2022). Other applications of diffusion models include density estimation, text-to-speech, and image editing (Kingma et al., 2021; Gao et al., 2023; Kawar et al., 2023).

After these successes in these domains, interest in developing diffusion models for time sequences has grown. Prominent recent large-scale works include, e.g., Imagen Video (Ho et al., 2022a), Stable Diffusion Video (StabilityAI, 2023). Other impressive results for generating video data have been achieved, e.g., by (Blattmann et al., 2023; Ge et al., 2023; Harvey et al., 2022; Singer et al., 2022; Ho et al., 2022b). Applications of sequential generative modeling outside video include, e.g., fluid mechanics or weather and climate modeling (Price et al., 2023; Meng et al., 2022; Lippe et al., 2023).

What’s common across many of these works is that they treat the temporal axis as an “extra spatial dimension”. That is, they treat the video as a 3D tensor of shape K×H×WK\times H\times W. This has several downsides. First, the memory and computational requirements can quickly grow infeasible if one wants to generate long sequences. Second, one is typically interested in being able to roll out generation for a variable number of time steps. Therefore, an alternative angle is a fully autoregressive approach by conditioning on a sequence of input frames and simulating a single output frame, which is then concatenated to the input frames, upon which the recursion can continue. In this case, one has to traverse the entire denoising diffusion chain for every single frame, which is computationally intensive. Additionally, iteratively sampling single frames leads to quick autoregressive error accumulation. A middle ground can be found by jointly generating blocks of frames. However, in this block-autoregressive case, a diffusion model would use the same number of denoising steps for every frame. This is suboptimal since, given a sequence of input frames, the uncertainty about the first upcoming few is much lower than the later ones. Finally, both methods sample frames only jointly with earlier frames, which is potentially a suboptimal parameterization.

In this work, we propose a new framework called Rolling Diffusion, a method that explicitly corrupts data from past to future. This is achieved by reparameterizing the global diffusion time to a local time for each frame. It turns out that by doing this, one can (apart from boundary conditions) completely focus on a local sliding window sequential denoising process. This has several temporal inductive biases, alleviating some of the abovementioned issues.

The model only has to predict the low frequencies (corresponding to their global visual features) for distant frames, whereas the high-frequency details are generated for temporally near frames.

Each frame is generated together with both a number of preceding and succeeding frames.

Due to the local sliding window point of view, every frame enjoys the same inductive bias and undergoes a similar sampling procedure regardless of its absolute position in the video.

These merits are empirically demonstrated in, among others, a video prediction experiment using the Kinetics-600 video dataset and in an experiment involving chaotic fluid mechanics simulations.

Background: Diffusion Models

where ata_{t} and σt2\sigma_{t}^{2} are strictly positive scalar functions of tt. We define their signal-to-noise ratio

to be monotonically decreasing in tt. Finally, we let αt2+σt2=1\alpha_{t}^{2}+\sigma_{t}^{2}=1, corresponding to a variance-preserving process which also implies at2∈(0,1]a_{t}^{2}\in(0,1] and σt2∈(0,1]\sigma_{t}^{2}\in(0,1].

Given the noising process, it can be shown (Sohl-Dickstein et al., 2015) that the true (i.e., optimal) denoising distribution for a single datapoint x{\bm{x}} from time tt to time ss is given by

2 Diffusion for temporal data

If one is interested in generation of temporal data, potentially indefinitely (or beyond typical hardware constraints), one must consider (autoregressive) conditional extension of previously generated data. I.e., given an initial sample xk∼q(x){\bm{x}}^{k}\sim q({\bm{x}}) at a temporal index kk, we want to estimate and sample the conditional distribution p(xk+1∣xk)p({\bm{x}}^{k+1}|{\bm{x}}^{k}). This process can then be extended to videos of arbitrary lengths. As discussed in Section 1, it is not yet clear what kinds of parameterization choices are optimal to estimate this conditional distribution. Further, no temporal inductive bias is typically baked into the denoising process.

Rolling Diffusion Models

We introduce rolling diffusion models, merging the arrow of time with the (de)noising process. To formalize this, we first have to discuss the global diffusion model. We will see that the only nontrivial parts of the global process take place locally. Defining the noise schedule locally is advantageous since the resulting model does not depend on the number of frames KK and can be unrolled indefinitely.

Note that we still require tk∈t_{k}\in for all k∈{0,…,K−1}{k\in\{0,\dots,K-1\}}. Furthermore, we still have a monotonically decreasing signal-to-noise schedule, ensuring a well-defined diffusion process. However, we now (effectively) have a different signal-to-noise schedule for each frame. In this work, we also always have tk≤tk+1t_{k}\leq t_{k+1} i.e., the local time of a frame is smaller than the local time of the next frame. This means we add more noise to future frames: a natural temporal inductive bias. Note that this is not strictly required; one could also have a reverse-time inductive bias or a mixture. An example of such a reparameterization is shown in Figure 1 (left). We depict a map that takes a global diffusion time tt (vertical axis) and a frame index kk (horizontal axis), and computes a local time tkt_{k}, indicated with a color intensity.

We now redefine the forward process using the local time:

where we can reuse the α\alpha and σ\sigma functions (now evaluated locally at tkt_{k}) from before. Here, xk{\bm{x}}^{k} denotes the kk-th frame of x{\bm{x}}.

True backward process and generative process

Given a tuple (s,t)(s,t), s∈s\in, t∈t\in, s≤ts\leq t, we can divide the frames k∈{0,…,K−1}k\in\{0,\dots,K-1\} into three categories:

This is helpful because we will see that the only frames that need to be modeled are in the window. Namely, the first factor has

In other words, if ztk{\bm{z}}^{k}_{t} is already noiseless, then zsk{\bm{z}}^{k}_{s} will also be noiseless. Regarding the second factor, we see that they are all independently normally distributed:

Simply put, in these cases zsk{\bm{z}}^{k}_{s} is independent noise and does not depend on data at all. Finally, the third factor has a true non-trivial denoising process:

where μtk→sk\mu_{t_{k}\to s_{k}} and σtk→sk2\sigma_{t_{k}\to s_{k}}^{2} are the analytical mean and variance functions. Note that we can then optimally (w.r.t. a KL-divergence) factorize the generative process similarly:

with p(zsclean∣zt):=∏k∈clean(s,t)δ(zsk∣ztk)p({\bm{z}}_{s}^{\text{clean}}|{\bm{z}}_{t}):=\prod_{k\in\text{clean}(s,t)}\delta({\bm{z}}^{k}_{s}|{\bm{z}}^{k}_{t}) and p(zsnoise∣zt):=∏k∈noise(s,t)N(zsk∣0,I)p({\bm{z}}_{s}^{\text{noise}}|{\bm{z}}_{t}):=\prod_{k\in\text{noise}(s,t)}\mathcal{N}({\bm{z}}^{k}_{s}|0,{\mathbf{I}}). The only ‘interesting’ parameterized part of the generative process then has

In other words, we can only focus the generative process on the frames that are in the sliding window. Finally, note that we can choose to not condition the model on all ztk{\bm{z}}^{k}_{t} that have tkt_{k} = 0, since frames that are far in the past are likely to be independent of the current frame, and this excessive conditioning would exceed computational constraints. As such, we get

In practice, we approximate pθ(zswin∣ztclean,ztwin):≈pθ(zswin∣zclean^,ztwin)p_{\theta}({\bm{z}}_{s}^{\text{win}}|{\bm{z}}^{\text{clean}}_{t},{\bm{z}}^{\text{win}}_{t}):\approx p_{\theta}({\bm{z}}_{s}^{\text{win}}|{\bm{z}}^{\widehat{\text{clean}}},{\bm{z}}^{\text{win}}_{t}), where zclean^{\bm{z}}^{\widehat{\text{clean}}} denotes a specific (potentially empty) subset of ztclean{\bm{z}}^{\text{clean}}_{t} used for extra conditioning. This typically includes a few frames slightly before the current sliding window.

In Appendix D, we formalize the above and show that the objective results in

where we suppress some arguments for notational convenience.

Observe Figure 1 again. After training is completed, we can essentially sample from the generative model by traversing the image with the sliding window from the top left to the bottom right. This allows us to completely focus on the local environment, as discussed in the next section.

2 A local perspective

In the last section, we saw that rolling diffusion allows us to focus the generative process as well as training solely on frames that are in the sliding window. As such, we now have twin(t):→t_{\text{win}}(t):\to denote a time reparameterization subject to the usual constraints. Note, though, that the shared time-component tt is now defined locally. Specifically, running the denoising chain from t=1t=1 to t=0t=0 will only sample a rolling window such that the first frame is completely noiseless and the following frames still contain some noise. In contrast, the global process described earlier denoises an entire video.

In principle, the design space of pure rolling diffusion models is enormous. For this reason, we make the following sensible assumptions in addition to the earlier constraints on the signal-to-noise schedule to allow for a sliding window sampling procedure. We would like twint_{\text{win}} to be:

local, meaning we allow for sharing and reusing the parameterization across various positions of the sliding window, independent of their absolute locations.

consistent under moving the window, meaning that twin(1)t_{\text{win}}(1) should equal twin(0)t_{\text{win}}(0) of the next frame.

Let W<KW<K be the size of the sliding window, and w∈{0,…,W−1}w\in\{0,\dots,W-1\} be the local indices of the frames. To satisfy the first assumption, we define the schedule in terms of the local index ww (see Figure 1 (right)).

where twW:→t_{w}^{W}:\to is a monotonically increasing function. For the second, we know twWt^{W}_{w} must have

for some monotonically increasing (in tt) function g:→g:\to. We will sometimes suppress WW for notational convenience. Note that due to the locality of the parameterization, the process can be unfolded indefinitely at test time.

where w∈{0,…,W−1}w\in\{0,\dots,W-1\}. See Figure 1 (right) for an illustration of how this local schedule is applied to each sequence of frames. Observe that

One can extend the linear local time to include clean conditioning frames. Let nclnn_{\text{cln}} denote the number of clean frames (chosen as a hyperparameters), then the local time for a frame ww is:

3 Boundary conditions

While framing rolling diffusion purely from a local perspective is convenient for training and sampling, it introduces complications at the boundaries of the sliding window. That is, given, e.g., the linear local time reparameterization twlint^{\text{lin}}_{w}, we have that given the diffusion time tt running from 11 to , the local times run from (1W,2W,…WW)(\frac{1}{W},\frac{2}{W},\ldots\frac{W}{W}) to (0W,1W,…,W−1W)(\frac{0}{W},\frac{1}{W},\ldots,\frac{W-1}{W}). Visually, in Figure 1, the rolling sampling procedure can be seen as moving the sliding window over the diagonal linearly from top left to bottom right such that the local times of the frames remain invariant upon shifting. However, this means that placing the window at the very left edge still results in having partially denoised frames. This means that in this local setting, the signal-to-noise ratios are never minimal, i.e., at full noise.

To account for this, we co-train the rolling diffusion model with an additional schedule that can handle this boundary condition.

This init noise schedule can start from random noise and generates a video in the “rolling state”. Note that this schedule cannot be used as a sliding window noise schedule, however, it contains twlint_{w}^{\text{lin}} as we will see later. That is, at diffusion time t=1t=1, this will put all frames to maximum noise, and at t=0t=0 the frames will be in the rolling state. To be precise, it starts from local times (1,1,…1)(1,1,\ldots 1) and denoises to (0,1W,2W,…,W−1W)(0,\frac{1}{W},\frac{2}{W},\dots,\frac{W-1}{W}), after which we can start the previously described local rolling diffusion process. From a visual perspective, in Figure 1, this corresponds to placing the window at the upper left corner and moving it down vertically, until it reaches the rolling state that can be used to continue diagonally.

On the domain [0,1W][0,\frac{1}{W}], this schedule contains the previous local schedule twlint_{w}^{\text{lin}} as a special case This means that the model could be trained solely with twinitt_{w}^{\text{init}} to handle the boundaries as well as being able to roll out indefinitely. Note, however, that the twlint^{\text{lin}}_{w} schedule only gets selected 1/W1/W of the time during training (assuming t∼U(0,1)t\sim U(0,1)). In contrast, this schedule is used almost exclusively at test time, with the exception being the boundary condition on the first WW frames. As such, we found it beneficial to include both schedules during training. This blend is achieved by sampling one schedule or the other based on a Bernoulli hyperparameter β\beta controlling the probability of selecting between the two schedules.

4 Local training

The training and sampling procedures are summarized in Algorithm 2, Algorithm 3, Algorithm 1. Furthermore, we provide a visual of the rolling sampling loop in Figure 6.

Related Work

Video diffusion has been studied and applied directly in pixel space (Ho et al., 2022a, b; Singer et al., 2022) and in latent space (Blattmann et al., 2023; Ge et al., 2023; He et al., 2022; Yu et al., 2023c), the latter typically empirically being slightly more effective. Furthermore, these videos usually extend the two-dimensional image setting to three (two spatial dimensions and one temporal dimension) without considering autoregressive extension.

Methods that specifically treat test-time unrolling of video generation include Yang et al. (2023); Harvey et al. (2022). It was shown that directly parameterizing the conditional distribution of future frames given past frames is preferable (Harvey et al., 2022; Tashiro et al., 2021) as opposed to adapting the denoising schedule of an unconditional diffusion model. Note that these previous approaches never explicitly introduced a notion of time in their training procedure, as opposed to Rolling Diffusion. Specifically, Harvey et al. (2022) compare various such conditioning schemes, but do not explicitly consider a temporally adapted noise schedule.

2 Other time-series diffusion models

Apart from video, sequential diffusion models have also been applied to other time-series data, such as audio (Kong et al., 2021), text (Li et al., 2022), but also scientifically to weather data (Price et al., 2023) or fluid mechanics (Kohl et al., 2023). Lippe et al. (2023) show that incorporating a diffusion-inspired denoising procedure can help recover high frequency information that typically gets lost when using learned numerical PDE solver emulators. Finally, Wu et al. (2023) also study autoregressive models with specialized noising schedules, focusing mostly on text generation.

Experiments

We conduct experiments using data from various domains and explore several conditioning settings. In all our experiments, we use the Simple Diffusion architecture (Hoogeboom et al., 2023) with equal parameters for both standard and rolling diffusion. Specifically, we use two-dimensional spatial convolution blocks after which, in the deepest layers, we have transformer blocks that operate (i.e., attend) both spatially and temporally.

First, we run an experiment on simulated fluid dynamics from JaxCFD (Kochkov et al., 2021; Dresdner et al., 2022). Specifically, we use the Kolmogorov flow, an instance of the incompressible Navier-Stokes equations. Recently, there has been increasing interest in emulating classical numerical PDE integrators with machine learning models. Various results have shown that these have the capacity to simulate from initial conditions complex systems to high precision (e.g., Li et al. (2020)). To similar ends, generative models are of increasing interest, as they provide several benefits. First, they provide a way to directly obtain marginal distributions over a future state of a physical system, as opposed to numerically rolling out an ensemble of initial conditions. This especially has use-cases in weather or climate modeling, fluid mechanics analyses, and stochastic differential equation studies. Second, they can improve modeling high data frequencies over approaches based on mean-squared error objectives.

where fˉ\bar{{\bm{f}}} and Σ{\bm{\Sigma}} denote the mean and covariance of the frequencies, respectively. We call this metric the Fréchet Spectral Distance (FSD).

In Figure 3 we present FSD computed from the horizontal velocity fuilds of the fluid. Regarding rolling diffusion, we use the twinit(ncln)t^{\text{init}}_{w}(n_{\text{cln}}) reparameterization with ncln=2n_{\text{cln}}=2, and use twlin(ncln)t^{\text{lin}}_{w}(n_{\text{cln}}) for long rollouts. It is clear that an autoregressive MSE-based model, as typically used in the literature, is not suitable for this task. For standard diffusion, we iteratively generate W−nclnW-n_{\text{cln}} frames, after which we concatenate these to the conditioning and continue the rollout. Rolling diffusion always shifts the window by one, sampling using the process described before. We see that the rolling diffusion is able to consistently outperform the standard diffusion methods, regardless of various conditioning settings and window sizes, denoted with ‘(ncln,W−ncln)(n_{\text{cln}},W-n_{\text{cln}})’. We provide an example rollout in Figure 2, where we plot the vorticity of the velocity field. We numerically present our results in Appendix C.

2 BAIR Robot Pushing Dataset

The Berkeley AI Research (BAIR) robot pushing dataset (Ebert et al., 2017) is a standard benchmark for video prediction. It contains 44 000 videos at 64×6464\times 64 of a robot arm pushing objects around. Following previous methods, we condition in on 1 frame and predict the next 15. We evaluate, consistently with previous works, using the Frechét Video Distance (FVD) (Unterthiner et al., 2019). For FVD, we use the I3D network (Carreira & Zisserman, 2017) by comparing 100×256100\times 256 model samples against the 256 examples in the evaluation set.

Regarding rolling diffusion, we use the twinit(ncln)t^{\text{init}}_{w}(n_{\text{cln}}) reparameterization to sample the W=16W=16 (ncln=1n_{\text{cln}}=1) frames to a partially denoised state, and then use twlin(ncln)t^{\text{lin}}_{w}(n_{\text{cln}}) to rollout and complete the sampling. Note that standard diffusion samples all 15 frames at once and might be at an advantage since we do not consider autoregressive extension.

The results are shown in Table 1. We observe that both standard diffusion and rolling diffusion using the same (Simple Diffusion) architecture outperform previous methods. Additionally, we see that there is no significant difference between the standard and rolling framework in this setting. This is because the sampled sequences are, in both cases, indistinguishable from the true data Figure 4. One could say that the models are can completely solve this task, yielding no significant difference in performance. However, it is still interesting that both models are able to outperform existing methods.

3 Kinetics-600

Finally, we evaluate video prediction on the Kinetics-600 benchmark (Kay et al., 2017; Carreira et al., 2018). It contains approximately 400 000 training videos depicting 600 different activities rescaled to 64×6464\times 64. We run two experiments using this dataset. The first is a baseline experiment in a setting equal to previously published works. The next one specifically tests Rolling Diffusion’s ability to autoregressively rollout for long sequences.

We compare against previous methods using 5 input frames and 11 output frames, and show the results in Table 2 The evaluation metric is again FVD. We note that many of the current SOTA methods are two-stage, meaning that they use an autoencoder and run diffusion in latent space. While empirically compelling (Rombach et al., 2022), this makes it hard to isolate the effect of the diffusion model itself. Note that it is not always clear whether the autoencoder parameters are included in the parameter count for two-stage methods, or on what data they are pretrained. Running diffusion using the standard diffusion U-ViT architecture achieves an FVD of 3.9, which is comparable to the best two-stage methods. Rolling diffusion has a strong disadvantage in this case: (1) the baseline generates all frames at once; and (2) with a stride of 1, there is very little dynamics in the 16 frames, mostly suitable for a standard diffusion model. Still, rolling diffusion achieves a FVD of 5.2.

Rollout

Next, we compare the models’ capabilities to autoregressively rollout and show the results in Table 3. All settings use a window size of W=16W=16 frames, where set nclnn_{\text{cln}} to various settings, denoted with ‘cond-gen’. Furthermore, we train rolling diffusion with a mix of the default schedule and a rescaled schedule. The default rolling diffusion schedule twlint^{\text{lin}}_{w} is oversampled at a rate of β\beta.

We analyze two settings, one with a stride (also known as frame-skip or frame-step) of 1, rolling out for 64 steps, and another setting with a stride of 8 rolling out for 24 steps, effectively predicting ahead up to the 192192th frame. In the first setting, standard diffusion performs better, quite possibly due to the invariability of the data. We oversample the linear rolling schedule with a rate of β=0.9\beta=0.9 to account for the high number of test-time steps. Rolling diffusion consistently wins in the second setting, which is much more dynamic. Note also that single-frame diffusion significantly underperforms here and that larger block autoregression is favorable. See an example rollout in Figure 5. From Appendix F, we get an indication of the effect of the oversampling rate on the performance of rolling diffusion. In this case, slightly oversampling twlint^{\text{lin}}_{w} yields the best result.

Conclusion

We presented Rolling Diffusion Models, a new DDPM framework that progressively noises (and denoises) data through time. Validating our method on video and fluid mechanics data, we observed that rolling diffusion’s natural inductive bias gets most effectively exploited when the data is highly dynamic. This allows for exciting future directions in, e.g., video, audio, and weather or climate modeling.

Broader Impact

Sequential generative models, including diffusion models, have a significant societal impact with applications in video generation and scientific research by enabling fast, highly detailed sampling. While they offer the upside of creating more accurate and compelling synthesis in fields ranging from climate modeling to medical imaging, there are notable downsides regarding content authenticity and originality of digital media.

References

Appendix A Hyperparameters

In this section we denote the hyperparameters for the different experiments. Throughout the experiments we use U-ViTs which are essentially U-Nets with MLP Blocks instead of convolutional layers when self-attention is used in a block. In the PDE experiments are relatively small architecture is used. For BAIR we used a larger architecture, increasing both the channel count and the number of blocks. For K600 we used even larger architectures, because this dataset turned out to be the most difficult to fit. For all the specifications see Table 4, 5 and 6.

Appendix B Simulation Details

In this section we present the hyperparameters to generate the Kolmogorov Flow simulation data, see Table 7. It is important to note that to introduce uncertainty into an otherwise deterministic system, we vary the viscosity and density parameters, which must then be inferred from the data. This would also make a standard solver very difficult to use in such a setting. Additionally, due to the chaotic nature of the system, it is not deterministically predictable up to arbitrary precision.

We use the “simple turbulence forcing” option, which combines a driving force with a damping term such that we simulate turbulence in two dimensions.

Appendix C Additional Results

We show in Table 8 the MSE and FSE errors at various time-steps. Note that the MSE model is always optimal in terms of MSE loss, which is as expected. However, in terms of matching the frequency distribution, as measured by FSD, standard diffusion, and in particular rolling diffusion are optimal.

Appendix D Rolling Diffusion Objective

where cc is a data entropy term. The prior and reconstruction loss terms are typically negligible. LD\mathcal{L}_{D} is the diffusion loss, which is defined as

Further, when T→∞T\rightarrow\infty, taking care that s→ts\to t we get continuous analog of Equation 30 (Kingma et al., 2021):

with w(λt)=1w(\lambda_{t})=1, where λt=log⁡SNR⁡(t)\lambda_{t}=\log\operatorname{SNR}(t). The weighting function w(λt)=1w(\lambda_{t})=1 is often changed to improve image quality, for instance by being the inverse −1/dλtdt-1/\frac{d\lambda_{t}}{dt} so that the objective is simply constant over ϵ{\bm{{\epsilon}}}-loss.

Rolling Diffusion Objective

In rolling diffusion, the signal-to-noise ratio is kept fixed but one has to account for the local time reparameterization. Recall that tk:=tk(t)t_{k}:=t_{k}(t) denotes the local time reparameterization. Concisely we can say SNR⁡k(t):=SNR⁡(tk)\operatorname{SNR}_{k}(t):=\operatorname{SNR}(t_{k}). The rolling continuous time objective is:

where λ(t):=log⁡SNR⁡(t)\lambda(t):=\log\operatorname{SNR}(t).

Recall the frame categorization of the main paper, i.e.,

we can see that tk′(t)=0t^{\prime}_{k}(t)=0 for k∈clean(t−dt,t)k\in\text{clean}(t-dt,t) and k∈noise(t−dt,t)k\in\text{noise}(t-dt,t) and thus the objective only has non-zero loss over the window:

This derivation shows a subtle difference between the global perspective (as used above) and the local perspective. If variational bounds are used, they are equivalent and the reparametrization of the time derivative takes into account the edge conditions such as fully noisy frames and fully determined frames. However, constant ϵ{\bm{{\epsilon}}} or vv losses on the global perspective do differ from the local perspective, as they do not vanish to zero at t=0t=0 and t=1t=1.

Empirically, ϵ{\bm{{\epsilon}}}-loss on the window achieved good performance against vv-loss. Our rolling diffusion objective could be viewed in two ways: As a constant ϵ{\bm{{\epsilon}}} defined on the window, or a globally defined constant ϵ{\bm{{\epsilon}}}-loss that masks out everything but the window, essentially taking into account the derivative of the time reparametrization outside of the window.

Appendix E Algorithms

We present algorithmic outlines for training of rolling diffusion as well as sampling at the boundary. The algorithm for autoregressive rollout was presented in the main text.

Appendix F Hyperparameter Search for β𝛽\beta

Appendix G Rescaled Noise Schedule

For our Kinetics-600 experiments, we used a different noise schedule which can sample from complete noise towards a “rolling state”, i.e., at diffusion times (1W,2W,…WW)(\frac{1}{W},\frac{2}{W},\ldots\frac{W}{W}). From there, we can roll out generation using, e.g., the linear rolling sampling schedule twlint_{w}^{\text{lin}}. The reason is that we hypothesized that the noise schedule twinitt_{w}^{\text{init}} uses a clip⁡\operatorname{clip} operation, which means that will be sampling in clean(s,t)\text{clean}(s,t), which is redundant as outlined in the main paper and Appendix D.

Another schedule that starts from complete noise and ends at the rolling state is the following:

Where we clearly have at t=1Wt=\frac{1}{W} that the local times are (1W,2W,…,WW)\left(\frac{1}{W},\frac{2}{W},\ldots,\frac{W}{W}\right), which is what we need. Note that this schedule is not directly proportional to tt.

Appendix H Example