Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion

Boyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz, Russ Tedrake, Vincent Sitzmann

Introduction

Probabilistic sequence modeling plays a crucial role in diverse machine learning applications including natural language processing , video prediction and decision making . Next-token prediction models in particular have a number of desirable properties. They enable the generation of sequences with varying length (generating only a single token or an “infinite” number of tokens via auto-regressive sampling), can be conditioned on varying amounts of history , support efficient tree search, and can be used for online feedback control .

Current next-token prediction models are trained via teacher forcing , where the model predicts the immediate next token based on a ground truth history of previous tokens. This results in two limitations: (1) there is no mechanism by which one can guide the sampling of a sequence to minimize a certain objective, and (2) current next-token models easily become unstable on continuous data. For example, when attempting to auto-regressively generate a video (as opposed to text or vector-quantized latents ) past the training horizon, slight errors in frame-to-frame predictions accumulate and the model diverges.

Full-sequence diffusion seemingly offers a solution. Commonly used in video generation and long-horizon planning, one directly models the joint distribution of a fixed number of tokens by diffusing their concatenation , where the noise level is identical across all tokens. They offer diffusion guidance to guide sampling to a desirable sequence, invaluable in decision-making (planning) applications . They further excel at generating continuous signals such as video . However, full-sequence diffusion is universally parameterized via non-causal, unmasked architectures. In addition to restricting sampling to full sequences, as opposed to variable length generation, we show that this limits the possibilities for both guidance and subsequence generation (Figure 1). Further, we demonstrate that a naive attempt at combining the best of both worlds by training a next-token prediction model for full-sequence diffusion leads to poor generations, intuitively because it does not model the fact that small uncertainty in an early token necessitates high uncertainty in a later one.

In this paper, we introduce Diffusion Forcing (DF), a training and sampling paradigm where each token is associated with a random, independent noise level, and where tokens can be denoised according to arbitrary, independent, per-token schedules through a shared next-or-next-few-token prediction model. Our approach is motivated by the observation that noising tokens is a form of partial masking—zero noise means a token is unmasked, and complete noise fully masks out a token. Thus, DF forces the model to learn to “unmask” any collection of variably noised tokens (Figure 2). Simultaneously, by parameterizing predictions as a composition of next-token prediction models, our system can flexibly generate varying length sequences as well as compositionally generalize to new trajectories (Figure 1).

We implement DF for sequence generation as Causal Diffusion Forcing (CDF), in which future tokens depend on past ones via a causal architecture. We train the model to denoise all tokens of a sequence at once, with an independent noise level per token. During sampling, CDF gradually denoises a sequence of Gaussian noise frames into clean samples where different frames may have different noise levels at each denoising step. Like next-token prediction models, CDF can generate variable-length sequences; unlike next-token prediction, it does so stabily from the immediate next token to thousands of tokens in the future – even for continuous tokens. Moreover, like full-sequence diffusion it accepts guidance towards high-reward generations. Synergistically leveraging causality, flexible horizon, and variable noise schedules, CDF enables a new capability, Monte Carlo Tree Guidance (MCTG), that dramatically improves the sampling of high-reward generations compared to non-causal full-sequence diffusion models. Fig. 1 overviews these capabilities.

In summary, our contributions are: (1) We propose Diffusion Forcing, a new probabilistic sequence model that has the flexibility of next-token prediction models while being able to perform long-horizon guidance like full-sequence diffusion models. (2) Taking advantage of Diffusion Forcing’s unique capabilities, we introduce a novel decision-making framework that allows us to use Diffusion Forcing as simultaneously a policy () and as a planner (). (3) We formally prove that, under appropriate conditions, optimizing our proposed training objective maximizes a lower bound on the likelihood of the joint distribution of all sub-sequences observed at training time. (4) We empirically evaluate CDF across diverse domains such as video generation, model-based planning, visual imitation learning, and time series prediction, and demonstrate CDF’s unique capabilities, such as stabilizing long-rollout autoregressive video generation, composing sub-sequences of those observed at training time with user-determined memory horizon, Monte Carlo Tree Guidance, and more.

Related Work and Preliminaries

We discuss related work and preliminaries for our core application, sequence generative modeling; see Appendix C for further literature review.

Given a Hidden Markov Model (HMM) defined by latent states zt\mathbf{z}_{t} and observations xt\mathbf{x}_{t}, a Bayes filter is a probabilistic method for estimating latent states recursively over time from incoming observations. A prior model p(zt+1∣zt)p(\mathbf{z}_{t+1}|\mathbf{z}_{t}) infers a belief over the next state given only the current state, and an observation model infers a belief over the next observation given the current latent state p(xt∣zt)p(\mathbf{x}_{t}|\mathbf{z}_{t}). When a new observation is made, a posterior model p(zt+1∣zt,xt+1)p(\mathbf{z}_{t+1}|\mathbf{z}_{t},\mathbf{x}_{t+1}) provides an updated estimation of the next latent state zt+1\mathbf{z}_{t+1}. When trained end-to-end with neural networks , latent states are not an estimate of any physical quantity, but a sufficiently expressive latent that summarizes past observations for predicting future observations (xt′)t′>t(\mathbf{x}_{t^{\prime}})_{t^{\prime}>t} in the sequence.

Diffusion models have proven to be highly expressive and reliable generative models. We review their essentials here. Let q(x)q(\mathbf{x}) denote a data distribution of interest, and let x0≡x∼q\mathbf{x}^{0}\equiv\mathbf{x}\sim q. We consider a forward diffusion process that gradually adds Gaussian noise to a data point over a series of time steps. This process is modeled as a Markov chain, where the data at each step kk is noised incrementally:

where N\mathcal{N} is the normal distribution and βk\beta_{k} is the variance of the noise added at each step controlled by a schedule {βk∈(0,1)}k=1K\{\beta_{k}\in(0,1)\}_{k=1}^{K}. The process continues until the data is converted into pure noise at xK\mathbf{x}^{K}. The reverse process is also a Markov chain and attempts to recreate the original data from the noise with a parameterized model pθp_{\theta}:

where the mean μ\bm{\mu} is model with a neural network, and where it is shown that one can set the covariance to the identity scaled by a fixed constant γk\gamma_{k} depending on kk. Adopting the standard exposition, we reparametrize the mean μ\bm{\mu} in terms of noise prediction ϵ=(1−αˉt)−1xtkt−αˉtμ\bm{\epsilon}=(\sqrt{1-\bar{\alpha}_{t}})^{-1}\mathbf{x}_{t}^{k_{t}}-\sqrt{\bar{\alpha}_{t}}\bm{\mu}. This leads to the following least squares objective:

where xk=αtˉx0+1−αtˉϵk\mathbf{x}^{k}=\sqrt{\bar{\alpha_{t}}}\mathbf{x}^{0}+\sqrt{1-\bar{\alpha_{t}}}\epsilon^{k} and ϵk∼N(0,I)\bm{\epsilon}^{k}\sim\mathcal{N}(0,\mathbf{I}) . One can then sample from this model via Langevin dynamics xk−1←1αk(xtk−1−αk1−αˉkϵθ(xtk,k))+σkw)\mathbf{x}^{k-1}\leftarrow\frac{1}{\sqrt{\alpha_{k}}}(\mathbf{x}_{t}^{k}-\frac{1-\alpha_{k}}{\sqrt{1-\bar{\alpha}_{k}}}\bm{\epsilon}_{\theta}(\mathbf{x}_{t}^{k},k))+\sigma_{k}\mathbf{w}) .

Guidance allows biasing diffusion generation towards desirable predictions at sampling time. We focus on classifier guidance : given a classifier c(y∣xk)c(y|\mathbf{x}^{k}) of some desired yy (e.g. class or success indicator), one modifies the Langevin sampling gradient ϵθ(xk,k)\bm{\epsilon}_{\theta}(\mathbf{x}^{k},k) to be ϵθ(xk,k)−1−αˉk∇xklog⁡c(y∣xk)\bm{\epsilon}_{\theta}(\mathbf{x}^{k},k)-\sqrt{1-\bar{\alpha}_{k}}\nabla_{x^{k}}\log c(y|\mathbf{x}^{k}). This allows sampling from the joint distribution of x\mathbf{x} and class label yy without the need to train a conditional model. Other energies such as a least-squares objective comparing the model output to a desirable ground-truth have been explored in applications such as decision making .

Next-token prediction models are sequence models that predict the next frame xt+1\mathbf{x}_{t+1} given past frames x1:t\mathbf{x}_{1:t}. At training time, one feeds a neural network with x1:t\mathbf{x}_{1:t} and minimizes ∣∣x^−x∣∣2||\hat{\mathbf{x}}-\mathbf{x}||^{2} for continuous data or a cross-entropy loss for discrete data . At sampling time, one samples the next frame x^t+1\hat{\mathbf{x}}_{t+1} following p(xt+1∣x1:t)p(\mathbf{x}_{t+1}|\mathbf{x}_{1:t}). If one treats x^t+1\hat{\mathbf{x}}_{t+1} as xt+1\mathbf{x}_{t+1}, one can use the same model to predict xt+2\mathbf{x}_{t+2} and repeat until a full sequence is sampled. Unlike full-sequence diffusion models, next-token models do not accept multi-step guidance, as prior frames must be fully determined to sample future frames.

Diffusion has been widely used in sequence modeling. use full-sequence diffusion models to achieve controllable text generation via guidance, such as generating text following specified parts of speech. trains full-sequence diffusion models to synthesize short videos and uses a sliding window to roll out longer conditioned on previously generated frames. uses full-sequence diffusion models as planners in offline reinforcement learning. This is achieved by training on a dataset of interaction trajectories with the environment and using classifier guidance at sampling time to sample trajectories with high rewards towards a chosen goal. modifies auto-regressive models to denoise the next token conditioned on previous tokens. It trains with teacher forcing and samples next-token auto-regressively for time series data. Most similar to our work is AR-Diffusion , which trains full-sequence text diffusion with a causal architecture with linearly dependent noise level along the time axis. We provide a detailed comparision between this approach and ours in Appendix C.

Diffusion Forcing

Recall that masking is the practice of occluding a subset of data, such as patches of an image or timesteps in a sequence , and training a model to recover unmasked portions. Without loss of generality, we can view any collection of tokens, sequential or not, as an ordered set indexed by tt. Training next-token prediction with teacher forcing can then be interpreted as masking each token xt\mathbf{x}_{t} at time tt and making predictions from the past x1:t−1\mathbf{x}_{1:t-1}. Restricted to sequences, we refer to all these practices as masking along the time axis. We can also view full-sequence forward diffusion, i.e., gradually adding noise to the data x1:T0≡x1:T\mathbf{x}^{0}_{1:T}\equiv\mathbf{x}_{1:T}, as a form of partial masking, which we refer to as masking along the noise axis. Indeed, after KK steps of noising, x1:TK\mathbf{x}^{K}_{1:T} is (approximately) pure white noise without information about the original data.

Diffusion Forcing (DF) is a framework for training and sampling arbitrary sequence lengths of noisy tokens (xtkt)1≤t≤T(\mathbf{x}_{t}^{k_{t}})_{1\leq t\leq T}, where critically, the noise level ktk_{t} of each token can vary by time step. In this paper, we focus on time series data, and thus instantiate Diffusion Forcing with causal architectures (where xtkt\mathbf{x}_{t}^{k_{t}} depends only on past noisy tokens), which we call Causal Diffusion Forcing (CDF). For simplicity, we focus on a minimal implementation with a vanilla Recurrent Neural Network (RNN) .

The RNN with weights θ\theta maintains latents zt\mathbf{z}_{t} capturing the influence of past tokens, and these evolve via dynamics zt∼pθ(zt∣zt−1,xtkt,kt)\mathbf{z}_{t}\sim p_{\theta}(\mathbf{z}_{t}|\mathbf{z}_{t-1},\mathbf{x}_{t}^{k_{t}},k_{t}) with a recurrent layer. When an incoming noisy observation xtkt\mathbf{x}_{t}^{k_{t}} is made, the hidden state is updated in a Markovian fashion zt∼pθ(zt∣zt−1,xtkt,kt)\mathbf{z}_{t}\sim p_{\theta}(\mathbf{z}_{t}|\mathbf{z}_{t-1},\mathbf{x}_{t}^{k_{t}},k_{t})We implement zt=pθ(zt∣zt−1,xtkt,kt)\mathbf{z}_{t}=p_{\theta}(\mathbf{z}_{t}|\mathbf{z}_{t-1},\mathbf{x}_{t}^{k_{t}},k_{t}) to be deterministic, with zt\mathbf{z}_{t} representing a distribution over beliefs rather than a sample from it. This allows training by backpropogating through the latent dynamics in Eq.(3.1).. When kt=0k_{t}=0, this is the posterior update in Bayes filtering; whereas when kt=Kk_{t}=K (and xtK\mathbf{x}_{t}^{K} is pure noise and thus uninformative), this is equivalent to modeling the “prior distribution” pθ(zt∣zt−1)p_{\theta}(\mathbf{z}_{t}\mid\mathbf{z}_{t-1}) in Bayes filtering. Given latent zt\mathbf{z}_{t}, an observation model pθ(xt0∣zt)p_{\theta}(\mathbf{x}_{t}^{0}|\mathbf{z}_{t}) predicts xt\mathbf{x}_{t}; this unit has the same input-output behavior as a standard conditional diffusion model, using a conditioning variable zt−1\mathbf{z}_{t-1} and a noisy token xtkt\mathbf{x}_{t}^{k_{t}} as input to predict the unnoised xt=xt0\mathbf{x}_{t}=\mathbf{x}_{t}^{0} and thus, indirectly, the noise ϵkt\epsilon^{k_{t}} via affine reparametrization . We can thus directly train (Causal) Diffusion Forcing with the conventional diffusion training objective. We parameterize the aforementioned unit in terms of noise prediction ϵθ(zt−1,xtkt,kt)\bm{\epsilon}_{\theta}(\mathbf{z}_{t-1},\mathbf{x}_{t}^{k_{t}},k_{t}). We then find parameters θ\theta by minimizing the loss

where we sample k1:Tk_{1:T} uniformly from [K]T[K]^{T}, x1:T\mathbf{x}_{1:T} from our training data, and ϵt∼N(0,σkt2I)\epsilon_{t}\sim\mathcal{N}(0,\sigma^{2}_{k_{t}}I) in accordance with the forward diffusion process (see Algorithm 1 for pseudocode). Importantly, the loss (3.1) captures essential elements of Bayesian filtering and conditional diffusion. In Appendix D.3, we further re-derive common techniques in diffusion model training for Diffusion Forcing. Finally, we prove the validity of this objective stated informally in the following Theorem 3.1 in Appendix A.

The Diffusion Forcing training procedure (Algorithm 1) optimizes a reweighting of an Evidence Lower Bound (ELBO) on the expected log-likelihoods ln⁡pθ((xtkt)1≤t≤T)\ln p_{\bm{\theta}}((\mathbf{x}_{t}^{k_{t}})_{1\leq t\leq T}), where the expectation is averaged over noise levels k1:T∼[K]Tk_{1:T}\sim[K]^{T} and xtkt\mathbf{x}_{t}^{k_{t}} noised according to the forward process. Moreover, under appropriate conditions, optimizing (3.1) also maximizes a lower bound on the likelihood for all sequences of noise levels, simultaneously.

We remark that a special case of ‘all sequences of noise levels’ are those for which either kt=0k_{t}=0 or kt=Kk_{t}=K; thus, one can mask out any prior token and DF will learn to sample from the correct conditional distribution, modeling the distribution of all possible sub-sequences of the training set.

1 Diffusion Forcing Sampling and Resulting Capabilities

Sampling is depicted in Algorithm 2 and is defined by prescribing a noise schedule on a 2D M×TM\times T grid K∈[K]M×T\mathcal{K}\in[K]^{M\times T}; columns correspond to time step tt and rows indexed by mm determine noise-level. Km,t\mathcal{K}_{m,t} represents the desired noise level of the time-step tt token for row mm. To generate a whole sequence of length TT, initialize the tokens x1:T\mathbf{x}_{1:T} to be white noise, corresponding to noise level k=Kk=K. We iterate down the grid row-by-row, denoising left-to-right across columns to the noise levels prescribed by K\mathcal{K}. By the last row m=0m=0, the tokens are clean, i.e. their noise level is K0,t≡0\mathcal{K}_{0,t}\equiv 0.

Section D.5 discusses corner cases of this scheme; the hyperparameters (αk,αˉk,σk)(\alpha_{k},\bar{\alpha}_{k},\sigma_{k}) are set to their standard values . We now explain new capabilities this sampling paradigm has to offer.

For high-dimensional, continuous sequences such as video, auto-regressive architectures are known to diverge, especially when sampling past the training horizon. In contrast, Diffusion Forcing can stably roll out long sequences even beyond the training sequence length by updating the latents using the previous latent associated with slightly “noisy tokens” for some small noise level 0<k≪K0<k\ll K. Our experiments (Sec. 4.1) illustrates the resulting marked improvements in long-horizon generation capabilities; App. B.2 provides further intuition.

Beginning from a sequence of white noise tokens [x1K,x2K,x3K]⊤[\mathbf{x}^{K}_{1},\mathbf{x}^{K}_{2},\mathbf{x}^{K}_{3}]^{\top}, we may denoise the first token fully and the second token partially, yielding [x10,x2K/2,x3K]⊤[\mathbf{x}^{0}_{1},\mathbf{x}^{K/2}_{2},\mathbf{x}^{K}_{3}]^{\top}, then [x10,x20,xK/23]⊤[\mathbf{x}^{0}_{1},\mathbf{x}^{0}_{2},\mathbf{x}^{K/2_{3}}]^{\top}, and finally denoising all tokens fully to [x10,x20,x30]⊤[\mathbf{x}^{0}_{1},\mathbf{x}^{0}_{2},\mathbf{x}^{0}_{3}]^{\top}. Interpreting the noise level as uncertainty, this “zig-zag” sampling scheme intuitively encodes the immediate future as more certain than the far future. Sec. 3.2 describes how this leads to more effective sequence guidance.

In Line 10 of Algorithm 2, one may add guidance to the partially diffused trajectory x1:T\mathbf{x}_{1:T} as in Sec. 2. Due to the dependency of future tokens on the past, guidance gradients from future tokens can propagate backwards in time. The unique advantage of Diffusion Forcing is that, because we can diffuse future tokens without fully diffusing the past, the gradient guides the sampling of past tokens, thereby achieving long-horizon guidance while respecting causality. We elaborate on implementation details in Appendix B.1. As we show in Section 4.2, planning in this manner significantly outperforms guided full-sequence diffusion models.

2 Diffusion Forcing for Flexible Sequential Decision Making

Diffusion Forcing (a) can be deployed on tasks of variable horizon, because each new action is selected sequentially, and (b) its lookahead window HH can be shortened to lower latency (using Diffusion Forcing as a policy), or lengthened to perform long-horizon planning (via guidance described below), without re-training or modifications of the architecture. Note that (a) is not possible for full-sequence diffusion models like Diffuser with full-trajectory generation horizons, whereas diffusion policies need fixed, small lookahead sizes, precluding (b).

As detailed in Appendix B.1, Diffusion Forcing can plan via guidance using any reward (in place of log⁡c\log c) specified over future steps: this includes dense per-time step rewards on the entire trajectory trajectory ∑t=1Trt\sum_{t=1}^{T}\mathbf{r}_{t}, dense rewards on a future lookahead ∑t′=tt+Hrt\sum_{t^{\prime}=t}^{t+H}\mathbf{r}_{t}, and sparse rewards indicating goal completion −∥oT−g∥2-\|\mathbf{o}_{T}-\mathbf{g}\|^{2}. Per-time step policies cannot take advantage of this latter, longer horizon guidance.

Causal Diffusion Forcing allows us to influence the generation of a token xtk\mathbf{x}_{t}^{k} by guidance on the whole distribution of future xt+1:T\mathbf{x}_{t+1:T}. Instead of drawing a single trajectory sample to calculate this guidance gradient, we can draw multiple samples and average their guidance gradients. We call this Monte Carlo Tree Guidance, where “tree” comes from the fact that the the denoising of the current xtk\mathbf{x}_{t}^{k} is influenced by gradients through many paths through the future. In the spirit of so-called shooting methods like MPPI , xtk\mathbf{x}_{t}^{k} is then guided by the expected reward over the distribution of all future outcomes instead of one particular outcome. The effect of MCTG is enhanced when combined with sampling schedules that keep the noise level of future tokens high when denoising immediate next tokens (e.g. the zig-zag schedule described in Sec. 3.1), accounting for greater uncertainty farther into the future. Appendix B.3 further justifies the significance of MCTG, and why Diffusion Forcing uniquely takes advantage of it.

Experiments

We extensively evaluate Diffusion Forcing’s merits as a generative sequence model across diverse applications in video and time series prediction, planning, and imitation learning. Please find dataset and reproducibility details in the Appendix, as well as video results on the project website.

We train a convolutional RNN implementation of Causal Diffusion Forcing for video generative modeling on videos of Minecraft gameplay and DMLab navigation . At sampling time, we perform auto-regressive rollout with stabilization proposed in Sec. 3.1. We consider two baselines, both leveraging the same exact RNN architecture: a next-frame diffusion baseline trained with teacher forcing as well as a causal full-sequence diffusion model. Figure 3 displays qualitative results of roll-outs generated by Diffusion Forcing and baselines starting from unseen frames for both datasets. While Diffusion Forcing succeeds at stably rolling out even far beyond its training horizon (e.g. 10001000 frames), teacher forcing and full-sequence diffusion baselines diverge quickly. Further, within the training horizon, we observe that full-sequence diffusion suffers from frame-to-frame discontinuity where video sequences jump dramatically, while Diffusion Forcing roll-outs show ego-motion through a consistent 3D environment. This highlights the ability of Diffusion Forcing to stabilize rollouts of high-dimensional sequences without compounding errors.

Decision-making uniquely benefits from Diffusion Forcing’s capabilities. We evaluate our proposed decision-making framework in a standard offline RL benchmark, D4RL . Specifically, we benchmark Diffusion Forcing on a set of 2D maze environments with sparse reward. An agent is tasked with reaching a designated goal position starting from a random starting position. In Appendix 16 we provide a detailed description of the environment. The benchmark provides a dataset of random walks through mazes (thus stochastic). We train one model per maze.

We benchmark the proposed decision-making framework 3.2 with state-of-the-art offline RL methods and the recently introduced Diffuser , a diffusion planning framework. See Fig. 4.1 for qualitative and quantitative reuslts: DF outperforms Diffuser and all baselines across all 66 environments.

The typical goal for an RL problem is to find actions that maximize the expected future rewards, which we achieve through MCTG. Full-sequence diffusion models such as Diffuser do not support sampling to maximize expected reward, as we formally derive in Section B.3. To understand MCTG’s importance, we ablate it in Section 4.1. Removing MCTG guidance degrades our performance, though Diffusion Forcing remains competitive even then.

Unlike pure generative modeling, sequential decision-making takes actions and receives feedback. Due to compounding uncertainty, the immediate next actions are more important than those in the far future. Though Diffuser and subsequent models are trained to generate sequences of action-reward-state tuples [at,rt,ot][\mathbf{a}_{t},\mathbf{r}_{t},\mathbf{o}_{t}], directly executing the actions will lead to a trajectory that deviates significantly from the generated states. In other words, the generated states and actions are not causally consistent with each other. To address this shortcoming, Diffuser’s implementation ignores the generated actions and instead relies on a hand-crafted PD controller to infer actions from generated states. In Table 4.1, we see that Diffuser’s performance drops dramatically when directly executing generated actions. In contrast, Diffusion Forcing’s raw action generations are self-consistent, outperforming even actions selected by combining Diffuser’s state predictions with a handcrafted PD controller.

3 Controllable Sequential Compositional Generation

Many RL tasks have a fixed horizon, requiring the planning horizon to shrink as an agent makes progress in the task. Diffusion Forcing accomplishes this by design, while full-sequence models like Diffuser perform poorly even with tweaks, as we explain in Appendix B.4.

We demonstrate that by only modifying the sampling scheme, we can flexibly compose sub-sequences of sequences observed at training time. We consider a dataset of trajectories on a 2D, square plane, where all trajectories start from one corner and end up in the opposite corner, forming a cross shape. As shown in Fig. 1, when no compositional behavior is desired, one can let DF keep full memory, replicating the cross-shaped distribution. When one desires compositionality, one can let the model generate shorter plans without memory using MPC, leading to stitching of the cross’s sub-trajectories, forming a V-shaped trajectory. Due to limited space, we defer the result to Appendix E.2.

Finally, we illustrate that Diffusion Forcing (DF) opens up new opportunities in visuomotor control of real-world robots. Imitation learning is a popular technique in robotic manipulation where one learns an observation-to-action mapping from expert demonstrations. However, the lack of memory often prevents imitation learning from accomplishing long-horizon tasks. DF not only alleviates this shortcoming but also provides a way to make imitation learning robust.

We collect a dataset of videos and actions by teleoperating a Franka robot. In the chosen task, one needs to swap the position of an apple and an orange, using a third slot. See Fig. 4 for an illustration. The initial positions of the fruits are randomized such that there are two possible goal states. As illustrated in Fig. 4, when one fruit is in the third slot, the desired outcome cannot be inferred from the current observation—a policy must remember the initial configuration to determine which fruit to move. In contrast to common behavior cloning methods, DF naturally incorporates memory in its latent state. We found that DF achieves 80%80\% success rate while diffusion policy , a state-of-the-art imitation learning algorithm without memory, fails.

Because it incorporates principles from Bayes filtering, Diffusion Forcing can perform imitation learning while being robust to noisy or missing observations. We demonstrate this by adding visual distractions and even fully occluding the camera during execution. DF allows us to easily indicate these observations as “noisy” by using k>0k>0, in which case DF relies heavily on its prior model to predict actions. Consequently, the succes rate is only lowered by 4%4\% to 76%76\%. In contrast, a next-frame diffusion model baseline attains a success rate of 48%48\%: it must treat perturbed observations as ground truth and suffers out-of-distribution error.

5 Time Series Forecasting: Diffusion Forcing is a Good General-purpose Sequence Model

Finally in parallel to generating actions, Fig. 4 illustrates that Diffusion Forcing is capable of generating a video of the robot performing the task given only an initial frame, unifying diffusion policy / imitation learning and video generative modeling and paving the way to pre-training on unlabeled video.

In Appendix E, we show that DF is competitive with prior diffusion and transformer-based work on multivariate time series forecasting, following the experimental setup of .

Our current causal implementation is based on a small RNN, and applications to higher-resolution video or more complex distributions likely require large transformer models. We do not investigate the scaling behavior of Diffusion Forcing to internet-scale datasets and tasks.

In this paper, we introduced Diffusion Forcing, a new training paradigm where a model is trained to denoise sets of tokens with independent, per-token noise levels. Applied to time series data, we show how a next-token prediction model trained with Diffusion Forcing combines benefits of both next-token models and full-sequence diffusion models. We introduced new sampling and guidance schemes that lead to dramatic performance gains when applied to tasks in sequential decision making. Future work may investigate the application of Diffusion Forcing to domains other than time series generative modeling, and scale up Diffusion Forcing to larger datasets.

References

This work was supported by the National Science Foundation under Grant No. 2211259, by the Singapore DSTA under DST00OECI20300823 (3D Self-Supervised Learning for Label-Efficient Vision), by the Intelligence Advanced Research Projects Activity (IARPA) via Department of Interior/ Interior Business Center (DOI/IBC) under 140D0423C0075, and by the Amazon Science Hub. Special thanks to Kiwhan Song (MIT) for 3D-Unet and transformer reimplementation of the paper after initial release.

Appendix A Theoretical Justification

In this section, we provide theoretical justification for the train of Diffusion Forcing. The main contributions can be summarized as follows:

We show that our training methods optimizes a reweighting of Evidence Lower Bound (ELBO) on the average log-likelihood of our data. We first establish this in full abstraction (Theorem A.1), and then specialize to the form of Gaussian diffusion (Corollary A.2). We show that the resulting terms decouple in such a fashion that, in the limit of a fully expressive latent and model, makes the reweighting terms immaterial.

We show that the expected likelihood over any distribution over sequences of noise levels can be lower bounded by a sum over nonnegative terms which, when reweighted, correspond to the terms optimized in the Diffusion Forcing training objective maximizes. Thus, for a fully expressive network that can drive all terms to their minimal value, Diffusion Forcing optimizes a valid surrogate of the likelihood for likelihood of all sequences of noise levels simultaneously.

We begin by stating an ELBO for general Markov forward processes q(⋅)q(\cdot), and generative models pθ(⋅)p_{\bm{\theta}}(\cdot), and then specialize to Gaussian diffusion, thereby recovering our loss. We denote our Markov forward process q(⋅)q(\cdot) as

We assume that pθp_{\bm{\theta}} satisfies the Markov property that

that is, the latent codes zt−1\mathbf{z}_{t-1} is a sufficient statistic for xkt\mathbf{x}^{k_{t}} given the history. We say that pθp_{\bm{\theta}} has deterministic latents if pθ(zt∣z1:t−1,(xsks)1≤s<t,xtkt)p_{\bm{\theta}}(\mathbf{z}_{t}\mid\mathbf{z}_{1:t-1},(\mathbf{x}_{s}^{k_{s}})_{1\leq s<t},\mathbf{x}_{t}^{k_{t}}) is a Dirac delta.

In order for pθp_{\bm{\theta}} to have deterministic latents and correspond to a valid probability distribution, we need to view the latents zt\mathbf{z}_{t} not as individual variables, but as a collections of variables zt(k1:t)\mathbf{z}_{t}(k_{1:t}) indexed by t∈[T]t\in[T] and and the history of noise levels k1:t∈{0,1,…,K}tk_{1:t}\in\{0,1,\dots,K\}^{t}.. In this case, simply setting zt(k1:t)=(k1:t,(xsks)1≤s≤t\mathbf{z}_{t}(k_{1:t})=(k_{1:t},(\mathbf{x}_{s}^{k_{s}})_{1\leq s\leq t} tautologically produces deterministic latents. The reason for indexing zt(k1:t)\mathbf{z}_{t}(k_{1:t}) with k1:tk_{1:t} then arises because, otherwise, pθ(zt∣((xsks)1≤s≤t,(xsks′)1≤s≤t)p_{\bm{\theta}}(\mathbf{z}_{t}\mid((\mathbf{x}_{s}^{k_{s}})_{1\leq s\leq t},(\mathbf{x}_{s}^{k_{s}^{\prime}})_{1\leq s\leq t}) would be ill-defined unless ks=ks′k_{s}=k_{s}^{\prime} for all 1≤s≤t1\leq s\leq t, and thus, pθp_{\bm{\theta}} would not correspond to a joint probability measure. The exposition and theorem that follows allows zt(k1:t)\mathbf{z}_{t}(k_{1:t}) to be indexed on past noise levels k1:tk_{1:t}, but suppresses dependence on k1:tk_{1:t} to avoid notational confusion.

We can now state our main theorem, which provides an evidence lower bound (ELBO) on the expected log likelihood of partially-noised sequences (xtkt)1≤t≤T)(\mathbf{x}_{t}^{k_{t}})_{1\leq t\leq T}), under uniform levels ktk_{t} and xtkt\mathbf{x}_{t}^{k_{t}} obtained by noising according to q(⋅)q(\cdot) as in (A.1). Notice that this formulation does not require an explicit for of q(⋅)q(\cdot) or pθp_{\bm{\theta}}, but we will specialize to Gaussian diffusion in the sequel.

Fix x1:T0\mathbf{x}_{1:T}^{0}. Define the expectation over the forward process with random noise level k1:Tk_{1:T} as

and the expectation over the latents under pθ(⋅)p_{\bm{\theta}}(\cdot) conditioned on k1:T,(xskt)1≤t≤Tk_{1:T},(\mathbf{x}_{s}^{k_{t}})_{1\leq t\leq T} as

Then, as long as pθp_{\bm{\theta}} satisfies the Markov property,

where C(x1:T0)C(\mathbf{x}_{1:T}^{0}) is a constant depending only on x1:T0\mathbf{x}_{1:T}^{0} (the unnoised data). Moreover, if the latents are deterministic (i.e. pθ(zt∣zt−1,xtkt)p_{\bm{\theta}}(\mathbf{z}_{t}\mid\mathbf{z}_{t-1},\mathbf{x}_{t}^{k_{t}}) is a Dirac distribution), then the inequality holds with inequality if and only if q(xtkt+1:T∣xtkt)≡pθ(xtkt+1:T∣xtkt,zt−1)q(\mathbf{x}_{t}^{k_{t}+1:T}\mid\mathbf{x}_{t}^{k_{t}})\equiv p_{\bm{\theta}}(\mathbf{x}_{t}^{k_{t}+1:T}\mid\mathbf{x}_{t}^{k_{t}},\mathbf{z}_{t-1}), i.e. the variational approximation is exact.

The proof of the amove theorem is given in Section A.2. Remarkably, it involves only two inequalities! The first holds with equality under deterministic latents and the second holds if and only if variational approximation is exact: q(xtkt+1:T∣xtkt)≡pθ(xtkt+1:T∣xtkt,zt−1)q(\mathbf{x}_{t}^{k_{t}+1:T}\mid\mathbf{x}_{t}^{k_{t}})\equiv p_{\bm{\theta}}(\mathbf{x}_{t}^{k_{t}+1:T}\mid\mathbf{x}_{t}^{k_{t}},\mathbf{z}_{t-1}). This tightness of the ELBO suggests that the expression in Theorem A.1 is a relatively strong surrogate objective for optimizing the likelihoods.

We now special Theorem A.1 to Gaussian diffusion. For now, we focus on the “x\mathbf{x}-prediction” formulation of diffusion, which is the one used in our implementation. The “ϵ\bm{\epsilon}-prediction” formalism, used throughout the main body of the text, can be derived similarly (see Section 2 of for a clean exposition). The following theorem follows directly by apply standard likelihood and KL-divergence computations for the DDPM to Theorem A.1.

and define αk=(1−βk)\alpha_{k}=(1-\beta_{k}), αˉk=∏j=1kαj\bar{\alpha}_{k}=\prod_{j=1}^{k}\alpha_{j}. Suppose that we parameterize pθ(xtj∣xtj+1,zt−1)=N(μθ(xtj+1,zt−1,j),σj2)p_{\bm{\theta}}(\mathbf{x}_{t}^{j}\mid\mathbf{x}_{t}^{j+1},\mathbf{z}_{t-1})=\mathcal{N}(\mu_{\bm{\theta}}(\mathbf{x}_{t}^{j+1},\mathbf{z}_{t-1},j),\sigma_{j}^{2}), where further,

Then, as long as pθp_{\bm{\theta}} satisfies the Markov property, we obtained

where above, we define cj=(1−αj)2αˉj−12σ2(1−αˉj)2c_{j}=\frac{(1-\alpha_{j})^{2}\bar{\alpha}_{j-1}}{2\sigma^{2}(1-\bar{\alpha}_{j})^{2}}.

As noted above, Corollary A.2 can also be stated for ϵ\bm{\epsilon}-prediction, or the so-called “v\mathbf{v}-prediction” formalism, as all are affinely related.

This leads to a striking finding: with expressive enough latents and pθp_{\theta}, we can view the maximization of each term in Corollary A.2 separately across time steps. The absence of this coupling means that the weighting terms are immaterial to the optimization, and thus can be ignored.

A.1.2 Capturing all subsequences

Theorem A.1 stipulates that, up to reweighting, the Diffusion Forcing objective optimizes a valid ELBO on the expected log-likelihoods over uniformly sampled noise levels. The following theorem can be obtained by a straightforward modification of the proof of Theorem A.1 generalizes this to arbitrary (possible temporally correlated) sequences of noise.

Let D\mathcal{D} be an arbitrary distribution over [K]T[K]^{T}, and define Pt(j∣k1:t−1):=Pr⁡D[kt=j∣k1:t−1]P_{t}(j\mid k_{1:t-1}):=\Pr_{\mathcal{D}}[k_{t}=j\mid k_{1:t-1}]. Fix x1:T0\mathbf{x}_{1:T}^{0}. Define the expectation over the forward process with random noise level k1:Tk_{1:T} as

and the expectation over the latents under pθ(⋅)p_{\bm{\theta}}(\cdot) conditioned on k1:T,(xskt)1≤t≤Tk_{1:T},(\mathbf{x}_{s}^{k_{t}})_{1\leq t\leq T} as

Then, as long as pθp_{\bm{\theta}} satisfies the Markov property,

where C(x1:T0)C(\mathbf{x}_{1:T}^{0}) is a constant depending only on x1:T0\mathbf{x}_{1:T}^{0} (the unnoised data), and where the inequality is an equality under the conditions that (a) pθ(zt∣zt−1,xtkt)p_{\bm{\theta}}(\mathbf{z}_{t}\mid\mathbf{z}_{t-1},\mathbf{x}_{t}^{k_{t}}) is a dirac distribution (determinstic latents), and (b) q(xtkt+1:T∣xtkt)≡pθ(xtkt+1:T∣xtkt,zt−1)q(\mathbf{x}_{t}^{k_{t}+1:T}\mid\mathbf{x}_{t}^{k_{t}})\equiv p_{\bm{\theta}}(\mathbf{x}_{t}^{k_{t}+1:T}\mid\mathbf{x}_{t}^{k_{t}},\mathbf{z}_{t-1}), i.e. the variational approximation is sharp.

In particular, in the Gaussian case of Corollary A.2, we have

A.2 Proof of Theorem A.1

Moreover, this lower bound holds with equality if zs∼p(zs∣zs−1,xsks)\mathbf{z}_{s}\sim p(\mathbf{z}_{s}\mid\mathbf{z}_{s-1},\mathbf{x}_{s}^{k_{s}}) is a Dirac distribution (i.e., deterministic latents).

Let’s fix a sequence k1:Tk_{1:T}. It holds that

We now unpack the terms obtained from the preceding claim.

where C1(x0,kt)C_{1}(\mathbf{x}_{0},k_{t}) is a constant depending only on x0\mathbf{x}_{0} and ktk_{t}, and where the inequality holds with equality if and only if q(xtkt+1:T∣xtkt)≡pθ(xtkt+1:T∣xtkt,zt−1)q(\mathbf{x}_{t}^{k_{t}+1:T}\mid\mathbf{x}_{t}^{k_{t}})\equiv p_{\bm{\theta}}(\mathbf{x}_{t}^{k_{t}+1:T}\mid\mathbf{x}_{t}^{k_{t}},\mathbf{z}_{t-1}).

where C2(x0,kt)C_{2}(\mathbf{x}_{0},k_{t}) is some other constant depending on x0\mathbf{x}_{0} and ktk_{t}.

The proof invokes similar manipulations to the standard ELBO derivation for diffusion, but with a few careful modifications to handle the fact that we only noise to level ktk_{t}. As is standard, we require the identity

where (i)(i) uses A.10, (ii)(ii) invokes a cancellation in the telescoping sum, and the final display follows from the computation

Observe that, because we don’t parameterize p(xtK∣zt−1)p(\mathbf{x}_{t}^{K}\mid\mathbf{z}_{t-1}), ln⁡(q(xtkt∣xtkt+1)1{kt≥1})+ln⁡p(xtK∣zt−1)q(xtK∣xtkt)\ln(q(\mathbf{x}_{t}^{k_{t}}\mid\mathbf{x}_{t}^{k_{t}+1})^{\mathbf{1}\{k_{t}\geq 1\}})+\frac{\ln p(\mathbf{x}_{t}^{K}\mid\mathbf{z}_{t-1})}{q(\mathbf{x}_{t}^{K}\mid\mathbf{x}_{t}^{k_{t}})} can be regarded as some constant C′(xtkt,xtkt+1,xtK)C^{\prime}(\mathbf{x}_{t}^{k_{t}},\mathbf{x}_{t}^{k_{t}+1},\mathbf{x}_{t}^{K}). Thus,

We can now simplify to taking expectations. Observe that

We are now ready to complete the proof. By combining the previous two claims, we have

since both terms only depend on k1:t−1,(xsks)1≤s≤t−1k_{1:t-1},(\mathbf{x}_{s}^{k_{s}})_{1\leq s\leq t-1} and z1:t−1\mathbf{z}_{1:t-1}. We conclude then that

as needed. Lastly, we recall that the above is an equality under the conditions that (a) pθ(zt∣zt−1,xtkt)p_{\bm{\theta}}(\mathbf{z}_{t}\mid\mathbf{z}_{t-1},\mathbf{x}_{t}^{k_{t}}) is a dirac distribution, and (b) q(xtkt+1:T∣xtkt)≡pθ(xtkt+1:T∣xtkt,zt−1)q(\mathbf{x}_{t}^{k_{t}+1:T}\mid\mathbf{x}_{t}^{k_{t}})\equiv p_{\bm{\theta}}(\mathbf{x}_{t}^{k_{t}+1:T}\mid\mathbf{x}_{t}^{k_{t}},\mathbf{z}_{t-1}), and we reindex j←j+1j\leftarrow j+1 to ensure consistency with indexing in standard expositions of the diffusion ELBO.

Appendix B Additional Intuitions and Explainations

As stated in Section 2, one can use the gradient of the logarithmic of a classifer log⁡c(y∣xtk)\log c(y|\mathbf{x}_{t}^{k}) to guide the sampling process of diffusion model towards samples with a desired attribute yy. For example, yy can refer to the indicator of a success event. However, we can consider the logarithmic of a more general energy function c(xtk)c(\mathbf{x}_{t}^{k}). This has the interpretation as Pr⁡(y∣xtk)\Pr(y|\mathbf{x}_{t}^{k}), where Pr⁡[y=1∣xtk]=ec(xtk)\Pr[y=1\mid\mathbf{x}_{t}^{k}]=e^{c(\mathbf{x}_{t}^{k})}. Some popular cadidate energies include

B.2 Noising and stabilizing long-horizon generations

B.3 Why Monte Carlo Tree Guidance relies on Diffusion Forcing

Monte Carlo Tree Guidance provides substantial variance reduction in our estimate of a cost-to-go guidance (B.1). This technique crucially relies on the ability to rollout future tokens from current ones to use these sample rollouts to get Monte Carlo estimates for gradients. This is not feasible with full-sequence diffusion, because this requires denoising all tokens in tandem; thus, for a given fixed noise level, there is no obvious source of randomness to use for the Monte Carlo estimate. It may be possible to achieve variable horizon via the trick proposed in the following subsection to simulate future rollouts, but to our knowledge, this approach is nonstandard.

B.4 Does the replacement technique lead to flexible horizons in full-sequence diffusion?

B.5 Further connection to Bayesian filtering

The core idea of Diffusion Forcing can be interpreted as using diffusion to construct an interpolation between prior distribution and posterior distribution of a Bayes filter. Consider the hybrid distribution p(zt∣zt−1,xtk)p(\mathbf{z}_{t}|\mathbf{z}_{t-1},\mathbf{x}_{t}^{k}). When k=0k=0, this hybrid distribution becomes the posterior p(zt∣zt−1,xt)p(\mathbf{z}_{t}|\mathbf{z}_{t-1},\mathbf{x}_{t}). On the other hand, when k=Kk=K, the hybrid distribution becomes p(zt∣zt−1,n)p(\mathbf{z}_{t}|\mathbf{z}_{t-1},\mathbf{n}) for n∼N(0,I)\mathbf{n}\sim\mathcal{N}(0,\mathbf{I}). Since the independent Gaussian noise term n\mathbf{n} contains no information about z\mathbf{z}, this is exactly the prior distribution p(zt∣zt−1)p(\mathbf{z}_{t}|\mathbf{z}_{t-1}). By varying kk between KK and , the same neural network can parameterize everything between prior and posterior.

B.6 Connection to other sequence training schemes

Noise as masking provides a unified view of different sequence training schemes. The following exposition uses a length 33 sequence as an example: We always start with fully masked sequence [x1K,x2K,x3K][\mathbf{x}_{1}^{K},\mathbf{x}_{2}^{K},\mathbf{x}_{3}^{K}] with the goal of denoising it a “clean sequence” of zero noise. [x10,x20,x30][\mathbf{x}_{1}^{0},\mathbf{x}_{2}^{0},\mathbf{x}_{3}^{0}]. Assume all diffusions are sampled with 33-step DDIM.

In teacher forcing, one trains a model to predict the next token conditioned on prior observations. One can train next-token diffusion models with teacher forcing such as : feed neural network with past observations as well as a current observation and ask it to predict clean current observation. A typical training pair can have the input of [x10,x20,x3K]⊤[\mathbf{x}_{1}^{0},\mathbf{x}_{2}^{0},\mathbf{x}_{3}^{K}]^{\top} and target of [x10,x20,x30]⊤[\mathbf{x}_{1}^{0},\mathbf{x}_{2}^{0},\mathbf{x}_{3}^{0}]^{\top}.

At sampling time, one fully diffuses the next token before adding the diffused observation to history to perform an autoregressive rollout. The diffusion process would thus look like

Notably, Diffusion Forcing can also perform this sampling scheme at sampling time for applications like imitation learning, when one wants to diffuse the next action as fast as possible.

Full sequence diffusion models accepts a noisy sequence and denoises level-by-level

Notably, Diffusion Forcing can also perform this sampling scheme at sampling time.

As shown in Figure 2, to model causal uncertainty, Diffusion Forcing keeps the far future more uncertain than the near future by having larger noise level kk, at any time of diffusion. An example pattern looks like this:

Notable, is the first one to propose such a linear uncertainty sampling scheme for causal diffusion models, although Diffusion Forcing provides a generalization of such scheme in combination of other abilities.

Previously we introduced the autoregressive sampling scheme that Diffusion Forcing can also do. However, such a scheme can accumulate single-step errors because it treats predicted x\mathbf{x} as ground truth observation. Diffusion Forcing addresses this problem by telling the model that generated images should be treated as noisy ground truth, as shown in 2.

Then, it feed the diffused x10\mathbf{x}_{1}^{0} into the model but tell it is of a slightly higher noise level, as x11\mathbf{x}_{1}^{1} to diffuse x2\mathbf{x}_{2}.

Then, it feed the diffused x20\mathbf{x}_{2}^{0} into the model but tell it is of a higher noise level, as x21\mathbf{x}_{2}^{1}.

Appendix C Extended Related Work

Masked Autoencoders for images and videos are a popular method for representation learning in pixel space. They have been extended to perform diffusion to generate masked patches conditioned on unmasked ones .

show that even generative modeling of non-sequential data, such as images, can be fruitfully cast as sequence generative modeling.

parameterize token-to-token transitions via a variational auto-encoder. This makes them probabilistic, but does not directly maximize the joint probability of sequences, but rather, enables sampling from the distribution of single-step transitions.

Most similar to our work is AR-Diffusion which similarly aim to train next-token prediction models for sequence diffusion. Key differences are that AR-Diffusion proposes a noise level that is linearly dependent on the position of each word in the sequence, while our critical contribution is to have each noise level be independent, as this uniquely enables our proposed sampling schemes, such as stabilizing auto-regressive generation and conditioning on corrupted observations. Further, AR-Diffusion only explores language modeling and does not explore guidance, while we investigate Diffusion Forcing as a broadly applicable sequence generative model with particular applications to sequential decision making. In particular, we introduce Monte-Carlo Tree Guidance as a novel guidance mechanism.

Appendix D Additional Method Details

We choose both the raw image x\mathbf{x} and latent state z\mathbf{z} to be 2D tensors with channel, width, and height. For simplicity, we use the same width and height for x\mathbf{x} and z\mathbf{z}. We then implement the transition model p(xtkt∣zt−1)p(\mathbf{x}_{t}^{k_{t}}|\mathbf{z}_{t-1}) with a typical diffusion U-net . We use the output of the U-net as the input to a gated recurrent unit (GRU) and use zt−1\mathbf{z}_{t-1} as the hidden state feed into a GRU. The output of GRU is treated as zt\mathbf{z}_{t}. For observation model p(xt∣zt)p(\mathbf{x}_{t}|\mathbf{z}_{t}), we use a 11-layer resnet followed by a conv layer. We combine these two models to create an RNN layer, where the latent of a particular time step is zt−1\mathbf{z}_{t-1}, input is xtkt\mathbf{x}_{t}^{k_{t}} and output is x^\hat{\mathbf{x}}. One can potentially obtain better results by training Diffusion Forcing with a causal transformer architecture. However, since RNN is more efficient for online decision-making, we also stick with it for video prediction and it already gives us satisfying results.

We choose the number of channels in z\mathbf{z} to be 1616 for DMlab and 3232 for Minecraft. In total, our Minecraft model consists of 3636 million parameters and our DMlab model consists of 2424 million parameters. We can potentially obtain a better Minecraft video prediction model with more parameters, but we defer that to future works to keep the training duration reasonable (<1<1 day). In maze planning, the number of total parameters is 4.334.33 million.

For non-spatial x\mathbf{x} that is not video nor images, we use residue MLPs instead of Unet as the backbone for the dynamics model. Residue MLP is basically the ResNet equivalent for MLP. Similar to video prediction, we feed the output of resMLP into a GRU along with zt−1\mathbf{z}_{t-1} to get zt\mathbf{z}_{t}. Another ResMLP serves as the observation model.

D.2 Diffusion parameterization

In diffusion models, there are three equivalent prediction objectives, x0\mathbf{x}_{0}, ϵ\epsilon , and vv parameterization . Different objectives lead to different reweighting of loss at different noise levels, together with SNR reweighting. For example, ϵ\epsilon parameterization and vv parameterization are essential in generating pixel data that favors high-frequency details.

In our experiments, we use vv parameterization for video prediction and found it essential to both convergence speed and quality.

We observe that x0\mathbf{x}_{0} parameterization is strongly favorable in planning and imitation learning, likely because they don’t favor an artificial emphasis on high-frequency details. We observe the benefits of v-parameterization in time-series prediction.

D.3 SNR reweighting

SNR reweighting is a widely used technique to accelerate the convergence of image diffusion models. In short, it reweighs the diffusion loss proportional to the signal-to-noise ratio (SNR) of noisy xk\mathbf{x}^{k}. In Diffusion Forcing, conditioning variable zt−1\mathbf{z}_{t-1} can also contain a non-trivial amount of information about xt\mathbf{x}_{t}, in addition to xtkt\mathbf{x}_{t}^{k_{t}}. For example, in a deterministic markovian system, if xt−1kt−1\mathbf{x}_{t-1}^{k_{t-1}} has its noise level kt−1=0k_{t-1}=0, the posterior state zt−1\mathbf{z}_{t-1} contains all the information needed to predict xt0\mathbf{x}^{0}_{t} regardless of the noise level of xtkt\mathbf{x}_{t}^{k_{t}}.

Therefore we re-derive SNR reweighting to reflect this change in Diffusion Forcing. We follow the intuition of original SNR reweighting to loosely define SNR in a sequence with independent levels of noises at different time steps. Denote StS_{t} as the normalized SNR reweighting factor for xtkt\mathbf{x}_{t}^{k_{t}} following its normal derivation in diffusion models. For example, if one uses min snr strategy , its reweighting factor will always fall between [0,C][0,C] which we divide by CC to get St∈S_{t}\in. Define signal decay factor 0<γ<10<\gamma<1, measuring what proportion of signal in xt−1kt−1\mathbf{x}_{t-1}^{k_{t-1}} contribute to denoising xtkt\mathbf{x}_{t}^{k_{t}}. This is the simple exponential decay model of sequential information. Now, define cumulated SNR recursively as the running mean of StS_{t}: Sˉt=γSˉt−1+(1−γ)St\bar{S}_{t}=\gamma\bar{S}_{t-1}+(1-\gamma)S_{t} to account for signals contributed by the entire noisy history to the denoising at time step tt. The other factor that contributes to the denoising is StS_{t} of noisy observation xtkt\mathbf{x}_{t}^{k_{t}}. To combine them, we use a simplified model for independent events. Notice StS_{t} and Sˉt\bar{S}_{t} always falls in range $,andthereforecanbereinterpretedasprobabilitiesofhavingallthesignaloneneedstoperfectdenoise, and therefore can be reinterpreted as probabilities of having all the signal one needs to perfect denoise\mathbf{x}_{t}^{k_{t}}.Sincethenoiselevelat. Since the noise level attisindependentofpriornoiselevels,wecanviewis independent of prior noise levels, we can viewS_{t}andand\bar{S}_{t-1}asprobabilitiesofindependentevents,andthuscancomposedtodefineajointprobabilityas probabilities of independent events, and thus can composed to define a joint probabilityS^{\prime}_{t}=1-(1-S_{t})(1-\bar{S}_{t-1}),andweusethis, and we use thisS^{\prime}_{t}$ as our fused SNR reweighting factor for diffusion training.

In our experiments, we choose to follow the min-SNR reweighting strategy to derive the SS. Our new SNR reweighting proves extremely useful to accelerate the convergence of video prediction, while we didn’t observe a boost on non-image domains so we didn’t use it there.

D.4 Noise schedule

We use sigmoid noise schedule for video prediction, linear noise schedule for maze planning, and cosine schedule for everything else.

D.5 Implementation Details of Sampling with Guidance

In our sampling algorithm, due to the flexibility of the scheduling matrix K\mathcal{K}, there are corner cases when xtkt\mathbf{x}_{t}^{k_{t}} is required to stay at its same noise level during a sampling step. The core question of this corner case is whether we should updatextkt\mathbf{x}_{t}^{k_{t}} at all. One option is just copying over the old value. The other option is to run a backward diffusion followed by a forward diffusion back to its old noise level to resample under the diffusion process. While we conclude this can be an open question, we prefer the later approach, resampling, and use it in Monte Carlo Guidance to generate multiple samples. We note that even if one take the first approach, the guidance gradient can still flow back in the time steps before tt as the dynamics model p(zt∣xtkt,zt−1)p(\mathbf{z}_{t}|\mathbf{x}_{t}^{k_{t}},\mathbf{z}_{t-1}) can still propagate the guidance gradient to zt−1\mathbf{z}_{t-1}.

Other than Monte Carlo Guidance, this corner case only happens when kt=0k_{t}=0 or kt=Kk_{t}=K throughout our experiments. That is, we chose our K\mathcal{K} such that once any token gets diffused slightly, it will keep diffusing. In the case of kt=Kk_{t}=K, keeping xtkt\mathbf{x}_{t}^{k_{t}} at the same noise level implies it will stay as white noise, and we don’t even need to sample another white noise. In case kt=0k_{t}=0, the time step is already completely diffused either approach should give us the same result so we just opt for copying over for simplicity.

In maze planning, our main baseline Diffuer discards the reward from the dataset and directly plans with the goal position and velocity. We adopt the same convention for Diffusion Forcing. One can perform guidance on goal position using log-likelihood ∣∣pT−g∣∣||\mathbf{p}_{T}-\mathbf{g}||, but a flexible horizon model should not require users to manually specify a TT to reach its goal, instead we want it to try to reach the goal for any possible horizon. Therefore we use the reward model ∑t∣∣pT−g∣∣\sum_{t}||\mathbf{p}_{T}-\mathbf{g}|| so any time step can be the final step to reach the goal. This objective is challenging due to the non-convex nature of 2D maze, but we found Diffusion Forcing can still reliably find plans without bumping into walls. However, we also observe that the agent tend to leave the goal location due to the nature of the provided dataset - the goal location is just one possible waypoint for the robot to pass through, and there are no trajectories that simply stay at the goal. We also tried this reward for guidance with Diffuser, but it didn’t work even with a good amount of tuning.

D.6 Performance Optimization

Accelerating the diffusion sampling of Diffusion Forcing is similar to that of normal diffusion models. We adopt DDIM sampling for the diffusion of each token. While we use K=1000K=1000 steps of diffusion, we sample with only 100100 DDIM for video prediction and 5050 for non-video domains.

While Diffusion Forcing can be implemented with transformers, we use an RNN as the backbone for Diffusion Forcing experiments it’s widely used in decision-making for its flexibility and efficiency in online decision-making systems. To further reduce training time and GPU memory usage, we use frame-stacking to stack multiple observed images as a single x\mathbf{x}. This is due to the fact that adjacent tokens can be very similar - e.g. recording the same motion at higher fps can lead to this. We deem that it’s wasteful if we rollout the dynamics model multiple times to generate almost identical tokens. For video datasets, we manually examine how many time steps it takes to require a minimal level of prediction power instead of copying frames over. There is another reason why we use frame stacking - many diffusion model techniques such as different noise schedules are designed to model x\mathbf{x} with correlated elements or redundancy. Low-dimensional systems may need drastically different hyperparameters when they lack the data redundancy these techniques are tested on. Frame stacking is thus also helpful for our non-image experiments so we can start with canonical hyperparameters of diffusion models. We use a frame stack of 44 for DMlab video prediction, 88 for Minecraft and 1010 for maze planning.

At sampling time, we also have a design choice to reduce compute usage, as reflected in line 8 of Algorithm 2. In line 8, we directly assign ztnew\mathbf{z}_{t}^{\text{new}} to zt\mathbf{z}_{t}, instead of recalculating zt\mathbf{z}_{t} with posterior model p(zt∣zt−1,xtnew,k−1)p(\mathbf{z}_{t}|\mathbf{z}_{t-1},\mathbf{x}_{t}^{\text{new}},k-1). Since the model is trained to condition on zt\mathbf{z}_{t} estimated from arbitrary noisy history, we recognize that both are valid approaches. The reason why the choose line 88 is two fold. First, it cuts the compute by half, avoiding computing posterior every step. Second, this happens to be what we want for stabilization - ztnew\mathbf{z}_{t}^{\text{new}} already contains the information of the clean xtnew\mathbf{x}_{t}^{\text{new}} under our simplified observation model, and happens to be estimated with k=ktk=k_{t}, a noise level higher than that of xtnew\mathbf{x}_{t}^{\text{new}}. This happens to implement the behavior we want for stabilization.

D.7 Sampling schedule for causal uncertainty

Inference is depicted in Algorithm 2 and Figure 2. In Equation D.1, we illustrate a specific instantiation of the K\mathcal{K} matrix we used for causal planning. For simplicity, we denote the case where a latent z0\mathbf{z}_{0} is given and aim to generate x1:H+1\mathbf{x}_{1:H+1}.

Diffusion Forcing begins by sampling our sequences as white noise with noise level KK. It then denoises along each row m=1,…,Mm=1,\dots,M of K\mathcal{K} in decreasing order. It does so by proceeding sequentially through frames t=1,…,Tt=1,\dots,T, updating the latent (Line 5 of Algorithm 2), and then partially applying the backward process to noise level k=Km,tk=\mathcal{K}_{m,t} dictated by the scheduling matrix K\mathcal{K} (Line 6-7 of Algorithm 2). We call a K\mathcal{K} like this pyramid scheduling, as the tokens in the far future are kept at higher noise level than near future.

D.8 Implementation Details of Timeseries Regression

We follow the implementation of pytorch-ts, where the validation set is a random subset of the training set with the same number of sequences as the test set. We use early stopping when validation crpssum hasn’t increased for 6 epochs. We leverage the same architecture (1 mlp and 4 grus) as well as a batch size of 32.

D.9 Compute Resources

All of our experiments use fp16fp16 mixed precision training. Time series, maze planning, compositionally, visual imitation experiments can be trained with a single 2080Ti2080Ti with 1111GB of memory. We tune the batch size such that we fully use the memory of GPUs. This translates to a batch size of 20482048 for maze planning and compositional experiments, and 3232 for visual imitation learning. While we use early stopping on the validation set for time series experiments, we did not carefully search for the minimal number of training steps required, though the model usually converges between 5050k to 100k100k steps. The above environments thus usually take 4−84-8 hours to train although there is without doubt a significant potential for speed up.

Video prediction is GPU intensive. We use 88 A100 GPUs for both video prediction datasets. We train for 50K50K steps with a batch size of 8×168\times 16. It usually take 1212 hours to converge at 40K40K steps of training (occasional validation time also included).

Appendix E Additional Experiment Results

To illustrate Diffusion Forcing’s new training objective does not degrade it as a generic sequence model, we evaluate Diffusion Forcing on high-dimensional and long-horizon sequence prediction tasks in time series prediction. We adopt multiple time series datasets with real-world applications from GluonTS and evaluate Diffusion Forcing with strong baselines with standard metrics in this domain. In this section, we mainly focus on the results and analysis. For a detailed description of datasets and the metric, we refer the reader to Appendix F.4.

over the samples in the prediction window. If we know the distribution in (E.1), we can sample forecast prediction sequences given some initial context from the evidence sequence. However, most time-dependent data generation processes in nature have complex dynamics and no tractable formulation of q(xt0:T∣x1:t0−1)q{\left(\mathbf{x}_{t_{0}:T}\mid\mathbf{x}_{1:t_{0}-1}\right)}. Instead, we construct a statistical model that approximates the generative process in (E.1) and estimates quantiles via Monte Carlo sampling of simulated trajectories. In this way, confidence levels or uncertainty measures can be calculated, and point forecasts can be produced as the mean or median trajectory .

[ caption = Results for time series forecasting. We report the test set CRPSsum\text{CRPS}_{\textbf{sum}} (the lower, the better) of comparable methods on six time series datasets. We measure the mean and standard deviation of our method from five runs trained with different seeds., label = tab:results_ts, pos = ht, doinside = , center, star ]lccccccc \FLMethod & Exchange Solar Electricity Traffic Taxi Wikipedia \MLVES 0.005 ±\pm 0.000 0.900 ±\pm 0.003 0.880 ±\pm 0.004 0.350 ±\pm 0.002 - - \NNVAR 0.005 ±\pm 0.000 0.830 ±\pm 0.006 0.039 ±\pm 0.001 0.290 ±\pm 0.001 - - \NNVAR-Lasso 0.012 ±\pm 0.000 0.510 ±\pm 0.006 0.025 ±\pm 0.000 0.150 ±\pm 0.002 - 3.100 ±\pm 0.004 \NNGARCH 0.023 ±\pm 0.000 0.880 ±\pm 0.002 0.190 ±\pm 0.001 0.370 ±\pm 0.001 - - \NNDeepAR - 0.336 ±\pm 0.014 0.023 ±\pm 0.001 0.055 ±\pm 0.003 - 0.127 ±\pm 0.042 \NNLSTM-Copula 0.007 ±\pm 0.000 0.319 ±\pm 0.011 0.064 ±\pm 0.008 0.103 ±\pm 0.006 0.326 ±\pm 0.007 0.241 ±\pm 0.033 \NNGP-Copula 0.007 ±\pm 0.000 0.337 ±\pm 0.024 0.025 ±\pm 0.002 0.078 ±\pm 0.002 0.208 ±\pm 0.183 0.086 ±\pm 0.004 \NNKVAE 0.014 ±\pm 0.002 0.340 ±\pm 0.025 0.051 ±\pm 0.019 0.100 ±\pm 0.005 - 0.095 ±\pm 0.012 \NNNKF - 0.320 ±\pm 0.020 0.016 ±\pm 0.001 0.100 ±\pm 0.002 - 0.071 ±\pm 0.002 \NNTransformer-MAF 0.005 ±\pm 0.003 0.301 ±\pm 0.014 0.021 ±\pm 0.000 0.056 ±\pm 0.001 0.179 ±\pm 0.002 0.063 ±\pm 0.003 \NNTimeGrad 0.006 ±\pm 0.001 0.287 ±\pm 0.020 0.021 ±\pm 0.001 0.044 ±\pm 0.006 0.114 ±\pm 0.020 0.049 ±\pm 0.002 \NNScoreGrad sub-VP SDE 0.006 ±\pm 0.001 0.256 ±\pm 0.015 0.019 ±\pm 0.001 0.041 ±\pm 0.004 0.101 ±\pm 0.004 0.043 ±\pm 0.002 \NNOurs 0.003 ±\pm 0.001 0.289 ±\pm 0.002 0.023 ±\pm 0.001 0.040 ±\pm 0.004 0.075 ±\pm 0.002 0.085 ±\pm 0.007 \LL We evaluate the effectiveness of Diffusion Forcing as a sequence model on the canonical task of multivariate time series forecasting by following the experiment setup of Concretely, we benchmark Diffusion Forcing on the datasets Solar, Electricity, Traffic, Taxi, and Wikipedia. These datasets have different dimensionality, domains, sampling frequencies, and capture seasonal patterns of different lengths. The features of each dataset are detailed in Table LABEL:tab:ts_data. We access the datasets from GluonTS , and set the context and prediction windows to the same length for each dataset. Additionally, we employ the same covariates as . We evaluate the performance of the model quantitatively by estimating the Summed Continuous Ranked Probability Score CRPS⁡sum\operatorname{CRPS}_{\text{sum}} via quantiles. As a metric, CRPS⁡sum\operatorname{CRPS}_{\text{sum}} measures how well a forecast distribution matches the ground truth distribution. We provide detailed descriptions of the metric in Appendix F.4. We benchmark with other diffusion-based methods in time series forecasting, such as TimeGrad and the transformer-based Transformer-MAF . In particular, the main baseline of interest, TimeGrad , is a next-token diffusion sequence model trained with teacher forcing. We track the CRPS⁡sum\operatorname{CRPS}_{\text{sum}} metric on the validation set and use early stopping when the metric has not improved for 6 consecutive epochs, while all epochs are fixed to 100 batches across datasets. We then measure the CRPS⁡sum\operatorname{CRPS}_{\text{sum}} on the test set at the end of training, which we report in Table LABEL:tab:results_ts. We use the exact same architecture and hyperparameters for all time series datasets and experiments. Diffusion Forcing outperforms all prior methods except for with which Diffusion Forcing is overall tied, except for the Wikipedia dataset, on which Diffusion Forcing takes fourth place. Note that time series is not the core application of Diffusion Forcing, and that we merely seek to demonstrate that the Diffusion Forcing objective is applicable to diverse domains with no apparent trade-off in performance over baseline objectives.

E.2 Additional results in compositional generation

Since Diffusion Forcing models the joint distribution of any subset of a sequence, we can leverage this unique property to achieve compositional behavior - i.e., Diffusion Forcing can sample from the distribution of subsets of the trajectory and compose these sub-trajectories into new trajectories.

In particular, we show that we can also have flexible control over how compositional Diffusion Forcing is. As shown in 7, consider a dataset of trajectories on a 2D, square plane, where all trajectories start from one corner and end up in the opposite corner, forming a cross shape. When no compositional behavior is desired, one can let the models replicate the cross-shaped distribution by allowing full memory of the HMM model. When one desires compositional such as generating a V-shaped trajectory, which stitches two sub trajectories together, one can let the model generate shorter plans with no-memory context using MPC. (Add figures).

E.3 Additional results in video prediction (wo/ cherry picking)

Diffusion Forcing can rollout longer than maximum training horizonwithout sliding window. That is, we run Diffusion Forcing’s RNN continuously without ever reinitializing z0\mathbf{z}_{0}. This is a surprising effect we observed from the rollout stabilization property of Diffusion Forcing. In Figure 8, 10, we use Diffusion Forcing to generate video sequences of length 180180 and visualize subsampled sequences. Notably, Diffusion Forcing used in these visualizations is trained with a maximum length of 7272 frames for Minecraft and 3636 frames for DMLab, illustrating it can rollout 2x-5x times longer than it’s trained on without sliding window. In addition, we also tried rolling these models out for 20002000 frames and without seeing the model blowing up on both datasets. There are occasional cases where the Minecraft agent gets stuck and the entire screen is the “dirt” block, but this is more of a dataset issue 16 and the agent is able to recover after it turns around.

We also present additional results where we only generate within our maximum training length. As shown in figure 13 12, Diffusion Forcing can generate consistent videos. Results are not cherry-picked.

E.4 Additional results in planning

We provide some additional visualizations of causal planning in 15. We also present additional visualization of Diffusion Forcing performing model predictive control in action. As shown in figure 14, Diffusion Forcing can generate plans of shorter horizons since it’s flexible horizon.

E.5 Real robot experiment setup

In Figure 16 we visualize our robot experiment setup with corruption on observation. The dataset is collected when the target bag isn’t present, while we test with such a bag in the scene zero-shot for the imitation learning experiment with observation corruption. The typical failure mode is when the robot no longer reacts to the visual clues of the randomized location of objects. We didn’t observe the robot to act wildly due to visual distractors.

Appendix F Additional details about datasets

We adopt the video prediction dataset Minecraft and DMlab used by TECO.

The Minecraft navigation dataset consists of first-person-view videos of random walks in the Minecraft ‘swamp‘ biome. The agent walks via a technique called ‘sprint jump‘ which allows it to jump across blocks without getting stuck at 1 block obstacles. The agent walks straight most of the time, with small chances of turning left or right. The height and width of the video is 128128 pixels and we trim long videos to subsequences of 7272 frames. The dataset comes with paired action data but we discard them to bring more stochasticity to the prediction task. Due to limited compute, we only train on about 10%10\% of the total subsequences.

One problem we noticed about the dataset is when the agent runs into obstacles with a height of 2 blocks or more. In this case, the agent will get stuck and the entire video sequence will consist of grey granite patterns or brown dirty patterns. This leads to a huge amount of frames with these patterns, making video models predict meaningless frames. Yet, we deem this as a problem of this dataset itself.

Deepmind Lab navigation dataset consists of random walks in a 3D maze environment. For DMLab, the resolution is 6464 pixels and we use subsequences of 4848 frames. We also disregard the provided actions due to training.

We note that the VQ-VAE latent that stable video diffusion diffuses is also only 128×128×3128\times 128\times 3, indicating Diffusion Forcing has the potential to scale up to higher resolution images with pre-trained image encoder and decoders. Due to the sheer size of the datasets, we only use about 10%10\% of the total data sequences for training due to limited computing, as we observe that doing so already allows us to make good generations from initial frames from the test set.

F.2 Dataset for planning

D4RL is a standard offline RL benchmark featuring a wide range of reinforcement learning environments. Each environment is associated with a provided dataset of offline interactions with the environment featuring state, action and reward trajectories.

Like Diffuer , we choose the 3 maze environments as they are challenging long-horizon, multi-modal, sparse reward problems uniquely suited for visualization and evaluating planning algorithms. The ID for the 3 used environments are “maze2d-medium-v1”, “maze2d-large-v1”, “maze2d-umaze-v1”. In each environment, one controls the acceleration of a robot to walk it towards a goal. The observation space is 44 dimensional, featuring 2D location and velocity. The action space is 2D acceleration. The agent always receives a random start location and the goal is to reach a fixed goal position for each maze. The agent receives a reward of 1 if it is within a circle of radius 0.5 centered at the goal state, and 0 otherwise.

The offline RL dataset for the maze environments consists of random walks in the maze. Specifically, the authors first designate all intersections and turns in the maze as waypoints and code an agent to navigate between waypoints with some randomization. As a result, the random walks are generated in a way that the path is collision-free with the walls. The random walks introduces stochasticity to the dataset, as trajectories in the dataset are never towards a specific goal.

There are a few choices adopted from our main baseline Diffuser : we disregard the reward in the dataset and plan with goals only. We also evaluate a multi-gaol variant of each environment (labeled as “multi” in Table 4.1), where the goal is randomized just like the starting position.

F.3 Dataset for robot learning

We choose a long horizon robotic manipulation task as described in Section 4.4: Consider a tabletop with three slots where we can place objects. One places an apple at slot A or slot B randomly, and then places an orange at the other slot between A and B. A robot is challenged to swap the position of two fruits using the third slot C. That is, it can only move a fruit to an empty slot at a time. For example, when the apple is at slot A and the orange is at slot B, it may move the apple to slot C, leaving slot A empty. Then move the orange to slot A and finally move the apple from slot C to slot B. In figure 4, we illustrate the non-markovian property of the task: When the apple is at slot B and the orange is at slot C, one cannot tell what the immediate action is without knowing the initial positions of objects.

We put stickers on the table indicating a circular region occupied by any slot. Each circular region is designed to be about double the diameter of a fruit. To make sure the task requires visual feedback, we also randomize the location of a fruit inside the slot. We collected 150150 expert demonstrations of a franka robot performing the task using VR teleoperation and impedance control. Among them, each initial slot configuration makes up half of the dataset. We record videos from two camera views, one from hand camera and one in the front capturing all three slots. Each demonstration also comes with 66 dof actions of the robot hand. During the data collection, since one successful demonstration will swap the position of two objects, its end configuration will naturally serve as the starting configuration of the other randomized location, which we leverage to save time.

Each demonstration comprises 500−600500-600 frames and actions. We train Diffusion Forcing on the entire sequence. However, since adjacent frames are visually close, we pad and downsample the videos to 4040 frames where each frame is bundled with 1515 actions.

F.4 Dataset for time series

Often, statistical models that approximate (E.1) benefit from manually curated features as additional input to the observations. A sequence of covariates C={ct}t=1T\boldsymbol{C}=\left\{\mathbf{c}_{t}\right\}_{t=1}^{T} can be constructed to help the model recognize seasonal patterns and other temporal dependencies. We follow the implementation in to construct the covariate sequence as a function of the frequency of each dataset in Table LABEL:tab:ts_data. As such, our covariates are composed of lagged inputs, as well as learned embeddings and handcrafted temporal features that encode information such as the hour of the day or the day of the month, depending on the sampling rate of the particular time series that is being modeled. Therefore, covariates are known for the entire interval [1,T]\left[1,T\right], even at inference. We can easily incorporate covariates into the probabilistic framework as

The benefit obtained from covariates is highly dependent on the characteristics of both the dataset and the model used, as well as the feature engineering practices followed.

The Continuous Ranked Probability Score (CRPS) is a scoring function that measures how good the forecast distribution matches the ground truth distribution:

as the average over the prediction window. The lower the CRPS⁡sum\operatorname{CRPS}_{\text{sum}} value, the better does the predicted distribution match the data distribution.

First, we manually sum the time series along the feature dimension and estimate the CDF F^sum(t)\hat{F}_{\text{sum}}(t) via 19 quantile levels at each time step tt from 100 sampled trajectories. We then use the implementation in GluonTs to compute the CRPS⁡\operatorname{CRPS}, which we report as CRPS⁡sum\operatorname{CRPS}_{\text{sum}} in Table LABEL:tab:results_ts. While we aggregate the data manually, we verify that the numerical error relative to the GluonTS implementation remains orders of magnitude below the precision threshold of the reported metric.