SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers

Nanye Ma, Mark Goldstein, Michael S. Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden, Saining Xie

Introduction

Contemporary success in image generation has come from a combination of algorithmic advances and improvements in model architecture and progress in scaling neural network models and data. State-of-the-art diffusion models proceed by incrementally transforming data into Gaussian noise as prescribed by an iterative stochastic process, which can be specified either in discrete or continuous time. At an abstract level, this corruption process can be viewed as defining a time-dependent distribution that is iteratively smoothed from the original data distribution into a standard normal distribution. Diffusion models learn to reverse this corruption process and push Gaussian noise backwards along this connection to obtain data samples. The objects learned to perform this transformation are conventionally either predicting the noise in the corruption process or predicting the score of the distribution connecting the data and the Gaussian , though alternatives of these choices exist .

The neural network architectures used to represent these objects have been shown to perform well on a variety of tasks. While diffusion models were originally built upon a U-Net backbone , recent work has highlighted that architectural advances in vision such as the Vision Transformer (ViT) can be incorporated into the standard diffusion model pipeline to improve performance . The aims of were to push improvements on the model side of the duality of algorithm and model.

Orthogonally, significant research effort has gone into exploring the structure of the noising process, which has been shown to lead to performance benefits . Yet, many of these efforts do not move past the notion of passing data through a diffusion process with an equilibrium distribution, which is a restricted type of connection between the data and the Gaussian. The recently-introduced stochastic interpolants lift such constraints and introduce more flexibility in the noise-data connection. In this paper, we further explore its performance in large scale image generation.

Intuitively, we expect that the difficulty of the learning problem can be related to both the specific connection chosen and the object that is learned. Our aim is to clarify these design choices, in order to simplify the learning problem and improve performance. To glean where potential benefits arise in the learning problem, we start with Denoising Diffusion Probabilistic Models (DDPMs) and sweep through adaptations of: (i) which object to learn, and (ii) which interpolant to choose to reveal best practices.

In addition to the learning problem, there is a sampling problem that must be solved at inference time. It has been acknowledged for diffusion models that sampling can be either deterministic or stochastic , and the choice of sampling method can be made after the learning process. Yet, the diffusion coefficients used for stochastic sampling are typically presented as intrinsically tied to the forward noising process, which need not be the case in general.

Throughout this paper, we explore how the design of the interpolant and the use of the resulting model as either a deterministic or a stochastic sampler impact performance. We gradually transition from a typical denoising diffusion model to an interpolant model by taking a series of orthogonal steps in the design space. As we progress, we carefully evaluate how each move away from the diffusion model impacts the performance. In summary, our main contributions are:

By moving from discrete to continuous time, changing the model prediction, interpolant, and the choice of sampler, we observe a consistent performance improvement over the Diffusion Transformer (DiT).

We systematically study where these improvements come from by addressing these factors one by one: learning in continuous time; learning a velocity as compared to a score; changing the interpolant connecting the the two distributions; and using the velocity in an SDE sampler with particular choices of diffusion coefficients.

We show that the SDE for the interpolant can be instantiated using just a velocity model, which we use to push the performance of these methods beyond previous results.

SiT: Scalable Interpolant Transformers

We begin by recalling the main ingredients for building flow-based and diffusion-based generative models.

In recent years, a flexible class of generative models based on turning noise ε∼N(0,I){\boldsymbol{\varepsilon}}\sim\mathsf{N}(0,\mathbf{I}) into data x∗∼p(x)\mathbf{x}_{*}\sim p(\mathbf{x}) have been introduced. These models use the time-dependent process

where αt\alpha_{t} is a decreasing function of tt and σt\sigma_{t} is an increasing function of tt. Stochastic interpolants and other flow matching methods restrict the process (1) on t∈t\in, and set α0=σ1=1\alpha_{0}=\sigma_{1}=1, α1=σ0=0\alpha_{1}=\sigma_{0}=0, so that xt{\mathbf{x}}_{t} interpolates exactly between x∗{\mathbf{x}}_{*} at time t=0t=0 and ε{\boldsymbol{\varepsilon}} and time t=1t=1. By contrast, score-based diffusion models set both αt\alpha_{t} and σt\sigma_{t} indirectly through different formulations of a stochastic differential equation (SDE) with N(0,I)\mathsf{N}(0,\mathbf{I}) as its equilibrium distribution. Moreover, they consider the process xt{\mathbf{x}}_{t} on an interval [0,T][0,T] with TT large enough that xT{\mathbf{x}}_{T} approximates a Gaussian distribution.

Common to both stochastic interpolants and score-based diffusion models is the observation that the process xt{\mathbf{x}}_{t} can be sampled dynamically using either an SDE or a probability flow ordinary differential equation (ODE). More precisely, the marginal probability distribution pt(x)p_{t}({\mathbf{x}}) of xt{\mathbf{x}}_{t} in (1) coincides with the distribution of the probability flow ODE with a velocity field

where v(x,t){\mathbf{v}}({\mathbf{x}},t) is given by the conditional expectation

Equation (3) is derived in Sec. A.1. By solving the probability flow ODE (2) backwards in time from XT=ε∼N(0,I){\mathbf{X}}_{T}={\boldsymbol{\varepsilon}}\sim\mathsf{N}(0,\mathbf{I}), we can generate samples from p0(x)p_{0}({\mathbf{x}}), which approximates the ground-truth data distribution p(x)p({\mathbf{x}}). We refer to (2) as a flow-based generative model.

Reverse-time SDE.

The time-dependent probability distribution pt(x)p_{t}({\mathbf{x}}) of xt{\mathbf{x}}_{t} also coincides with the distribution of the reverse-time SDE

where Wˉt\bar{\mathbf{W}}_{t} is a reverse-time Wiener process, wt>0w_{t}>0 is an arbitrary time-dependent diffusion coefficient, v(x,t){\mathbf{v}}({\mathbf{x}},t) is the velocity defined in (3), and where s(x,t)=∇log⁡pt(x){\mathbf{s}}({\mathbf{x}},t)=\nabla\log p_{t}({\mathbf{x}}) is the score. Similar to v{\mathbf{v}}, this score is given by the conditional expectation

This equation is derived in Sec. A.3. Similarly, solving the reverse SDE (4) backwards in time from XT=ε∼N(0,I){\mathbf{X}}_{T}={\boldsymbol{\varepsilon}}\sim\mathsf{N}(0,\mathbf{I}) enables generating samples from the approximated data distribution p0(x)∼p(x)p_{0}(\mathbf{x})\sim p({\mathbf{x}}). We refer to (2) as a diffusion-based generative model.

Design choices.

Score-based diffusion models typically tie the choice of αt\alpha_{t}, σt\sigma_{t}, and wtw_{t} in (4) to the drift and diffusion coefficients used in the forward SDE that generates xt{\mathbf{x}}_{t} (see (10) below). The stochastic interpolant framework decouples the formulation of xt{\mathbf{x}}_{t} from the forward SDE and shows that there is more flexibility in the choices of αt\alpha_{t}, σt\sigma_{t}, and wtw_{t}. Below, we will exploit this flexibility to construct generative models that outperform score-based diffusion models on standard benchmarks in image generation task.

2 Estimating the score and the velocity

Practical use of the probability flow ODE (2) and the reverse-time SDE (4) as generative models relies on our ability to estimate the velocity v(x,t){\mathbf{v}}({\mathbf{x}},t) and/or score s(x,t){\mathbf{s}}({\mathbf{x}},t) fields that enter these equations. The key observation made in score-based diffusion models is that the score can be estimated parametrically as sθ(x,t){\mathbf{s}}_{\theta}({\mathbf{x}},t) using the loss

This loss can be derived by using (5) along with standard properties of the conditional expectation. Similarly, the velocity in (3) can be estimated parametrically as vθ(x,t){\mathbf{v}}_{\theta}({\mathbf{x}},t) via the loss

We note that any time-dependent weight can be included under the integrals in both (6) and (7). These weight factors are key in the context of score-based models when TT becomes large ; in contrast, with stochastic interpolants where T=1T=1 without any bias, these weights are less important and might impose numerical stability issue (see Appendix B).

We observed that only one of the two quantities sθ(x,t){\mathbf{s}}_{\theta}({\mathbf{x}},t) and vθ(x,t){\mathbf{v}}_{\theta}({\mathbf{x}},t) needs to be estimated in practice. This follows directly from the constraint

which can be used to re-express the score (5) in terms of the velocity (3) as

We will use this relation to specify our model prediction. Conversely, we can also express v(x,t){\mathbf{v}}({\mathbf{x}},t) in terms of s(x,t){\mathbf{s}}({\mathbf{x}},t). In our experiments, we typically learn the velocity field v(x,t){\mathbf{v}}({\mathbf{x}},t) and use it to express the score s(x,t){\mathbf{s}}({\mathbf{x}},t) when using an SDE for sampling. We include a detailed derivation in Sec. A.4.

Note that by our definitions α˙t<0\dot{\alpha}_{t}<0 and σ˙t>0\dot{\sigma}_{t}>0, so that the denominator of (9) is never zero. Yet, σt\sigma_{t} vanishes at t=0t=0, making the σt−1\sigma_{t}^{-1} in (9) appear to cause a singularity thereWe remark that s(x,t){\mathbf{s}}({\mathbf{x}},t) can be shown to be non-singular at t=0t=0 analytically if the data distribution p(x)p({\mathbf{x}}) has a smooth density , though this singularity appears in numerical implementations and losses in general.. This suggests the choice wt=σtw_{t}=\sigma_{t} in (4) to cancel this singularity (see Sec. A.3), for which we will explore the performance in the numerical experiments.

3 Specifying the interpolating process

In Score-Based Diffusion Models (SBDM), the choice of αt\alpha_{t} and σt\sigma_{t} in (1) is typically determined by the choice of the forward SDE used to define this process, though recent work has tried to reconsider this . For example, if we use the standard variance-preserving (VP) SDE

for some βt>0\beta_{t}>0, it can be shown (see Appendix B) that the solution to (10) has the same probability distribution pt(x)p_{t}({\mathbf{x}}) as the process xt{\mathbf{x}}_{t} defined in (1) for the choice

The only design flexibility in (11) comes from the choice of βt\beta_{t}, because it determines both αt\alpha_{t} and σt\sigma_{t} VP is the only linear scalar SDE with an equilibrium distribution ; interpolants extend beyond αt2+σt2=1\alpha_{t}^{2}+\sigma_{t}^{2}=1 by foregoing the requirement of an equilibrium distribution. . For example, setting βt=1\beta_{t}=1 leads to αt=e−t\alpha_{t}=e^{-t} and σt=1−e−2t\sigma_{t}=\sqrt{1-e^{-2t}}. This choice necessitates taking TT sufficiently large or searching for more appropriate choices of βt\beta_{t} to reduce the bias induced by the fact that the solution to the SDE (10) only converges to ε∼N(0,I){\boldsymbol{\varepsilon}}\sim\mathsf{N}(0,\mathbf{I}) as t→∞t\to\infty.

General interpolants.

In the stochastic interpolant framework, the process (1) is defined explicitly and without any reference to a forward SDE, creating more flexibility in the choice of αt\alpha_{t} and σt\sigma_{t}. Specifically, any choice satisfying:

αt2+σt2>0\alpha_{t}^{2}+\sigma_{t}^{2}>0 for all t∈t\in;

αt\alpha_{t} and σt\sigma_{t} are differentiable for all t∈t\in;

α1=σ0=0\alpha_{1}=\sigma_{0}=0, α0=σ1=1\alpha_{0}=\sigma_{1}=1;

gives a process that interpolates without bias between xt=0=x∗{\mathbf{x}}_{t=0}={\mathbf{x}}_{*} and xt=1=ε{\mathbf{x}}_{t=1}={\boldsymbol{\varepsilon}}. In our numerical experiments, we exploit this design flexibility to test, in particular, the choices

where GVP refers to a generalized VP which has constant variance across time for any endpoint distributions with the same variance. We note that the fields v(x,t){\mathbf{v}}({\mathbf{x}},t) and s(x,t){\mathbf{s}}({\mathbf{x}},t) entering (2) and (4) depend on the choice of αt\alpha_{t} and σt\sigma_{t}, and typically must be specified before learningThe requirement to learn and sample under one choice of path specified by αt,σt\alpha_{t},\sigma_{t}, at training time may be relaxed and is explored in .. This is in contrast to the diffusion coefficient w(t)w(t), as we now describe.

4 Specifying the diffusion coefficient

As stated earlier, the SBDM diffusion coefficient used in the reverse SDE (4) is usually taken to match that of the forward SDE (10). That is, one sets wt=βtw_{t}=\beta_{t}. In the stochastic interpolant framework, this choice is again subject to greater flexibility: any wt>0w_{t}>0 can be used. Interestingly, this choice can be made after learning, as it does not affect the velocity v(x,t){\mathbf{v}}({\mathbf{x}},t) or the score s(x,t){\mathbf{s}}({\mathbf{x}},t). In our experiments, we exploit this flexibility by considering the choices listed in Table 2.

5 Time-discretization and link with DDPM

During inference, continuous time models must be discretized when solving the probability flow ODE (2) and the reverse-time SDE (4). This allows us to make a link with DDPMs .

Assuming that we discretize time using a grid 0=t0<t1<t2<…<tN=T0=t_{0}<t_{1}<t_{2}<\ldots<t_{N}=T, the process (1) can be evaluated at each grid point, xti=αtix∗+σtiε{\mathbf{x}}_{t_{i}}=\alpha_{t_{i}}{\mathbf{x}}_{*}+\sigma_{t_{i}}{\boldsymbol{\varepsilon}}, and both the velocity and the score can be estimated on these points via the losses

Moreover, only the learned sθ(x,ti){\mathbf{s}}_{\theta}({\mathbf{x}},{t_{i}}) or vθ(x,ti){\mathbf{v}}_{\theta}({\mathbf{x}},{t_{i}}) is needed to integrate the probability flow ODE (2) and the reverse-time SDE (4) on the same grid. The resulting procedure, in which we define xt{\mathbf{x}}_{t} iteratively on the grid, is a generalization of DDPM. Starting from xt0=x∗{\mathbf{x}}_{t_{0}}={\mathbf{x}}_{*}, we set for i≥0i\geq 0,

where h=ti+1−tih=t_{i+1}-t_{i} and where we assume that the grid is uniform. Because 1−hβti=1−12hβti+o(h)\sqrt{1-h\beta_{t_{i}}}=1-\frac{1}{2}h\beta_{t_{i}}+o(h), it is easy to see that (15) is a consistent time-discretization of the forward SDE (10). Our results show that it is not necessary to specify the time discretized process xti{\mathbf{x}}_{t_{i}} using (15), but instead we can directly use (1) on the time grid.

6 Interpolant Transformer Architecture

The backbone architecture and capacity of generative models are also crucial for producing high-quality samples. In order to eliminate any confounding factors and focus on our exploration, we strictly follow the standard Diffusion Transformer (DiT) and its configurations. This way, we can also test the scalability of our model across various model sizes.

Here we briefly introduce the model design. Generating high-resolution images with diffusion models can be computationally expensive. Latent diffusion models (LDMs) address this by first downsampling images into a smaller latent embedding space using an encoder EE, and then training a diffusion model on z=E(x)z=E(x). New images are created by sampling zz from the model and decoding it back to images using a decoder x=D(z)x=D(z).

Similarly, SiT is also a latent generative model and we use the same pre-trained VAE encoder and decoder models originally used in Stable Diffusion . SiT processes a spatial input zz (shape 32×32×432\times 32\times 4 for 256×256×3256\times 256\times 3 images) by first ‘patchifying’ it into TT linearly embedded tokens of dimension dd. We always use a patch size of 2 in these models as they achieve the best sample quality. We then apply standard ViT sinusoidal positional embeddings to these tokens. We use a series of NN SiT transformer blocks, each with hidden dimension dd.

Our model configurations—SiT-{S,B,L,XL}—vary in model size (parameters) and compute (flops), allowing for a model scaling analysis. For class-conditional generation on ImageNet, we use the AdaLN-Zero block to process additional conditional information (times and class labels). SiT architectural details are listed in Table 3.

Experiments

To provide a more detailed answer to the question raised in Tab. 1 and make a fair comparison between DiT and SiT, we gradually transition from a DiT model (discretized, score prediction, VP interpolant) to a SiT model (continuous, velocity prediction, Linear interpolant) in the following four subsections, and present the impacts on performance. Throughout our experiments in each subsection, we uses a DiT-B model at 400K training steps as our backbone. For solving the ODE (2), we adopt a fixed Heun integrator; for solving the SDE (4), we used an Euler-Maruyama integrator. With both solver choices we limit the number of function evaluations (NFE) to be 250250 to match the number of sampling steps used in DiT. All numbers presented in the following sections are FID-50K scores evaluated on the ImageNet256 training set.

To understand the role of continuous-time versus discrete-time models, we study discrete-time DDPM against continuous-time SBDM-VP with estimation of the score. The results are presented in Table 4, where we find a marginal improvement in FID scores when going from a discrete-time denoiser to a continuous-time score field.

2 Model parameterizations

To clarify the role of the model parameterization in the context of SBDM-VP, we now compare learning (i) a score model using (6), (ii) a weighted score model (see Sec. A.3), or (iii) a velocity model using (7). The results are shown in Table 5, where we find that we obtain a significant performance improvement by learning a weighted score model or a velocity model.

3 Choice of interpolant

Section 2 highlights that there are many possible ways to build a connection between the data distribution and a Gaussian by varying the choice of αt\alpha_{t} and σt\sigma_{t} in the definition of the interpolant (1). To understand the role of this choice, we now study the benefits of moving away from the commonly-used SBDM-VP setup. We consider learning a velocity model v(x,t){\mathbf{v}}({\mathbf{x}},t) with the Linear and GVP interpolants presented in (12), which make the interpolation between the Gaussian and the data distribution exact on $$. We benchmark these models against the SBDM-VP in Table 6, where we find that both the GVP and Linear interpolants obtain significantly improved performance.

4 Deterministic vs stochastic sampling

As shown in Sec. 2, given a learned model, we can sample using either the probability flow equation (2) or an SDE (4). In SBDM we conventionally take as diffusion coefficient wt=βtw_{t}=\beta_{t}. For Linear and GVP interpolant, we follow the derivation in Appendix B to express βt\beta_{t} in terms of αt\alpha_{t} and σt\sigma_{t}.

Our results are shown in Tab. 7, where we find performance improvements by sampling with an SDE over the ODE, which is in line with the bounds given in : the SDE has better control over the KL divergence between the time-dependent density at t=0t=0 and the ground truth data distribution. We note that the performance of ODE and SDE integrators may differ under different computation budgets. As shown in Fig. 5, the ODE converges faster with fewer NFE, while the SDE is capable of reaching a much lower final FID score when given a larger computational budget.

Tunable diffusion coefficient.

Motivated by the improved performance of SDE sampling, we now consider the effect of tuning the diffusion coefficient in a manner that is distinct from the choices made in SBDM, as detailed in Tab. 2. As shown in Tab. 8, we find that the optimal choice for sampling is both model prediction and interpolant dependent. We picked the three functions with the best performance and sweep through all different combinations of our model prediction and interpolant, and present the result in Tab. 8.

We also note that the influences of different diffusion coefficients can vary across different model sizes. Empirically, we observe the best choice for our SiT-XL is a velocity model with Linear interpolant and sampled with σt\sigma_{t} coefficient.

5 Classifier-free guidance

Classifier-free guidance (CFG) often leads to improved performance for score-based models. In this section, we give a concise justification for adopting it on the velocity model, and then empirically show that the drastic gains in performance for DiT case carry across to SiT.

Guidance for a velocity field means that: (i) that the velocity model vθ(x,t;y){\mathbf{v}}_{\theta}({\mathbf{x}},t;{\mathbf{y}}) takes class labels yy during training, where yy is occasionally masked with a null token ∅\emptyset; and (ii) during sampling the velocity used is vθζ(x,t;y)=ζvθ(x,t;y)+(1−ζ)vθ(x,t;∅){\mathbf{v}}_{\theta}^{\zeta}({\mathbf{x}},t;{\mathbf{y}})=\zeta{\mathbf{v}}_{\theta}({\mathbf{x}},t;{\mathbf{y}})+(1-\zeta){\mathbf{v}}_{\theta}({\mathbf{x}},t;\emptyset) for a fixed ζ>0\zeta>0. In Appendix C, we show that this indeed corresponds to sampling the tempered density p(xt)p(y∣xt)ζp({\mathbf{x}}_{t})p({\mathbf{y}}|{\mathbf{x}}_{t})^{\zeta} as proposed in . Given this observation, one can leverage the usual argument for classifier-free guidance of score-based models.

For a CFG scale of ζ=1.5\zeta=1.5, DiT-XL sees an improvement in FID from 9.6 (non-CFG) down to 2.27 (CFG). We observed similar performance improvement with our largest SiT-XL model under identical computation budget and CFG scale. Sampled with an ODE, the FID-50K score improves from 9.4 to 2.15 (Tab. 9 and Appendix D); with an SDE, the FID improves from 8.6 to 2.06 (Tab. 1 and Tab. 9). This shows that SiT benefits from the same training and sampling choices explored previously, and can surpass DiTs performance in each training setting, not only with respect to model size, but also with respect to sampling choices.

Related Work

The transformer architecture has emerged as a powerful tool for application domains as diverse as vision , language , quantum chemistry , active matter systems , and biology . Several works have built on DiT and have made improvements by modifying the architecture to internally include masked prediction layers ; these choices are orthogonal to the transition from DiT to SiT studied in this work; they may be fruitfully combined in future work.

Training and Sampling in Diffusions.

Diffusion models arose from and have close historical relationship with denoising methods . Various efforts have gone into improving the sampling algorithms behind these methods in the context of DDPM and SBDM ; these are also orthogonal to our studies and may be combined to push for better performance in future work. Improved Diffusion ODE also studies several combinations of model parameterizations (velocity versus noise) and paths (VP versus Linear) for sampling an ODE; they report best results for velocity model with smoother probability flow; they focus on lower dimensional experiments, benchmark with likelihoods, and do not consider SDE sampling. In our work, we explore the effects of changing between VP, Linear, and GVP interpolants, as well as score and velocity parameterizations in depth and show how these choices individually improve performance on the larger scale ImageNet256. We also document how FIDs change with respect to a family of sampling algorithms including black-box ODEs and SDEs indexed by a choice of diffusion coefficients, and show that the best coefficient choice may depend on the model and interpolant. This brings the observations about the flexibility and trade-offs of sampling from into practice.

Interpolants and flow matching.

Velocity field parameterizations using the Linear interpolant were also studied in , and were generalized to the manifold setting in . A trade-off in bounds on the KL divergence between the target distribution and the model arises when considering sampling with SDEs versus ODE; shows that minimizing the objectives presented in this work controls KL for SDEs, but not for ODEs. Error bounds for SDE-based sampling with score-based diffusion models are studied in . Error bounds on the ODE are also explored in , in addition to the Wasserstein bounds provided in .

Other related works make improvements by changing how noise and data are sampled during training. compute mini-batch optimal couplings between the Gaussian and data distribution to reduce the transport cost and gradient variance; instead build the coupling by flowing directly from the conditioning variable to the data for image-conditional tasks such as super-resolution and in-painting. Finally, various work considers learning a stochastic bridge connecting two arbitrary distributions . These directions are compatible with our investigations; they specify the learning problem for which one can then vary the choices of model parameterizations, interpolant schedules, and sampling algorithms.

Diffusion in Latent Space.

Generative modeling in latent space is a tractable approach for modeling high-dimensional data. The approach has been applied beyond images to video generation , which is a yet-to-be explored and promising application area for velocity trained models. also train velocity models in the latent space of the pre-trained Stable Diffusion VAE. They demonstrate promising results for the DiT-B backbone with a final FID-50K of 4.46; their study was one motivation for the investigation in this work regarding which aspects of these models contribute to the gains in performance over DiT.

Conclusion

In this work, we have presented Scalable Interpolant Transformers, a simple and powerful framework for image generation tasks. Within the framework, we explored the tradeoffs between a number of key design choices: the choice of a continuous or discrete-time model, the choice of interpolant, the choice of model prediction, and the choice of diffusion coefficient. We highlighted the advantages and disadvantages of each choice and demonstrated how careful decisions can lead to significant performance improvements. Many concurrent works explore similar approaches in a wide variety of downstream tasks, and we leave the application of SiT to these tasks for future works.

We would like to thank Adithya Iyer, Sai Charitha Akula, Fred Lu, Jitao Gu, and Edwin P. Gerber for helpful discussions and feedback. The research is partly supported by the Google TRC program.

References

Appendix A Proofs

where ∇x⋅[vpt]=∑i=1d∂∂xi[vipt]\nabla_{\mathbf{x}}\cdot[{\mathbf{v}}p_{t}]=\sum_{i=1}^{d}\frac{\partial}{\partial x_{i}}[v_{i}p_{t}] is the divergence operator and we take advantage of the divergence theorem and used integration by parts to get the second equality. By the properties of Fourier transform, Eq. 23 implies that pt(x)p_{t}({\mathbf{x}}) satisfies the transport equation

Solving this equation by the method of characteristic leads to probability flow ODE (2).

A.2 Proof of the SDE (4)

We show that the SDE (4) has marginal density pt(x)p_{t}({\mathbf{x}}) with any choice of wt≥0w_{t}\geq 0. To this end, recall that solution to the SDE

has a PDF that satisfies the Fokker-Planck equation

where Δx\Delta_{{\mathbf{x}}} is the Laplace operator defined as Δx=∇x⋅∇x=∑i=0d∂2∂xi2\Delta_{\mathbf{x}}=\nabla_{\mathbf{x}}\cdot\nabla_{\mathbf{x}}=\sum_{i=0}^{d}\frac{\partial^{2}}{\partial x_{i}^{2}}. Reorganizing the equation and usng the definition of the score s(x,t)=∇xlog⁡pt(x)=pt−1(x)∇xpt(x){\mathbf{s}}({\mathbf{x}},t)=\nabla_{{\mathbf{x}}}\log p_{t}({\mathbf{x}})=p_{t}^{-1}({\mathbf{x}})\nabla_{{\mathbf{x}}}p_{t}({\mathbf{x}}), we have

By definition of Laplace operator, the last equation holds for any wt≥0w_{t}\geq 0. When wt=0w_{t}=0, the Fokker-Planck equation reduces to a continuity equation, and the SDE reduces to an ODE, so the connection trivially holds.

A.3 Proof of the expression for the score in Eq. 5

Since ε∼N(0,I){\boldsymbol{\varepsilon}}\sim\mathsf{N}(0,\mathbf{I}), we can compute the expectation explicitly to obtain

Since x∗{\mathbf{x}}_{*} and ε{\boldsymbol{\varepsilon}} are independent random variable, we have

where p^t(k)\hat{p}_{t}({\mathbf{k}}) is the characteristic function of xt=αtx∗+σtε{\mathbf{x}}_{t}=\alpha_{t}{\mathbf{x}}_{*}+\sigma_{t}{\boldsymbol{\varepsilon}} defined in Eq. 16. The left hand-side of this equation can also be written as:

where we again used divergence theorem and integration by parts to get the third equality, and again the definition of the score to get the last. Comparing Eq. 33 and Eq. 37 we deduce that, when σt≠0\sigma_{t}\not=0,

Further, setting wtw_{t} to σt\sigma_{t} in Eq. 4 gives

for all t∈t\in. This bypass the constraint of σt≠0\sigma_{t}\neq 0 and effectively eliminate the singularity at t=0t=0.

A.4 Proof of Eq. 9

We note that there exists a straightforward connection between v(x,t){\mathbf{v}}({\mathbf{x}},t) and s(x,t){\mathbf{s}}({\mathbf{x}},t). From Eq. 1, we have

Given Eq. 44 is linear in terms of s{\mathbf{s}}, reverting it will lead to Eq. 9.

Appendix B Connection with Score-based Diffusion

As shown in Song et al. , the reverse-time SDE from Eq. 10 is

Let us show this SDE is Eq. 4 for the specific choice wt=βtw_{t}=\beta_{t}. To this end, notice that the solution Xt{\mathbf{X}}_{t} to Eq. 51 for the initial condition Xt=0=x∗{\mathbf{X}}_{t=0}={\mathbf{x}}_{*} with x∗{\mathbf{x}}_{*} fixed is Gaussian distributed with mean and variance given respectively by

Using Eq. 44, the velocity of the score-based diffusion model can therefore be expressed as

we see that 2λtσt2\lambda_{t}\sigma_{t} is precisely βt\beta_{t}, making λt\lambda_{t} correspond to the square of maximum likelihood weighting proposed in Song et al. . Further, if we plug Eq. 55 into Eq. 4 with wt=βtw_{t}=\beta_{t}, we arrive at Eq. 51.

Appendix C Sampling with Guidance

Let pt(x∣y)p_{t}({\mathbf{x}}|{\mathbf{y}}) be the density of xt=αtx∗+σtε{\mathbf{x}}_{t}=\alpha_{t}{\mathbf{x}}_{*}+\sigma_{t}{\boldsymbol{\varepsilon}} conditioned on some extra variable y{\mathbf{y}}. By argument similar to the one given in Sec. A.1, it is easy to see that pt(x∣y)p_{t}({\mathbf{x}}|{\mathbf{y}}) satisfies the transport equation (compare Eq. 24)

Proceeding as in Sec. A.3 and Sec. A.4, it is also easy to see that the score s(x,t∣y)=∇xlog⁡pt(x∣y){\mathbf{s}}({\mathbf{x}},t|{\mathbf{y}})=\nabla_{{\mathbf{x}}}\log p_{t}({\mathbf{x}}|{\mathbf{y}}) is given by (compare Eq. 5)

and that v(x,t∣y){\mathbf{v}}({\mathbf{x}},t|{\mathbf{y}}) and s(x,t∣y){\mathbf{s}}({\mathbf{x}},t|{\mathbf{y}}) are related via (compare Eq. 44)

where we have used the fact ∇xlog⁡pt(x∣y)=∇xlog⁡pt(y∣x)+∇xlog⁡pt(x)\nabla_{{\mathbf{x}}}\log p_{t}({\mathbf{x}}|{\mathbf{y}})=\nabla_{{\mathbf{x}}}\log p_{t}({\mathbf{y}}|{\mathbf{x}})+\nabla_{{\mathbf{x}}}\log p_{t}({\mathbf{x}}) that follows from pt(x∣y)p(y)=pt(y∣x)pt(x)p_{t}({\mathbf{x}}|{\mathbf{y}})p({\mathbf{y}})=p_{t}({\mathbf{y}}|{\mathbf{x}})p_{t}({\mathbf{x}}), and ζ\zeta to be some constant greater than 11. Eq. 64 shows that using the score mixture sζ(x,t∣y)=(1−ζ)s(x,t)+ζs(x,t∣y){\mathbf{s}}^{\zeta}({\mathbf{x}},t|{\mathbf{y}})=(1-\zeta){\mathbf{s}}({\mathbf{x}},t)+\zeta{\mathbf{s}}({\mathbf{x}},t|{\mathbf{y}}), and the velocity mixture associated with it, namely,

allows one to to construct generative models that sample the tempered distribution pt(xt)ptζ(y∣xt)p_{t}({\mathbf{x}}_{t})p^{\zeta}_{t}({\mathbf{y}}|{\mathbf{x}}_{t}) following classifier guidance . Note that pt(x)ptζ(y∣x)∝ptζ(x∣y)pt1−ζ(x)p_{t}({\mathbf{x}})p^{\zeta}_{t}({\mathbf{y}}|{\mathbf{x}})\propto p^{\zeta}_{t}({\mathbf{x}}|{\mathbf{y}})p^{1-\zeta}_{t}({\mathbf{x}}), so we can also perform classifier free guidance sampling . Empirically, we observe significant performance boost by applying classifier free guidance, as showed in Tab. 1 and Tab. 9.

Appendix D Sampling with ODE and SDE

In the main body of the paper, we used a Heun integrator for solving the ODE in Eq. 2 and an Euler-Maruyama integrator for solving the SDE in Eq. 4. We summarize all results in Tab. 1, and present the implementations below.

It is feasible to use either a velocity model vθ{\mathbf{v}}_{\theta} or a score model sθ{\mathbf{s}}_{\theta} in applying the above two samplers. If learning the score for the deterministic Heun sampler, we could always convert the learned sθ{\mathbf{s}}_{\theta} to vθ{\mathbf{v}}_{\theta} following Sec. A.4. However, as there exists potential numerical instability (depending on interpolants) in σ˙t\dot{\sigma}_{t}, αt−1\alpha_{t}^{-1} and λt\lambda_{t}, it’s recommended to learn vθ{\mathbf{v}}_{\theta} in sampling with deterministic sampler instead of sθ{\mathbf{s}}_{\theta}. For the stochastic sampler, it’s required to have both vθ{\mathbf{v}}_{\theta} and sθ{\mathbf{s}}_{\theta} in integration, so we always need to convert from one (either learning velocity or score) to obtain the other. Under this scenario, the numerical issue from Sec. A.4 can only be avoided by clipping the time interval near t=0t=0. Empirically we found clipping the interval by h=0.04h=0.04 and doing a long last step from t=0.04t=0.04 to can greatly benefit the performance. A detailed summary of sampler configuration is provided in Appendix E.

Additionally, we could replace vθ{\mathbf{v}}_{\theta} and sθ{\mathbf{s}}_{\theta} by vθζ{\mathbf{v}}_{\theta}^{\zeta} and sθζ{\mathbf{s}}_{\theta}^{\zeta} presented in Appendix C as inputs of the two samplers and enjoy the performance improvements coming along with guidance. As guidance requires evaluating both conditional and unconditional model output in a single step, it will impose twice the computational cost when sampling.

We primarily investigate and report the performance comparison between DDPM and Euler-Maruyama samplers. We set our Euler sampler’s number of steps to be 250 to match that of DDPM during evaluation. This comparison is made direct and fair, as the DDPM method is equivalent to a discretized Euler’s method.

Comparison between DDIM and Heun

We also investigate the performance difference produced by deterministic samplers between DiT and our models. In Fig. 1, we show the FID-50K results for both DiT models sampled with DDIM and SiT models sampled with Heun. We note that this is not directly an apples-to-apples comparison, as DDIM can be viewed as a discretized version of the first order Euler’s method, while we use the second order Heun’s method in sampling SiT models, due to the large discretization error with Euler’s method in continuous time. Nevertheless, we control the NFEs for both DDIM (250 sampling steps) and Heun (250 NFE).

Higher order solvers

The performances of an adaptive deterministics dopri5 solver and a second order stochastic Heun Sampler are also tested. For dopri5, we set atol and rtol to 1e-6 and 1e-3, respectively; for Heun, we again maintain the NFE to be 250 to match that of DDPM. In both solvers we do not observe performance increment; under the CFG scale of ζ=1.5\zeta=1.5, dopri5 and stochastic Heun gives FID-50K of 2.15 and 2.07, respectively.

We also note that our models are compatible with other samplers specifically tuned to diffusion models as well as sampling distillation . We do not include the evaluations of those methods in our work for the sake of apples-to-apples comparison with the DDPM model, and we leave the investigation of potential performance improvements to future work.

Appendix E Additional Implementation Details

We implemented our models in JAX following the DiT PyTorch codebase by Peebles and Xie https://github.com/facebookresearch/DiT, and referred to Albergo et al. https://github.com/malbergo/stochastic-interpolants, Song et al. https://github.com/yang-song/score_sde, and Dockhorn et al. https://github.com/nv-tlabs/CLD-SGM for our implementation of the Euler-Maruyama sampler. For the Heun sampler, we directly used the one from diffrax https://github.com/patrick-kidger/diffrax, a JAX-based numerical differential equation solver library.

We trained all of our models following identical structure and hyperparameters retained from DiT . We used AdamW as optimizer for all models. We use a constant learning rate of 1×10−41\times 10^{-4} and a batch size of 256256. We used random horizontal flip with probability of 0.50.5 in data augmentation. We did not tune the learning rates, decay/warm up schedules, AdamW parameters, nor use any extra data augmentation or gradient clipping during training. Our largest model, SiT-XL, trains at approximately 6.86.8 iters/sec on a TPU v4-64 pod following the above configurations. This speed is slightly faster compared to DiT-XL, which trains at 6.46.4 iters/sec under identical settings. We also gather the training speed of other model sizes and summarize them below.

Sampling configurations

We maintain an exponential moving average (EMA) of all models weights over training with a decay of 0.99990.9999. All results are sampled from the EMA checkpoints, which is empirically observed to yield better performance. We summarize the start and end points of our deterministic and stochastic samplers with different interpolants below, where each t0t_{0} and tNt_{N} are carefully tuned to optimize performance and avoid numerical instability during integration.

FID calculation

We calculate FID scores between generated images (10K or 50K) and all available real images in ImageNet training dataset. We observe small performance variations between TPU-based FID evaluation and GPU-based FID evaluation (ADM’s TensorFlow evaluation suite https://github.com/openai/guided-diffusion/tree/main/evaluations). To ensure consistency with the basline DiT, we sample all of our models on GPU and obtain FID scores using the ADM evaluation suite.

Appendix F Additional Visual results