Fast ODE-based Sampling for Diffusion Models in Around 5 Steps
Zhenyu Zhou, Defang Chen, Can Wang, Chun Chen
Introduction
Diffusion models have been attracting growing attentions in recent years due to their impressive generative capability . Given a noise input, they are able to generate a realistic output by performing iterative denoising steps with the score function . This process can be interpreted as applying a certain numerical discretization on a stochastic differential equation (SDE), or more commonly, its corresponding probability flow ordinary differential equation (PF-ODE) . Comparing to other generative models such as GANs and VAEs , diffusion models have the advantages in high sample quality and stable training, but suffer from slow sampling speed, which poses a great challenge to their applications.
Existing methods for accelerating diffusion sampling fall into two main streams. One is designing faster numerical solvers to increase step size while maintaining small truncation errors . They can be further categorized as single-step solvers and multi-step solvers . The former computes the next step solution only using information from the current time step, while the latter uses multiple past time steps. These methods have successfully reduced the number of function evaluations (NFE) from 1000 to less than 20, almost without affecting the sample quality.
Another kind of methods aim to build a one-to-one mapping between the data distribution and the pre-specified noise distribution , based on the idea of knowledge distillation. With a well-trained student model in hand, high-quality generation can be achieved with only one NFE. However, training such a student model either requires pre-generation of millions of images , or huge training cost with carefully modified training procedure . Besides, distillation-based models cannot guarantee the increase of sample quality given more NFE and they have difficulty in likelihood evaluation.
In this paper, we further boost ODE-based sampling for diffusion models in around 5 steps. Based on the geometric property that each sampling trajectory approximately lies on a two-dimensional subspace embedded in the high-dimension space, we propose Approximate MEan-Direction Solver (AMED-Solver), a single-step ODE solver that learns to predict the optimal intermediate time step in every sampling step. A comparison of various ODE solvers is illustrated in Fig. 2. We also extend our method to any ODE solvers as a plugin. Based on the previous observation that the sampling trajectory generated by diffusion models is nearly straight , we perform analytical first step (AFS) to save one NFE almost without affecting the sample quality. When applying AMED-Plugin on the improved PNDM method, we achieve FID of 7.14 on CIFAR-10, 13.75 on ImageNet 6464, and 12.79 on LSUN Bedroom. Our main contributions are as follows:
We introduce AMED-Solver, a new single-step ODE solver for diffusion models that eliminates truncation errors by design.
We propose AMED-Plugin that can be applied to any ODE solvers with a small training overhead and a negligible sampling overhead.
Extensive experiments on various datasets validate the effectiveness of our method in fast image generation.
Background
The forward diffusion process can be formalized as a SDE:
sharing the same marginals with the reverse SDE , and is known as the score function . Generally, this PF-ODE is preferred in practice for its conceptual simplicity, sampling efficiency and unique encoding . Throughout this paper, we follow the configuration of EDM by setting , and . In this case, the reciprocal of equals to the signal-to-noise ratio and the perturbation kernel is
To simulate the PF-ODE, we usually train a U-Net predicting to approximate the intractable . There are mainly two parameterizations in the literature. One uses a noise prediction model predicting the Gaussian noise added to at time , and another uses a data prediction model predicting the denoising output of from time to . They have the following relationship in our setting:
The training of diffusion models in the noise prediction notation is performed by minimizing a weighted combination of the least squares estimations:
We then plug the learned score function Eq. 4 into Eq. 2 to obtain a simple formulation for the PF-ODE
The sampling trajectory is obtained by first draw and then solve Eq. 6 with steps following time schedule .
2 Categorization of Previous Fast ODE Solvers
To accelerate diffusion sampling, various fast ODE solvers have been proposed, which can be classified into single-step solvers or multi-step solvers . Single-step solvers including DDIM , EDM and DPM-Solver which only use the information from the current time step to compute the solution for the next time step, while multi-step solvers including PNDM and DEIS utilize multiple past time steps to compute the next time step (see Fig. 2 for an intuitive comparison). We emphasize that one should differ single-step ODE solvers from single-step (NFE=1) distillation-based methods .
The advantages of single-step methods lie in the easy implementation and the ability to self-start since no history record is required. However, as illustrated in Fig. 3, they suffer from fast degradation of sample quality especially when the NFE budget is limited. The reason may be that the actual sampling steps of multi-step solvers are twice as much as those of single-step solvers with the same NFE, enabling them to adjust directions more frequently and flexibly. We will show that our proposed AMED-Solver can largely fix this issue with learned intermediate time steps.
Our Proposed AMED-Solver
In this section, we propose AMED-Solver, a single-step ODE solver for diffusion models that releases the potential of single-step solvers in extremely small NFE, enabling them to match or even surpass the performance of multi-step solvers. Furthermore, our proposed method can be generalized as a plugin on any ODE solver, yielding promising improvement across various datasets. Our key observation is that the sampling trajectory generated by Eq. 6 nearly lies in a two-dimensional subspace embedded in high-dimensional space, which motivates us to minimize the discretization error with the mean value theorem.
The sampling trajectory generated by solving Eq. 6 exhibits an extremely simple geometric structure and has an implicit connection to the annealed mean shift, as revealed in the previous work . Each sample starting from the noise distribution approaches the data manifold along a smooth, quasi-linear sampling trajectory in a monotone likelihood increasing way . Besides, all trajectories from different initial points share the similar geometric shape, whether in the conditional or unconditional generation case.
In this paper, we further point out that the sampling trajectory generated by ODE solvers almost lies in a two-dimensional plane embedded in a high-dimensional space. We experimentally validate this claim by performing Principal Component Analysis (PCA) for 1k sampling trajectories on different datasets including CIFAR10 3232 , FFHQ 6464 , ImageNet 6464 and LSUN Bedroom 256256 . As illustrated in Fig. 4, the relative projection error using two principal components is no more than 8 and always stays in a small level. Besides, the sample variance can be fully explained using only two principal components. Given the vast image space with dimensions of 3072(=3), 12288(=3), or 196608(=3), the sampling trajectories show intriguing property that their dynamics can almost be described using only two principal components.
2 Approximate Mean-Direction Solver
With the geometric intuition, we next explain our methods in more detail. The exact solution of Eq. 6 is:
Various numerical approximations to the integral above correspond to different types of fast ODE-based samplers. For instance, direct explicit rectangle method gives DDIM , linear multi-step method yields PNDM , Taylor expansion yields DPM-Solver and polynomial extrapolation recovers DEIS . Different from these works, we derive our method more directly by expecting that the mean value theorem holds for the integral involved so that we can find an intermediate time step with
Although the well-known mean value theorem for real-valued functions does not hold in vector-valued case , the remarkable geometric property that the sampling trajectory almost lies in a two-dimensional subspace guarantees our use. By choosing the intermediate time step, we are able to achieve an approximation of Eq. 7 by
The formulation above is actually a single-step ODE solver. The typical DPM-Solver-2 can be recovered by setting . In Tab. 1, we compare various single-step solvers by generalizing as the gradient term.
For the choose of intermediate time steps, we train a shallow neural network based on distillation with small training and negligible sampling costs. Briefly, given samples on the teacher sampling trajectory with higher NFE budget and on the student sampling trajectory, gives the intermediate time step that minimizes where is given by Eq. 9 and is a distance metric. Since we seek to find a mean-direction that best approximates the integral in Eq. 8, we name our proposed single-step ODE solver Eq. 9 Approximate MEan-Direction Solver, dubbed as AMED-Solver. Before specifying the training and sampling details, we proceed to generalize our idea to any fast ODE solvers.
3 AMED as A Plugin
The idea of AMED can be used to further improve existing fast ODE solvers for diffusion models. Throughout the paper, we adapt the polynomial time schedule :
Note that another usually used uniform logSNR schedule is actually the limit of Eq. 10 as approaches .
Given a time schedule , the AMED-Solver is obtained by performing extra model evaluations at . Under the same manner, we are able to improve any ODE solvers by predicting that best aligns the student and teacher sampling trajectories.
We validate this by an experiment where we first generate a ground truth trajectory using Heun’s second method with 80 NFE and extract samples at to get . For every ODE solver, we generate a baseline trajectory by performing evaluations at following the original time schedule. We then apply a grid search on , giving and a searched trajectory . We define the relative alignment to be . As shown in Fig. 5, the relative alignment keeps positive in most cases, meaning that appropriate choose of intermediate time steps can further improve fast ODE solvers. Therefore, as described in Sec. 3.2, we can also train a neural network to appropriately predict intermediate time steps. As this process still has the meaning of searching for the direction pointing to the ground truth, we name this method as AMED-Plugin and apply it on various fast ODE solvers.
Through Fig. 5, we obtain a direct comparison between DPM-Solver-2 and our proposed AMED-Solver since they share the same baseline trajectory when is fixed to . The AMED-Solver provides better alignment to the ground truth and the location of searched intermediate time steps is much more stable than that of DPM-Solver-2. We speculate that for DPM-Solver-2, the direction of gradient term is restricted by the combination where the coefficients is fixed by (see Tab. 1). Instead, the gradient term of AMED-Solver is directly given by the intermediate time step, providing more flexibility than DPM-Solver-2.
4 Training and Sampling
As samples from different sampling trajectories approach the asymmetric data manifold from different directions, the location of current points should contribute to the shape of the sampling trajectories. To recognize the location of one sample without extra computation overhead, we extract the bottleneck feature of the pre-trained U-Net model every time after its evaluation. We then take the current time step along with the bottleneck feature as the inputs of the shallow neural network to predict the intermediate time step . Formally, we have
The network architecture is shown in Fig. 6.
When sampling from time to , we first perform one U-Net evaluation at and extract the bottleneck feature to predict and step to . For AMED-Solver, another U-Net evaluation at gives the gradient term and derives with Eq. 9. When applying AMED-Plugin on other ODE solvers, we step from to following the solver’s original sampling procedure. The total NFE is thus . We denote such a sampling process by
where is the set of intermediate time steps injected in , .
The training of is based on distillation, where student and teacher sampling trajectories evaluated at are required. Denote the sampling process that generates student and teacher trajectories by with and with respectively. For , since teacher trajectories require more NFE to give reliable reference, we set it to be a smooth interpolation of intermediate time steps between and following the original polynomial time schedule, i.e., where
Denote the student sampling trajectory and teacher sampling trajectory extracted at by and , we train using a distance metric between samples on both trajectories with predicted by :
In one training loop, we first generate a batch of noise images at and then the teacher trajectories. We then calculate the loss progressively from time to . Hence, backpropagations are applied in one training loop. We specify the algorithms for training and sampling in Algorithm 1 and Algorithm 2.
Consistent with the geometric property that the sampling trajectory is nearly straight when is large , we notice that the gradient term at time shares almost the same direction as . We thus simply use as the direction in the first sampling step to save one NFE, which is important when the budget is limited. This trick is called analytical first step (AFS) . We find that the application of AFS yields little degradation or even increase of sample quality for datasets with small resolutions.
5 Comparing with Distillation-based Methods
Though being a solver-based method, our proposed AMED-Solver shares a similar principal with distillation-based methods. The main difference is that distillation-based methods finally build a mapping from noise to data distribution by fine-tuning the pre-trained model or training a new prediction model from scratch, while our AMED-Solver still follows the nature of solving an ODE, building a flow of distributions from noise to image. Distillation-based methods have shown impressive results of performing high-quality generation by only one NFE. However, these methods require huge efforts on training. One should carefully design the training details and it takes a large amount of time to train the model (usually several or even tens of GPU days). Moreover, as distillation-based models directly build the mapping like typical generative models, they suffer from the inability of interpolating between two disconnected modes . Therefore, distillation-based methods may fail in some downstream tasks requiring such an interpolation.
The training of AMED-Solver is also related to distillation, but the training goal is to properly predict the location of intermediate time steps instead of directly predicting high-dimensional samples at next time. Therefore, our architecture is very simple and is easy to train given the geometric property of sampling trajectories. The solver we obtain still maintain the nature of solving an ODE like solver-based methods, which do not suffer from obvious internal imperfection for downstream tasks.
Experiments
Datasets. We employ AMED-Solver and AMED-Plugin on a wide range of datasets with image resolutions ranging from 32 to 512, including CIFAR10 3232 , FFHQ 6464 , ImageNet 6464 and LSUN Bedroom 256256 . We also give qualitative results generated by stable-diffusion with resolution of 512.
Models. The pre-trained models are pixel-space models from and and latent-space model from .
Solvers. We reimplement several representative fast ODE solvers including DDIM , DPM-Solver-2 , multi-step DPM-Solver++ , UniPC and improved PNDM (iPNDM) . It is worth mentioning that we find iPNDM achieves very impressive results and outperforms other ODE solvers in many cases.
Time schedule. We mainly use the polynomial schedule with , which is the default setting in , except for DPM-Solver++ and UniPC where we use logSNR schedule recommended in original papers for better results.
Training. We train for 20k images, which takes 5-20 minutes on CIFAR10 using a single NVIDIA A100 GPU. For datasets with resolution of 256, it takes 1-2 hours on 4 NVIDIA A100 GPUs. For the distance metric in Eq. 14, we use L2 norm in most experiments except for those on CIFAR10 where we use LPIPS for slightly better performance. For the generation of teacher trajectories for AMED-Solver, we use DPM-Solver-2 or EDM with doubled NFE. For AMED-Plugin, we use the same solver that generates student trajectories with or (we provide a detailed discussion in Sec. C.2).
Sampling. Due to designation, our AMED-Solver or AMED-Plugin naturally create solvers with even NFE. Once AFS is used, the total NFE becomes odd. With the goal of designing fast solvers in small NFE, we mainly test our method on NFE where AFS is applied. There are also results on NFE without AFS.
Evaluation. We measure the sample quality via Fréchet Inception Distance (FID) with 50k samples.
2 Image Generation
In this section, we show the results of image generation on various datasets. For datasets with small resolution such as CIFAR10 3232, FFHQ 6464 and ImageNet 6464, we report the results of AMED-Plugin applied on iPNDM solver for its leading results. For large-resolution datasets like LSUN Bedroom, we report the results of AMED-Plugin applied on DPM-Solver-2 since the use of AFS causes inferior results for multi-step solvers in this case (see Sec. C.3 for a detailed discussion). We implement DPM-Solver++ and UniPC with order of 3, iPNDM with order of 4. To report results of DPM-Solver-2 and EDM with odd NFE, we apply AFS in their first steps.
The results are shown in Tab. 2. Our AMED-Solver exhibits promising results and outperforms other single-step methods. Besides, the AMED-Solver performs on par with multi-step solvers in datasets with small resolution, and even beats them in large-resolution cases. For the AMED-Plugin, we find it showing large boost when applied on various solvers especially for DPM-Solver-2 as shown in Tab. 2(d). Notably, the AMED-Plugin improves the FID by 6.45, 5.24 and 4.63 on CIFAR10 3232, ImageNet 6464 and FFHQ 6464 in 5 NFE. Our methods achieve state-of-the-art results in solver-based methods in around 5 NFE.
In Fig. 7, we show the learned coefficient of for AMED-Solver, where each is predicted by and . The dashed line recovers the default setting of DPM-Solver-2. We include more quantitative as well as qualitative results in Appendix C.
3 Ablation Study
Sample-wise and Time-wise Intermediate Steps. As the U-Net bottleneck input to varies for different samples, the learned intermediate time steps are sample-wise, meaning that different trajectories have different time schedules. Since sampling trajectories from different starting points share similar geometric shapes, the effectiveness of inputting bottleneck might be limited. We also notice that during training, the standard deviation of learned intermediate time steps in one batch is small. Therefore, it should be safe to replace the U-Net bottleneck with zero matrix and get time-wise intermediate steps where is shared across all sampling trajectories. In Tab. 3, we provide results of sample-wise intermediate time steps learned with U-Net bottleneck as input, time-wise steps with zero matrix as input, and optimal time-wise steps obtained by a grid search of 20k samples from 0.02 to 0.98 with step size of 0.02. The results show the efficiency of AMED-Solver and again confirm the aforementioned geometric property.
Teacher solvers. In image generation, we simply let the student and teacher ODE solver to be the same, but given the property of unique encoding of ODE solvers, any ODE solver is competent to be the teacher. In Tab. 4, we use different teacher solvers for the training of to obtain different AMED-Solvers. It turns out that the best results are achieved when teacher solver resembles the student solver in the sampling process. We hence recommend that one should keep the teacher and student solvers to be the same.
Loss metric. As shown in Fig. 8, the use of LPIPS metric in the training loss Eq. 5 provides slightly better results on CIFAR10. However, applying LPIPS on datasets with large resolution or in conditional case yields worse results. We use L2 metric in these cases.
Time schedule. During experiments, we observe that different fast ODE solvers have different preference on time schedules. This preference even depends on the dataset. We conduct experiments on time schedules for DPM-Solver++(3M) on CIFAR10 in Tab. 5 and find that it prefers logSNR schedule where the interval between the first and second time step is larger than other schedules. We also show the results of AMED-Plugin applied on these cases. The AMED-Plugin largely and consistently improves the results no matter what the time schedule is.
Conclusion
In this paper, we introduce a single-step ODE solver called AMED-Solver to minimize the discretization error for fast diffusion sampling. Our key observation is that each sampling trajectory generated by existing ODE solvers approximately lies on a two-dimensional subspace, and thus the mean value theorem guarantees the learning of such a mean direction. The AMED-Solver effectively mitigates the problem of rapid sample quality degradation, which is commonly encountered in single-step methods, and shows promising results on datasets with large resolution. We also generalize the idea of AMED-Solver to AMED-Plugin, a plugin that can be applied on any fast ODE solvers to further improve the sample quality. We validate our methods through extensive experiments and achieve state-of-the-art results in extremely small NFE (around 5). We hope our attempt can inspire future works to further release the potential of fast solvers for diffusion models.
Limitation and Future Work. Fast ODE solvers for diffusion models are highly sensitive to time schedules especially when the NFE budget is limited (See Tab. 5). Through experiments, we observe that any fixed time schedule fails to perform well in all situations. In fact, our AMED-Plugin can be treated as adjusting half of the time schedule to partly alleviates but not avoids this issue. We believe that a better designation of time schedule requires further knowledge to the geometric shape of the sampling trajectory . We leave this to the future work.
References
Appendix A Related Works
Ever since the birth of diffusion models , their generation speed has become a major drawback compared to other generative models . To address this issue, efforts have been taken to accelerate the sampling of diffusion models, which fall into two main streams.
One is designing faster solvers. In early works , the authors speed up the generation from 1000 to less than 50 NFE by reducing the number of time steps with systematic or quadratic sampling. Analytic-DPM , provides an analytic form of optimal variance in sampling process and improves the results. More recently, with the knowledge of interpreting the diffusion process as a PF-ODE , there is a class of ODE solvers based on numerical methods that accelerate the sampling process to around 10 NFE. The authors in EDM achieve several improvements on training and sampling of diffusion models and propose to use Heun’s second method. PDNM uses linear multi-step method to solve the PF-ODE with Runge-Kutta algorithms for the warming start. The authors in recommend to use lower order linear multi-step method for warming start and propose iPNDM. Given the semi-linear structure of the PF-ODE, DPM-Solver and DEIS are proposed by approximate the integral involved in the analytic solution of PF-ODE with Taylor expansion and polynomial extrapolation respectively. DPM-Solver is further extend to both single-step and multi-step methods in DPM-Solver++ . UniPC gives a unified predictor-corrector solver and improved results compared to DPM-Solver++.
Besides training-free fast solvers above, there are also solvers requiring additional training. proposes a series of reparameterization for a generalized family of DDPM with a KID loss. More related to our method, in GENIE , the authors apply the second truncated Taylor method to the PF-ODE and distill a new model to predict the higher-order gradient term. Different from this, in AMED-Solver, we train a network that only predict the intermediate time steps, instead of a high-dimensional output.
Another mainstream is training-based distillation methods, which attempt to build a direct mapping from noise distribution to implicit data distribution. This idea is first introduced in as an offline method where one needs to pre-construct a dataset generated by the original model. Rectified flow also introduces an offline distillation based on optimal transport. For online distillation methods, one can progressively distill a diffusion model from more than 1k steps to 1 step , or utilize the consistency property of PF-ODE trajectory to tune the denoising output .
Appendix B Experimental Details
Datasets. We employ AMED-Solver and AMED-Plugin on a wide range of datasets and settings. We report results on settings including unconditional generation in both pixel and latent space, conditional generation with or without guidance. Datasets are chosen with image resolutions ranging from 32 to 512, including CIFAR10 3232 , FFHQ 6464 , ImageNet 6464 and LSUN Bedroom 256256 . We give qualitative results generated by stable-diffusion with resolution of 512. To further evaluate the effectiveness of our methods, in Appendix C, we also include more results on LSUN Cat 256256 , ImageNet 256256 with classifier guidance, and latent-space LSUN Bedroom 256256 .
Models. The pre-trained models we use throughout our experiments are pixel-space models from and as well as , and latent-space models from . Our code architecture is mainly based on the implementation in .
Fast ODE solvers. To give fair comparison, we reimplement several representative fast ODE solvers including DDIM , DPM-Solver-2 , multi-step DPM-Solver++ , UniPC and improved PNDM (iPNDM) . Through the implementation, we obtain better or on par FID results compared with original papers. It is worth mentioning that during the implementation, we find iPNDM achieves very impressive results and outperforms other ODE solvers in many cases.
Time schedule. We notice that different ODE solvers have different preference on time schedule. We mainly use the polynomial time schedule with , which is the default setting in , except for DPM-Solver++ and UniPC where we use logSNR time schedule recommended in original papers for better results. As illustrated in Sec. 3.4, the logSNR schedule is the limit case of polynomial schedule as approaches . For experiments on ImageNet with guidance, stable-diffusion and latent-space LSUN Bedroom, uniform time schedule gives best results. The reason may lie in their different training process.
Training. Since the number of parameters is small, the training of does not cause much computational cost. The training spends its main time on generating student and teacher trajectories. We train for 20k images, which takes 5-20 minutes for experiments on CIFAR10 using a single NVIDIA A100 GPU. For datasets with resolution of 256, it takes 1-2 hours on 4 NVIDIA A100 GPUs. For the distance metric in Eq. 14, we use L2 norm in most experiments except for those on CIFAR10 where we use LPIPS metric for slightly better performance. For the generation of teacher trajectories for AMED-Solver, we use DPM-Solver-2 with doubled NFE. For AMED-Plugin, we use the same solver that generates student sampling trajectories with 1 or 2.
Sampling. Due to designation, our AMED-Solver or AMED-Plugin naturally create solvers with even NFE. Therefore once AFS is used, the total NFE become odd. With the goal of designing fast ODE solvers in extremely small NFE, we mainly test our method on NFE where AFS is applied. There are also results on NFE without using AFS.
Evaluation. We measure the sample quality via Frechet Inception Distance (FID) , which is a well-known quantitative evaluation metric for image quality that aligns well with human perception. For all the experiments involved, we calculate FID with 50k samples using the implementation in .
Appendix C Additional Results
In pilot experiments, we find that the fast degradation of single-step solvers can be alleviated by appropriate choose of the intermediate time steps. As shown in Tab. 6, the performance of DPM-Solver-2 is very sensitive to the choice of its hyperparameter . We apply AMED-Plugin to learn the appropriate and find that it achieves similar results with the best searched while adding little training overhead and negligible sampling overhead.
C.2 Ablation Study on Intermediate Steps
As illustrated in Sec. 3.4, for the teacher sampling trajectory, intermediate time steps are injected in every sampling step. By using smooth interpolation, we mean that the teacher time schedule given by the original schedule combined with , is equivalent to the time schedule obtained by simply setting the total number of time steps to in Eq. 10 under the same . In this way, we can easily extract samples on teacher trajectories at to get reference samples .
Here we take unconditional generation on CIFAR10 using iPNDM with AMED-Plugin as an example and provide an ablation study on the choose of . The student and teacher solvers are set to be the same. The results are shown in Tab. 7. Throughout our experiments, we set .
C.3 Ablation Study on AFS
The trick of analytical first step (AFS) is first introduced in to reduce one NFE, where the authors replace the U-Net output in the first sampling step with the direction of . In Tab. 9 and Tab. 10, we provide extended results of Tab. 2(a) and Tab. 2(b) as well as the ablations between AFS and our proposed AMED-Plugin. We find that the use of AFS provides consistent improvement on datasets with resolutions of 32 and 64. The results show that in most cases they can be considered as two independent components that can together boost the performance of various ODE solvers. However, for datasets with large resolutions, applying AFS usually causes a large degradation (see Tab. 8).
C.4 More Quantitative Results
In this section, we provide additional quantitative results on more datasets including LSUN Cat 256256 , latent-space LSUN Bedroom 256256 and ImageNet 256256 with classifier guidance . The results are shown in Tab. 11, Tab. 12 and Tab. 13.
C.5 More Qualitative Results
We give more qualitative results generated by stable-diffusion-v1 with a default classifier-free guidance scale 7.5 in Fig. 9. Results on various datasets with NFE of 3 and 5 are provided from Fig. 11 to Fig. 16.
Appendix D Theoretical Analysis
In Sec. 3.1, we experimentally showed that the sampling trajectory of diffusion models generated by an ODE solver almost lies in a two-dimensional subspace embedded in the ambient space. This is the core condition for the mean value theorem to approximately hold in the vector-valued function case. However, the sampling trajectory would not necessarily lie in a plane. In this section, we analyze to what extent will this affects our AMED-Solver, where we set the two-dimensional subspace to be the place spanned by the first two principal components.
Define the scaled logistic function to be
Assume that the parallel component is optimally learned and for the perpendicular component, we have
There exists such a s.t. with high probability that
Under assumption 1 and 4, let , then concentrates at a thin shell with radius
Since the SDE Eq. 17 has zero drift coefficient, its perturbation kernel is a Gaussian with zero mean . The covariance is given by
By the well-known concentration of measure , there exists a constant s.t. for any , we have
Given , under the assumptions and Lemma 1 above, with high probability we have
Under assumptions and Lemma 1 above, we have