Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, Mikhail Kudinov
Introduction
Deep generative modelling proved to be effective in various machine learning fields, and speech synthesis is no exception. Modern text-to-speech (TTS) systems often consist of two parts designed as deep neural networks: the first part converts the input text into time-frequency domain acoustic features (feature generator), and the second one synthesizes raw waveform conditioned on these features (vocoder). Introduction of the conventional state-of-the-art autoregressive models such as Tacotron2 (Shen et al., 2018) used for feature generation and WaveNet (van den Oord et al., 2016) used as vocoder marked the beginning of the neural TTS era. Later, other popular generative modelling frameworks such as Generative Adversarial Networks (Goodfellow et al., 2014) and Normalizing Flows (Rezende & Mohamed, 2015) were used in the design of TTS engines for a parallel generation with comparable quality of the synthesized speech.
Since the publication of the WaveNet paper (2016), there have been various attempts to propose a parallel non-autoregressive vocoder, which could synthesize high-quality speech. Popular architectures based on Normalizing Flows like Parallel WaveNet (van den Oord et al., 2018) and WaveGlow (Prenger et al., 2019) managed to accelerate inference while keeping synthesis quality at a very high level but demonstrated fast synthesis on GPU devices only. Eventually, parallel GAN-based vocoders such as Parallel WaveGAN (Yamamoto et al., 2020), MelGAN (Kumar et al., 2019), and HiFi-GAN (Kong et al., 2020) greatly improved the performance of waveform generation on CPU devices. Furthermore, the latter model is reported to produce speech samples of state-of-the-art quality outperforming WaveNet.
Among feature generators, Tacotron2 (Shen et al., 2018) and Transformer-TTS (Li et al., 2019) enabled highly natural speech synthesis. Producing acoustic features frame by frame, they achieve almost perfect mel-spectrogram reconstruction from input text. Nonetheless, they often suffer from computational inefficiency and pronunciation issues coming from attention failures. Addressing these problems, such models as FastSpeech (Ren et al., 2019) and Parallel Tacotron (Elias et al., 2020) substantially improved inference speed and pronunciation robustness by utilizing non-autoregressive architectures and building hard monotonic alignments from estimated token lengths. However, in order to learn character duration, they still require pre-computed alignment from the teacher model. Finally, the recently proposed Non-Attentive Tacotron framework (Shen et al., 2020) managed to learn durations implicitly by employing the Variational Autoencoder concept.
Glow-TTS feature generator (Kim et al., 2020) based on Normalizing Flows can be considered as one of the most successful attempts to overcome pronunciation and computational latency issues typical for autoregressive solutions. Glow-TTS model made use of Monotonic Alignment Search algorithm (an adoption of Viterbi training (Rabiner, 1989) finding the most likely hidden alignment between two sequences) proposed to map the input text to mel-spectrograms efficiently. The alignment learned by Glow-TTS is intentionally designed to avoid some of the pronunciation problems models like Tacotron2 suffer from. Also, in order to enable parallel synthesis, Glow-TTS borrows encoder architecture from Transformer-TTS (Li et al., 2019) and decoder architecture from Glow (Kingma & Dhariwal, 2018). Thus, compared with Tacotron2, Glow-TTS achieves much faster inference making fewer alignment mistakes. Besides, in contrast to other parallel TTS solutions such as FastSpeech, Glow-TTS does not require an external aligner to obtain token duration information as Monotonic Alignment Search (MAS) operates in an unsupervised way.
Lately, another family of generative models called Diffusion Probabilistic Models (DPMs) (Sohl-Dickstein et al., 2015) has started to prove its capability to model complex data distributions such as images (Ho et al., 2020), shapes (Cai et al., 2020), graphs (Niu et al., 2020), handwriting (Luhman & Luhman, 2020). The basic idea behind DPMs is as follows: we build a forward diffusion process by iteratively destroying original data until we get some simple distribution (usually standard normal), and then we try to build a reverse diffusion parameterized with a neural network so that it follows the trajectories of the reverse-time forward diffusion. Stochastic calculus offers a continuous easy-to-use framework for training DPMs (Song et al., 2021) and, which is perhaps more important, provides a number of flexible inference schemes based on numerical differential equation solvers.
As far as text-to-speech applications are concerned, two vocoders representing the DPM family showed impressive results in raw waveform reconstruction: WaveGrad (Chen et al., 2021) and DiffWave (Kong et al., 2021) were shown to reproduce the fine-grained structure of human speech and match strong autoregressive baselines such as WaveNet in terms of synthesis quality while at the same time requiring much fewer sequential operations. However, despite such a success in neural vocoding, no feature generator based on diffusion probabilistic modelling is known so far.
This paper introduces Grad-TTS, an acoustic feature generator with a score-based decoder using recent diffusion probabilistic modelling insights. In Grad-TTS, MAS-aligned encoder outputs are passed to the decoder that transforms Gaussian noise parameterized by these outputs into a mel-spectrogram. To cope with the task of reconstructing data from Gaussian noise with varying parameters, we write down a generalized version of conventional forward and reverse diffusions. One of the remarkable features of our model is that it provides explicit control of the trade-off between output mel-spectrogram quality and inference speed. In particular, we find that Grad-TTS is capable of generating mel-spectrograms of high quality with only as few as ten iterations of reverse diffusion, which makes it possible to outperform Tacotron2 in terms of speed on GPU devices. Additionally, we show that it is possible to train Grad-TTS as an end-to-end TTS pipeline (i.e., vocoder and feature generator are combined in a single model) by replacing its output domain from mel-spectrogram to raw waveform.
Diffusion probabilistic modelling
Loosely speaking, a process of the diffusion type is a stochastic process that satisfies a stochastic differential equation (SDE)
where is the standard Brownian motion, for some finite time horizon , and coefficients and (called drift and diffusion correspondingly) satisfy certain measurability conditions. A rigorous definition of the diffusion type processes, as well as other notions from stochastic calculus we use in this section, can be found in (Liptser & Shiryaev, 1978).
It is easy to find such a stochastic process that terminal distribution converges to standard normal when for any initial data distribution ( is identity matrix and is data dimensionality). In fact, there are lots of such processes as it follows from the formulae given later in this section. Any process of the diffusion type with such property is called forward diffusion and the goal of diffusion probabilistic modelling is to find a reverse diffusion such that its trajectories closely follow those of the forward diffusion but in reverse time order. This is, of course, a much harder task than making Gaussian noise out of data, but in many cases it still can be accomplished if we parameterize reverse diffusion with a proper neural network. In this case, generation boils down to sampling random noise from and then just solving the SDE describing dynamics of the reverse diffusion with any numerical solver (usually a simple first-order Euler-Maruyama scheme (Kloeden & Platen, 1992) is used). If forward and reverse diffusion processes have close trajectories, then the distribution of resulting samples will be very close to that of the data . This approach to generative modelling is summarized in Figure 1.
Until recently, score-based and denoising diffusion probabilistic models were formalized in terms of Markov chains (Sohl-Dickstein et al., 2015; Song & Ermon, 2019; Ho et al., 2020; Song & Ermon, 2020). A unified approach introduced by Song et al. (2021) has demonstrated that these Markov chains actually approximated trajectories of stochastic processes satisfying certain SDEs. In our work, we follow this paper and define our DPM in terms of SDEs rather than Markov chains. As one can see later in Section 3, the task we are solving suggests generalizing DPMs described in (Song et al., 2021) in such a way that for infinite time horizon forward diffusion transforms any data distribution into instead of for any given mean and diagonal covariance matrix . So, the rest of this section contains the detailed description of the generalized forward and reverse diffusions we utilize as well as the loss function we optimize to train the reverse diffusion. All corresponding derivations can be found in Appendix.
First, we need to define a forward diffusion process that transforms any data into Gaussian noise given infinite time horizon . If -dimensional stochastic process satisfies the following SDE:
for non-negative function , which we will refer to as noise schedule, vector , and diagonal matrix with positive elements, then its solution (if it exists) is given by
Note that the exponential of a diagonal matrix is just an element-wise exponential. Let
By properties of Itô’s integral conditional distribution of given is Gaussian:
It means that if we consider infinite time horizon then for any noise schedule such that we have
So, random variable converges in distribution to independently of , and it is exactly the property we need: forward diffusion satisfying SDE (2) transforms any data distribution into Gaussian noise .
2 Reverse diffusion
While in earlier works on DPMs reverse diffusion was trained to approximate the trajectories of forward diffusion, Song et al. (2021) proposed to use the result by Anderson (1982), who derived an explicit formula for reverse-time dynamics of a wide class of stochastic processes of the diffusion type. In our case, this result leads to the following SDE for the reverse diffusion:
where is the reverse-time Brownian motion and is the probability density function of random variable . This SDE is to be solved backwards starting from terminal condition .
Moreover, Song et al. (2021) have shown that instead of SDE (8), we can consider an ordinary differential equation
Forward Kolmogorov equations corresponding to (2) and (9) are identical, which means that the evolution of probability density functions of stochastic processes given by (2) and (9) is the same.
Thus, if we have a neural network that estimates the gradient of the log-density of noisy data , then we can model data distribution by sampling from and numerically solving either (8) or (9) backwards in time.
3 Loss function
Estimating gradients of log-density of noisy data is often referred to as score matching, and in recent papers (Song & Ermon, 2019, 2020) loss was used to approximate these gradients with a neural network. So, in our paper, we use the same type of loss.
Due to the formula (6), we can sample noisy data given only initial data without sampling intermediate values . Moreover, is Gaussian, which means that its log-density has a very simple closed form. If we sample from and then put
in accordance with (6), then the gradient of log-density of noisy data in this point is given by
where is the probability density function of the conditional distribution (6). Thus, loss function corresponding to estimating the gradient of log-density of data corrupted with noise accumulated by time is
where is sampled from and is calculated by formula (10).
Grad-TTS
The acoustic feature generator we propose consists of three modules: encoder, duration predictor, and decoder. In this section, we will describe their architectures as well as training and inference procedures. The general approach is illustrated in Figure 2. Grad-TTS has very much in common with Glow-TTS (Kim et al., 2020), a feature generator based on Normalizing Flows. The key difference lies in the principles the decoder relies on.
The output sequence is then passed to the decoder, which is a Diffusion Probabilistic Model. A neural network with parameters defines an ordinary differential equation (ODE)
which is solved backwards in time using the first-order Euler scheme. The sequence is also used to define the terminal condition . Noise schedule and time horizon are some pre-defined hyperparameters whose choice mostly depends on the data, while step size in the Euler scheme is a hyperparameter that can be chosen after Grad-TTS is trained. It expresses the trade-off between the quality of output mel-spectrograms and inference speed.
Reverse diffusion in Grad-TTS evolves according to equation (13) for the following reasons:
We obtained better results in practice when using dynamics (9) instead of (8): for small values of step size , they performed equally well, while for larger values the former led to much better sounding results.
We chose to simplify the whole feature generation pipeline.
We used as an additional input to the neural network . It follows from (11) that the neural network essentially tries to predict Gaussian noise added to data observing only noisy data . So, if for every time we supply with an additional knowledge of how the limiting noise looks like (note that it is different for different text input), then this network can make more accurate predictions of noise at time .
We also found it beneficial for the model performance to introduce a temperature hyperparameter and to sample terminal condition from instead of . Tuning can help to keep the quality of output mel-spectrograms at the same level when using larger values of step size .
2 Training
One of Grad-TTS training objectives is to minimize the distance between aligned encoder output and target mel-spectrogram because the inference scheme that has just been described suggests to start decoding from random noise . Intuitively, it is clear that decoding is easier if we start from noise, which is already close to the target in some sense.
The encoder loss has to be optimized with respect to both encoder parameters and alignment function . Since it is hard to do a joint optimization, we apply an iterative approach proposed by Kim et al. (2020). Each iteration of optimization consists of two steps: (i) searching for an optimal alignment given fixed encoder parameters; (ii) fixing this alignment and taking one step of stochastic gradient descent to optimize loss function with respect to encoder parameters. We use Monotonic Alignment Search at the first step of this approach. MAS utilizes the concept of dynamic programming to find an optimal (from the point of view of loss function ) monotonic surjective alignment. This algorithm is described in detail in (Kim et al., 2020).
To estimate the optimal alignment at inference, Grad-TTS employs the duration predictor network. As in (Kim et al., 2020), we train the duration predictor with Mean Square Error (MSE) criterion in logarithmic domain:
As for the loss related to the DPM, it is calculated using formulae from Section 2. As already mentioned, we put , so the distribution of noisy data (6) simplifies, and its covariance matrix becomes just an identity matrix multiplied by a scalar
The overall diffusion loss function is the expectation of weighted losses associated with estimating gradients of log-density of noisy data at different times :
where stands for target mel-spectrogram sampled from training data, is sampled from uniform distribution on , – from and the formula
To sum it up, the training procedure consists of the following steps:
Fix the encoder, duration predictor, and decoder parameters and run MAS algorithm to find the alignment that minimizes .
Fix the alignment and minimize with respect to encoder, duration predictor, and decoder parameters.
Repeat the first two steps till convergence.
3 Model architecture
As for the encoder and duration predictor, we use exactly the same architectures as in Glow-TTS, which in its turn borrows the structure of these modules from Transformer-TTS (Li et al., 2019) and FastSpeech (Ren et al., 2019) correspondingly. The duration predictor consists of two convolutional layers followed by a projection layer that predicts the logarithm of duration. The encoder is composed of a pre-net, Transformer blocks with multi-head self-attention, and the final linear projection layer. The pre-net consists of layers of convolutions followed by a fully-connected layer.
The decoder network has the same U-Net architecture (Ronneberger et al., 2015) used by Ho et al. (2020) to generate images, except that we use twice fewer channels and three feature map resolutions instead of four to reduce model size. In our experiments we use -dimensional mel-spectrograms, so operates on resolutions , and . We zero-pad mel-spectrograms if the number of frames is not a multiple of . Aligned encoder output is concatenated with U-Net input as an additional channel.
Experiments
LJSpeech dataset (Ito, 2017) containing approximately hours of English female voice recordings sampled at kHz was used to train the Grad-TTS model. The test set contained around short audio recordings (duration less than seconds each). The input text was phonemized before passing to the encoder; as for the output acoustic features, we used conventional -dimensional mel-spectrograms. We tried training both on original and normalized mel-spectrograms and found that the former performed better. Grad-TTS was trained for iterations on a single GPU (NVIDIA RTX Ti with GB memory) with mini-batch size . We chose Adam optimizer and set the learning rate to .
We would like to mention several important things about Grad-TTS training:
We chose , , and .
As in (Bińkowski et al., 2020; Donahue et al., 2021), we use random mel-spectrogram segments of fixed length ( seconds in our case) as training targets to allow for memory-efficient training. However, MAS and the duration predictor still use whole mel-spectrograms.
Although diffusion loss seems to converge very slowly after the beginning epochs, as shown on Figure 3, such long training is essential to get a good model because the neural network has to learn to estimate gradients accurately for all . Two models with almost equal diffusion losses can produce mel-spectrograms of very different quality: inaccurate predictions for a small subset may have a small impact on but be crucial for the output mel-spectrogram quality if ODE solver involves calculating in at least one point belonging to .
Once trained, Grad-TTS enables the trade-off between quality and inference speed due to the ability to vary the number of steps the decoder takes to solve ODE (13) at inference. So, we evaluate four models which we denote by Grad-TTS-N where . We use at synthesis for all four models. As baselines, we take an official implementation of Glow-TTS (Kim et al., 2020), the model which resembles ours to the most extent among the existing feature generators, FastSpeech (Ren et al., 2019), and state-of-the-art Tacotron2 (Shen et al., 2018). Recently proposed HiFi-GAN (Kong et al., 2020) is known to provide excellent sound quality, so we use this vocoder with all models we compare.
To make subjective evaluation of TTS models, we used the crowdsourcing platform Amazon Mechanical Turk. For Mean Opinion Score (MOS) estimation we synthesized sentences from the test set with each model. The assessors were asked to estimate the quality of synthesized speech on a nine-point Likert scale, the lowest and the highest scores being point (“Bad”) and points (“Excellent”) with a step of point. To ensure the reliability of the obtained results, only Master assessors were assigned to complete the listening test. Each audio was evaluated by assessors. A small subset of speech samples used in the test is available at https://grad-tts.github.io/.
MOS results with confidence intervals are presented in Table 2. It demonstrates that although the quality of the synthesized speech gets better when we use more iterations of the reverse diffusion, the quality gain becomes marginal starting from a certain number of iterations. In particular, there is almost no difference between Grad-TTS-1000 and Grad-TTS-10 in terms of MOS, while the gap between Grad-TTS-10 and Grad-TTS-4 ( was the smallest number of iterations leading to satisfactory quality) is much more significant. As for other feature generators, Grad-TTS-10 is competitive with all compared models, including state-of-the-art Tacotron2. Furthermore, Grad-TTS-1000 achieves almost natural synthesis with MOS being less than that for ground truth recordings by only . We would like to note that the relatively low results of FastSpeech could possibly be explained by the fact that we used its unofficial implementation https://github.com/xcmyz/FastSpeech.
To verify the benefits of the proposed generalized DPM framework we trained the model with the same architecture as Grad-TTS to reconstruct mel-spectrograms from instead of . The preference test provided in Table 1 shows that Grad-TTS-10 is significantly better ( in sign test) than this model taking , and even iterations of the reverse diffusion. It demonstrates that the model trained to generate from needs more steps of ODE solver to get high-quality mel-spectrograms than Grad-TTS we propose. We believe this is because the task of reconstructing mel-spectrogram from pure noise is more difficult than the one of reconstructing it from its noisy copy . One possible objection could be that the model trained with as terminal distribution can just add to this noise at the first step of sampling (it is possible because has as its input) and then repeat the same steps as our model to generate data from . In this case, it would generate mel-spectrograms of the same quality as our model taking only one step more. However, this argument is wrong, since reverse diffusion removes noise not arbitrarily, but according to the reverse trajectories of the forward diffusion. Since forward diffusion adds noise gradually, reverse diffusion has to remove noise gradually as well, and the first step of the reverse diffusion cannot be adding to Gaussian noise with zero mean because the last step of the forward diffusion is not a jump from to zero.
We also made an attempt to estimate what kinds of mistakes are characteristic of certain models. We compared Tacotron2, Glow-TTS, and Grad-TTS-10 as the fastest version of our model with high synthesis quality. Each record was estimated by assessors. Figure 4 demonstrates the results of the multiple-choice test whose participants had to choose which kinds of errors (if any) they could hear: sonic artifacts like clicking sounds or background noise (“sonic” in the figure), mispronunciation of words/phonemes (“mispron”), unnatural pauses (“pause”), monotone speech (“monotonic”), robotic voice (“robotic”), wrong word stressing (“stress”) or others. It is clear from the figure that Glow-TTS frequently stresses words in a wrong way, and the sound it produces is perceived as “robotic” in around a quarter of cases. These are the major factors that make Glow-TTS performance inferior to that of Grad-TTS and Tacotron2, which in their turn have more or less the same drawbacks in terms of synthesis quality.
2 Objective evaluation
Although DPMs can be shown to maximize weighted variational lower bound (Ho et al., 2020) on data log-likelihood, they do not explicitly optimize exact data likelihood. In spite of this, Song et al. (2021) show that it is still possible to calculate it using the instantaneous change of variables formula (Chen et al., 2018) if we look at DPMs from the “continuous” point of view. However, it is necessary to use Hutchinson’s trace estimator to make computations feasible, so in Table 2 log-likelihood for Grad-TTS comes with a confidence interval.
We randomly chose sentences from the test set and calculated their average log-likelihood under two probabilistic models we consider – Glow-TTS and Grad-TTS. Interestingly, Grad-TTS achieves better log-likelihood than Glow-TTS even though the latter has a decoder with x larger capacity and was trained to maximize exact data likelihood. Similar phenomena were observed by Song et al. (2021) in the image generation task.
3 Efficiency estimation
We assess the efficiency of the proposed model in terms of Real-Time Factor (RTF is how many seconds it takes to generate one second of audio) computed on GPU and the number of parameters. Table 2 contains efficiency information for all models under comparison. Additional information regarding absolute inference speed dependency on the input text length is given in Figure 5.
Due to its flexibility at inference, Grad-TTS is capable of real-time synthesis on GPU: if the number of decoder steps is less than , it reaches RTF . Moreover, although it cannot compete with Glow-TTS and FastSpeech in terms of inference speed, it still can be approximately twice faster than Tacotron2 if we use decoder iterations sufficient for getting high-fidelity mel-spectrograms. Besides, Grad-TTS has around parameters, thus being significantly smaller than other feature generators we compare.
4 End-to-end TTS
The results of our preliminary experiments show that it is also possible to train an end-to-end TTS model as a DPM. In brief, we moved from U-Net to WaveGrad (Chen et al., 2021) in Grad-TTS decoder: the overall architecture resembles WaveGrad conditioned on the aligned encoder output instead of ground truth mel-spectrograms as in original WaveGrad. Although synthesized speech quality is fair enough, it cannot compete with the results reported above, so we do not include our end-to-end model in the listening test but provide demo samples at https://grad-tts.github.io/. Encoder and duration predictor parameters are calculated together.
Future work
End-to-end speech synthesis results reported above show that it is a promising future research direction for text-to-speech applications. However, there is also much room for investigating general issues regarding DPMs.
In the analysis in Section 2, we always assume that both forward and reverse diffusion processes exist, i.e., SDEs (2) and (8) have strong solutions. It applies some Lipschitz-type constraints (Liptser & Shiryaev, 1978) on noise schedule and, what is more important, on the neural network . Wasserstein GANs offer an encouraging example of incorporating Lipschitz constraints into neural networks training (Gulrajani et al., 2017), suggesting that similar techniques may improve DPMs.
Little attention has been paid so far to the choice of the noise schedule – most researchers use a simple linear schedule. Also, it is mostly unclear how to choose weights for losses (12) at time in the global loss function optimally. A thorough investigation of such practical questions is crucial as it can facilitate applying DPMs to new machine learning problems.
Conclusion
We have presented Grad-TTS, the first acoustic feature generator utilizing the concept of diffusion probabilistic modelling. The main generative engine of Grad-TTS is the diffusion-based decoder that transforms Gaussian noise parameterized with the encoder output into mel-spectrogram while alignment is performed with Monotonic Alignment Search. The model we propose allows to vary the number of decoder steps at inference, thus providing a tool to control the trade-off between inference speed and synthesized speech quality. Despite its iterative decoding, Grad-TTS is capable of real-time synthesis. Moreover, it can generate mel-spectrograms twice faster than Tacotron2 while keeping synthesis quality competitive with common TTS baselines.
References
Appendix
We include an appendix with detailed derivations, proofs and additional information. Our proposed diffusion probabilistic framework employs generalized terminal distribution instead of as proposed by Song et al. (2021). The derivation for the solution (3) of SDE (2) that transforms the original data distribution to the terminal distribution is described in Appendix A. In Appendix B we also derive the distribution which the solution (3) for the diffused data follows. Then, the goal of diffusion probabilistic modelling is to reconstruct the reverse-time trajectories of the forward diffusion process, and Song et al. (2021) showed that these dynamics can follow two different differential equations: either SDE (8) proposed by Anderson (1982) or ODE (9). So, Appendix C contains these differential equations for serving as terminal distribution. They depend on time-dependent gradient field supposed to be modelled using neural network. In order to train it, we show how to compute the gradient in Appendix D.
Exponential of a diagonal matrix is just element-wise exponential, so we can rewrite it in multidimensional form as
or writing this down in terms of :
where is identity matrix.
Let . It is a diagonal matrix and its -th diagonal element equals . Assume for each . Itô’s integral is defined as the limit of integral sums when mesh of partition tends to zero:
where the first equality in distribution holds due to the properties of Brownian motion and the fact that are deterministic (implying that are independent normal random variables with mean and variance ) and the second equality in distribution follows from Lévy’s continuity theorem (it is easy to check that the sequence of characteristic functions of random variables on the left-hand side converges point-wise to the characteristic function of the random variable on the right-hand side). Then, simple integration gives
It implies that in multidimensional case we have:
C Reverse dynamics
The result by Anderson (1982) implies that if -dimensional process of the diffusion type satisfies
where is the probability density function of random variable and is the reverse-time standard Brownian motion such that is independent of its past increments for . Reverse-time dynamics means that all the integrals associated with reverse-time differentials have as their lower limit (e.g. relates to ). Anderson’s result is obtained under the assumption that Kolmogorov equations (for probability density functions) associated with all considered processes have unique smooth solutions. On the other hand, Song et al. (2021) argued that SDE (28) has the same forward Kolmogorov equation as the following ODE:
which means that processes following (28) and (30) are equal in distribution if they start from the same initial distribution . In our case and , so we have two equivalent reverse diffusion dynamics:
where both differential equations are to be solved backwards.
D Score estimation
If is known, then (27) implies that
where is the probability density function of conditional distribution . So, if we sample by the formula where , then . In the simplified case when we have where . In this case gradient of noisy data log-density reduces to . If , then we have