Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech

Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, Mikhail Kudinov

Introduction

Deep generative modelling proved to be effective in various machine learning fields, and speech synthesis is no exception. Modern text-to-speech (TTS) systems often consist of two parts designed as deep neural networks: the first part converts the input text into time-frequency domain acoustic features (feature generator), and the second one synthesizes raw waveform conditioned on these features (vocoder). Introduction of the conventional state-of-the-art autoregressive models such as Tacotron2 (Shen et al., 2018) used for feature generation and WaveNet (van den Oord et al., 2016) used as vocoder marked the beginning of the neural TTS era. Later, other popular generative modelling frameworks such as Generative Adversarial Networks (Goodfellow et al., 2014) and Normalizing Flows (Rezende & Mohamed, 2015) were used in the design of TTS engines for a parallel generation with comparable quality of the synthesized speech.

Since the publication of the WaveNet paper (2016), there have been various attempts to propose a parallel non-autoregressive vocoder, which could synthesize high-quality speech. Popular architectures based on Normalizing Flows like Parallel WaveNet (van den Oord et al., 2018) and WaveGlow (Prenger et al., 2019) managed to accelerate inference while keeping synthesis quality at a very high level but demonstrated fast synthesis on GPU devices only. Eventually, parallel GAN-based vocoders such as Parallel WaveGAN (Yamamoto et al., 2020), MelGAN (Kumar et al., 2019), and HiFi-GAN (Kong et al., 2020) greatly improved the performance of waveform generation on CPU devices. Furthermore, the latter model is reported to produce speech samples of state-of-the-art quality outperforming WaveNet.

Among feature generators, Tacotron2 (Shen et al., 2018) and Transformer-TTS (Li et al., 2019) enabled highly natural speech synthesis. Producing acoustic features frame by frame, they achieve almost perfect mel-spectrogram reconstruction from input text. Nonetheless, they often suffer from computational inefficiency and pronunciation issues coming from attention failures. Addressing these problems, such models as FastSpeech (Ren et al., 2019) and Parallel Tacotron (Elias et al., 2020) substantially improved inference speed and pronunciation robustness by utilizing non-autoregressive architectures and building hard monotonic alignments from estimated token lengths. However, in order to learn character duration, they still require pre-computed alignment from the teacher model. Finally, the recently proposed Non-Attentive Tacotron framework (Shen et al., 2020) managed to learn durations implicitly by employing the Variational Autoencoder concept.

Glow-TTS feature generator (Kim et al., 2020) based on Normalizing Flows can be considered as one of the most successful attempts to overcome pronunciation and computational latency issues typical for autoregressive solutions. Glow-TTS model made use of Monotonic Alignment Search algorithm (an adoption of Viterbi training (Rabiner, 1989) finding the most likely hidden alignment between two sequences) proposed to map the input text to mel-spectrograms efficiently. The alignment learned by Glow-TTS is intentionally designed to avoid some of the pronunciation problems models like Tacotron2 suffer from. Also, in order to enable parallel synthesis, Glow-TTS borrows encoder architecture from Transformer-TTS (Li et al., 2019) and decoder architecture from Glow (Kingma & Dhariwal, 2018). Thus, compared with Tacotron2, Glow-TTS achieves much faster inference making fewer alignment mistakes. Besides, in contrast to other parallel TTS solutions such as FastSpeech, Glow-TTS does not require an external aligner to obtain token duration information as Monotonic Alignment Search (MAS) operates in an unsupervised way.

Lately, another family of generative models called Diffusion Probabilistic Models (DPMs) (Sohl-Dickstein et al., 2015) has started to prove its capability to model complex data distributions such as images (Ho et al., 2020), shapes (Cai et al., 2020), graphs (Niu et al., 2020), handwriting (Luhman & Luhman, 2020). The basic idea behind DPMs is as follows: we build a forward diffusion process by iteratively destroying original data until we get some simple distribution (usually standard normal), and then we try to build a reverse diffusion parameterized with a neural network so that it follows the trajectories of the reverse-time forward diffusion. Stochastic calculus offers a continuous easy-to-use framework for training DPMs (Song et al., 2021) and, which is perhaps more important, provides a number of flexible inference schemes based on numerical differential equation solvers.

As far as text-to-speech applications are concerned, two vocoders representing the DPM family showed impressive results in raw waveform reconstruction: WaveGrad (Chen et al., 2021) and DiffWave (Kong et al., 2021) were shown to reproduce the fine-grained structure of human speech and match strong autoregressive baselines such as WaveNet in terms of synthesis quality while at the same time requiring much fewer sequential operations. However, despite such a success in neural vocoding, no feature generator based on diffusion probabilistic modelling is known so far.

This paper introduces Grad-TTS, an acoustic feature generator with a score-based decoder using recent diffusion probabilistic modelling insights. In Grad-TTS, MAS-aligned encoder outputs are passed to the decoder that transforms Gaussian noise parameterized by these outputs into a mel-spectrogram. To cope with the task of reconstructing data from Gaussian noise with varying parameters, we write down a generalized version of conventional forward and reverse diffusions. One of the remarkable features of our model is that it provides explicit control of the trade-off between output mel-spectrogram quality and inference speed. In particular, we find that Grad-TTS is capable of generating mel-spectrograms of high quality with only as few as ten iterations of reverse diffusion, which makes it possible to outperform Tacotron2 in terms of speed on GPU devices. Additionally, we show that it is possible to train Grad-TTS as an end-to-end TTS pipeline (i.e., vocoder and feature generator are combined in a single model) by replacing its output domain from mel-spectrogram to raw waveform.

Diffusion probabilistic modelling

Loosely speaking, a process of the diffusion type is a stochastic process that satisfies a stochastic differential equation (SDE)

where WtW_{t} is the standard Brownian motion, t∈[0,T]t\in[0,T] for some finite time horizon TT, and coefficients bb and aa (called drift and diffusion correspondingly) satisfy certain measurability conditions. A rigorous definition of the diffusion type processes, as well as other notions from stochastic calculus we use in this section, can be found in (Liptser & Shiryaev, 1978).

It is easy to find such a stochastic process that terminal distribution Law(XT)Law(X_{T}) converges to standard normal N(0,I)\mathcal{N}(0,I) when T→∞T\to\infty for any initial data distribution Law(X0)Law(X_{0}) (II is n×nn\times n identity matrix and nn is data dimensionality). In fact, there are lots of such processes as it follows from the formulae given later in this section. Any process of the diffusion type with such property is called forward diffusion and the goal of diffusion probabilistic modelling is to find a reverse diffusion such that its trajectories closely follow those of the forward diffusion but in reverse time order. This is, of course, a much harder task than making Gaussian noise out of data, but in many cases it still can be accomplished if we parameterize reverse diffusion with a proper neural network. In this case, generation boils down to sampling random noise from N(0,I)\mathcal{N}(0,I) and then just solving the SDE describing dynamics of the reverse diffusion with any numerical solver (usually a simple first-order Euler-Maruyama scheme (Kloeden & Platen, 1992) is used). If forward and reverse diffusion processes have close trajectories, then the distribution of resulting samples will be very close to that of the data Law(X0)Law(X_{0}). This approach to generative modelling is summarized in Figure 1.

Until recently, score-based and denoising diffusion probabilistic models were formalized in terms of Markov chains (Sohl-Dickstein et al., 2015; Song & Ermon, 2019; Ho et al., 2020; Song & Ermon, 2020). A unified approach introduced by Song et al. (2021) has demonstrated that these Markov chains actually approximated trajectories of stochastic processes satisfying certain SDEs. In our work, we follow this paper and define our DPM in terms of SDEs rather than Markov chains. As one can see later in Section 3, the task we are solving suggests generalizing DPMs described in (Song et al., 2021) in such a way that for infinite time horizon forward diffusion transforms any data distribution into N(μ,Σ)\mathcal{N}(\mu,\Sigma) instead of N(0,I)\mathcal{N}(0,I) for any given mean μ\mu and diagonal covariance matrix Σ\Sigma. So, the rest of this section contains the detailed description of the generalized forward and reverse diffusions we utilize as well as the loss function we optimize to train the reverse diffusion. All corresponding derivations can be found in Appendix.

First, we need to define a forward diffusion process that transforms any data into Gaussian noise given infinite time horizon TT. If nn-dimensional stochastic process XtX_{t} satisfies the following SDE:

for non-negative function βt\beta_{t}, which we will refer to as noise schedule, vector μ\mu, and diagonal matrix Σ\Sigma with positive elements, then its solution (if it exists) is given by

Note that the exponential of a diagonal matrix is just an element-wise exponential. Let

By properties of Itô’s integral conditional distribution of XtX_{t} given X0X_{0} is Gaussian:

It means that if we consider infinite time horizon then for any noise schedule βt\beta_{t} such that lim⁡t→∞e−∫0tβsds=0\lim_{t\to\infty}e^{-\int_{0}^{t}\beta_{s}ds}=0 we have

So, random variable XtX_{t} converges in distribution to N(μ,Σ)\mathcal{N}(\mu,\Sigma) independently of X0X_{0}, and it is exactly the property we need: forward diffusion satisfying SDE (2) transforms any data distribution Law(X0)Law(X_{0}) into Gaussian noise N(μ,Σ)\mathcal{N}(\mu,\Sigma).

2 Reverse diffusion

While in earlier works on DPMs reverse diffusion was trained to approximate the trajectories of forward diffusion, Song et al. (2021) proposed to use the result by Anderson (1982), who derived an explicit formula for reverse-time dynamics of a wide class of stochastic processes of the diffusion type. In our case, this result leads to the following SDE for the reverse diffusion:

where W~t\widetilde{W}_{t} is the reverse-time Brownian motion and ptp_{t} is the probability density function of random variable XtX_{t}. This SDE is to be solved backwards starting from terminal condition XTX_{T}.

Moreover, Song et al. (2021) have shown that instead of SDE (8), we can consider an ordinary differential equation

Forward Kolmogorov equations corresponding to (2) and (9) are identical, which means that the evolution of probability density functions of stochastic processes given by (2) and (9) is the same.

Thus, if we have a neural network sθ(Xt,t)s_{\theta}(X_{t},t) that estimates the gradient of the log-density of noisy data ∇log⁡pt(Xt)\nabla\log{p_{t}(X_{t})}, then we can model data distribution Law(X0)Law(X_{0}) by sampling XTX_{T} from N(μ,Σ)\mathcal{N}(\mu,\Sigma) and numerically solving either (8) or (9) backwards in time.

3 Loss function

Estimating gradients of log-density of noisy data XtX_{t} is often referred to as score matching, and in recent papers (Song & Ermon, 2019, 2020) L2L_{2} loss was used to approximate these gradients with a neural network. So, in our paper, we use the same type of loss.

Due to the formula (6), we can sample noisy data XtX_{t} given only initial data X0X_{0} without sampling intermediate values {Xs}s<t\{X_{s}\}_{s<t}. Moreover, Law(Xt∣X0)Law(X_{t}|X_{0}) is Gaussian, which means that its log-density has a very simple closed form. If we sample ϵt\epsilon_{t} from N(0,λ(Σ,t))\mathcal{N}(0,\lambda(\Sigma,t)) and then put

in accordance with (6), then the gradient of log-density of noisy data in this point XtX_{t} is given by

where p0t(⋅∣X0)p_{0t}(\cdot|X_{0}) is the probability density function of the conditional distribution (6). Thus, loss function corresponding to estimating the gradient of log-density of data X0X_{0} corrupted with noise accumulated by time tt is

where ϵt\epsilon_{t} is sampled from N(0,λ(Σ,t))\mathcal{N}(0,\lambda(\Sigma,t)) and XtX_{t} is calculated by formula (10).

Grad-TTS

The acoustic feature generator we propose consists of three modules: encoder, duration predictor, and decoder. In this section, we will describe their architectures as well as training and inference procedures. The general approach is illustrated in Figure 2. Grad-TTS has very much in common with Glow-TTS (Kim et al., 2020), a feature generator based on Normalizing Flows. The key difference lies in the principles the decoder relies on.

The output sequence μ=μ1:F\mu=\mu_{1:F} is then passed to the decoder, which is a Diffusion Probabilistic Model. A neural network sθ(Xt,μ,t)s_{\theta}(X_{t},\mu,t) with parameters θ\theta defines an ordinary differential equation (ODE)

which is solved backwards in time using the first-order Euler scheme. The sequence μ\mu is also used to define the terminal condition XT∼N(μ,I)X_{T}\sim\mathcal{N}(\mu,I). Noise schedule βt\beta_{t} and time horizon TT are some pre-defined hyperparameters whose choice mostly depends on the data, while step size hh in the Euler scheme is a hyperparameter that can be chosen after Grad-TTS is trained. It expresses the trade-off between the quality of output mel-spectrograms and inference speed.

Reverse diffusion in Grad-TTS evolves according to equation (13) for the following reasons:

We obtained better results in practice when using dynamics (9) instead of (8): for small values of step size hh, they performed equally well, while for larger values the former led to much better sounding results.

We chose Σ=I\Sigma=I to simplify the whole feature generation pipeline.

We used μ\mu as an additional input to the neural network sθ(Xt,μ,t)s_{\theta}(X_{t},\mu,t). It follows from (11) that the neural network sθs_{\theta} essentially tries to predict Gaussian noise added to data X0X_{0} observing only noisy data XtX_{t}. So, if for every time tt we supply sθs_{\theta} with an additional knowledge of how the limiting noise lim⁡T→∞Law(XT∣X0)\lim_{T\to\infty}Law(X_{T}|X_{0}) looks like (note that it is different for different text input), then this network can make more accurate predictions of noise at time t∈[0,T]t\in[0,T].

We also found it beneficial for the model performance to introduce a temperature hyperparameter τ\tau and to sample terminal condition XTX_{T} from N(μ,τ−1I)\mathcal{N}(\mu,\tau^{-1}I) instead of N(μ,I)\mathcal{N}(\mu,I). Tuning τ\tau can help to keep the quality of output mel-spectrograms at the same level when using larger values of step size hh.

2 Training

One of Grad-TTS training objectives is to minimize the distance between aligned encoder output μ\mu and target mel-spectrogram yy because the inference scheme that has just been described suggests to start decoding from random noise N(μ,I)\mathcal{N}(\mu,I). Intuitively, it is clear that decoding is easier if we start from noise, which is already close to the target yy in some sense.

The encoder loss Lenc\mathcal{L}_{enc} has to be optimized with respect to both encoder parameters and alignment function AA. Since it is hard to do a joint optimization, we apply an iterative approach proposed by Kim et al. (2020). Each iteration of optimization consists of two steps: (i) searching for an optimal alignment A∗A^{*} given fixed encoder parameters; (ii) fixing this alignment A∗A^{*} and taking one step of stochastic gradient descent to optimize loss function with respect to encoder parameters. We use Monotonic Alignment Search at the first step of this approach. MAS utilizes the concept of dynamic programming to find an optimal (from the point of view of loss function Lenc\mathcal{L}_{enc}) monotonic surjective alignment. This algorithm is described in detail in (Kim et al., 2020).

To estimate the optimal alignment A∗A^{*} at inference, Grad-TTS employs the duration predictor network. As in (Kim et al., 2020), we train the duration predictor DPDP with Mean Square Error (MSE) criterion in logarithmic domain:

As for the loss related to the DPM, it is calculated using formulae from Section 2. As already mentioned, we put Σ=I\Sigma=I, so the distribution of noisy data (6) simplifies, and its covariance matrix becomes just an identity matrix II multiplied by a scalar

The overall diffusion loss function Ldiff\mathcal{L}_{diff} is the expectation of weighted losses associated with estimating gradients of log-density of noisy data at different times t∈[0,T]t\in[0,T]:

where X0X_{0} stands for target mel-spectrogram yy sampled from training data, tt is sampled from uniform distribution on [0,T][0,T], ξt\xi_{t} – from N(0,I)\mathcal{N}(0,I) and the formula

To sum it up, the training procedure consists of the following steps:

Fix the encoder, duration predictor, and decoder parameters and run MAS algorithm to find the alignment A∗A^{*} that minimizes Lenc\mathcal{L}_{enc}.

Fix the alignment A∗A^{*} and minimize Lenc+Ldp+Ldiff\mathcal{L}_{enc}+\mathcal{L}_{dp}+\mathcal{L}_{diff} with respect to encoder, duration predictor, and decoder parameters.

Repeat the first two steps till convergence.

3 Model architecture

As for the encoder and duration predictor, we use exactly the same architectures as in Glow-TTS, which in its turn borrows the structure of these modules from Transformer-TTS (Li et al., 2019) and FastSpeech (Ren et al., 2019) correspondingly. The duration predictor consists of two convolutional layers followed by a projection layer that predicts the logarithm of duration. The encoder is composed of a pre-net, 66 Transformer blocks with multi-head self-attention, and the final linear projection layer. The pre-net consists of 33 layers of convolutions followed by a fully-connected layer.

The decoder network sθs_{\theta} has the same U-Net architecture (Ronneberger et al., 2015) used by Ho et al. (2020) to generate 32×3232\times 32 images, except that we use twice fewer channels and three feature map resolutions instead of four to reduce model size. In our experiments we use 8080-dimensional mel-spectrograms, so sθs_{\theta} operates on resolutions 80×F80\times F, 40×F/240\times F/2 and 20×F/420\times F/4. We zero-pad mel-spectrograms if the number of frames FF is not a multiple of 44. Aligned encoder output μ\mu is concatenated with U-Net input XtX_{t} as an additional channel.

Experiments

LJSpeech dataset (Ito, 2017) containing approximately 2424 hours of English female voice recordings sampled at 22.0522.05kHz was used to train the Grad-TTS model. The test set contained around 500500 short audio recordings (duration less than 1010 seconds each). The input text was phonemized before passing to the encoder; as for the output acoustic features, we used conventional 8080-dimensional mel-spectrograms. We tried training both on original and normalized mel-spectrograms and found that the former performed better. Grad-TTS was trained for 1.7m1.7m iterations on a single GPU (NVIDIA RTX 20802080 Ti with 1111GB memory) with mini-batch size 1616. We chose Adam optimizer and set the learning rate to 0.00010.0001.

We would like to mention several important things about Grad-TTS training:

We chose T=1T=1, βt=β0+(β1−β0)t\beta_{t}=\beta_{0}+(\beta_{1}-\beta_{0})t, β0=0.05\beta_{0}=0.05 and β1=20\beta_{1}=20.

As in (Bińkowski et al., 2020; Donahue et al., 2021), we use random mel-spectrogram segments of fixed length (22 seconds in our case) as training targets yy to allow for memory-efficient training. However, MAS and the duration predictor still use whole mel-spectrograms.

Although diffusion loss Ldiff\mathcal{L}_{diff} seems to converge very slowly after the beginning epochs, as shown on Figure 3, such long training is essential to get a good model because the neural network sθs_{\theta} has to learn to estimate gradients accurately for all t∈t\in. Two models with almost equal diffusion losses can produce mel-spectrograms of very different quality: inaccurate predictions for a small subset S⊂S\subset may have a small impact on Ldiff\mathcal{L}_{diff} but be crucial for the output mel-spectrogram quality if ODE solver involves calculating sθs_{\theta} in at least one point belonging to SS.

Once trained, Grad-TTS enables the trade-off between quality and inference speed due to the ability to vary the number of steps NN the decoder takes to solve ODE (13) at inference. So, we evaluate four models which we denote by Grad-TTS-N where N∈N\in. We use τ=1.5\tau=1.5 at synthesis for all four models. As baselines, we take an official implementation of Glow-TTS (Kim et al., 2020), the model which resembles ours to the most extent among the existing feature generators, FastSpeech (Ren et al., 2019), and state-of-the-art Tacotron2 (Shen et al., 2018). Recently proposed HiFi-GAN (Kong et al., 2020) is known to provide excellent sound quality, so we use this vocoder with all models we compare.

To make subjective evaluation of TTS models, we used the crowdsourcing platform Amazon Mechanical Turk. For Mean Opinion Score (MOS) estimation we synthesized 4040 sentences from the test set with each model. The assessors were asked to estimate the quality of synthesized speech on a nine-point Likert scale, the lowest and the highest scores being 11 point (“Bad”) and 55 points (“Excellent”) with a step of 0.50.5 point. To ensure the reliability of the obtained results, only Master assessors were assigned to complete the listening test. Each audio was evaluated by 1010 assessors. A small subset of speech samples used in the test is available at https://grad-tts.github.io/.

MOS results with 95%95\% confidence intervals are presented in Table 2. It demonstrates that although the quality of the synthesized speech gets better when we use more iterations of the reverse diffusion, the quality gain becomes marginal starting from a certain number of iterations. In particular, there is almost no difference between Grad-TTS-1000 and Grad-TTS-10 in terms of MOS, while the gap between Grad-TTS-10 and Grad-TTS-4 (44 was the smallest number of iterations leading to satisfactory quality) is much more significant. As for other feature generators, Grad-TTS-10 is competitive with all compared models, including state-of-the-art Tacotron2. Furthermore, Grad-TTS-1000 achieves almost natural synthesis with MOS being less than that for ground truth recordings by only 0.10.1. We would like to note that the relatively low results of FastSpeech could possibly be explained by the fact that we used its unofficial implementation https://github.com/xcmyz/FastSpeech.

To verify the benefits of the proposed generalized DPM framework we trained the model with the same architecture as Grad-TTS to reconstruct mel-spectrograms from N(0,I)\mathcal{N}(0,I) instead of N(μ,I)\mathcal{N}(\mu,I). The preference test provided in Table 1 shows that Grad-TTS-10 is significantly better (p<0.005p<0.005 in sign test) than this model taking 1010, 2020 and even 5050 iterations of the reverse diffusion. It demonstrates that the model trained to generate from N(0,I)\mathcal{N}(0,I) needs more steps of ODE solver to get high-quality mel-spectrograms than Grad-TTS we propose. We believe this is because the task of reconstructing mel-spectrogram from pure noise N(0,I)\mathcal{N}(0,I) is more difficult than the one of reconstructing it from its noisy copy N(μ,I)\mathcal{N}(\mu,I). One possible objection could be that the model trained with N(0,I)\mathcal{N}(0,I) as terminal distribution can just add μ\mu to this noise at the first step of sampling (it is possible because sθs_{\theta} has μ\mu as its input) and then repeat the same steps as our model to generate data from N(μ,I)N(\mu,I). In this case, it would generate mel-spectrograms of the same quality as our model taking only one step more. However, this argument is wrong, since reverse diffusion removes noise not arbitrarily, but according to the reverse trajectories of the forward diffusion. Since forward diffusion adds noise gradually, reverse diffusion has to remove noise gradually as well, and the first step of the reverse diffusion cannot be adding μ\mu to Gaussian noise with zero mean because the last step of the forward diffusion is not a jump from μ\mu to zero.

We also made an attempt to estimate what kinds of mistakes are characteristic of certain models. We compared Tacotron2, Glow-TTS, and Grad-TTS-10 as the fastest version of our model with high synthesis quality. Each record was estimated by 55 assessors. Figure 4 demonstrates the results of the multiple-choice test whose participants had to choose which kinds of errors (if any) they could hear: sonic artifacts like clicking sounds or background noise (“sonic” in the figure), mispronunciation of words/phonemes (“mispron”), unnatural pauses (“pause”), monotone speech (“monotonic”), robotic voice (“robotic”), wrong word stressing (“stress”) or others. It is clear from the figure that Glow-TTS frequently stresses words in a wrong way, and the sound it produces is perceived as “robotic” in around a quarter of cases. These are the major factors that make Glow-TTS performance inferior to that of Grad-TTS and Tacotron2, which in their turn have more or less the same drawbacks in terms of synthesis quality.

2 Objective evaluation

Although DPMs can be shown to maximize weighted variational lower bound (Ho et al., 2020) on data log-likelihood, they do not explicitly optimize exact data likelihood. In spite of this, Song et al. (2021) show that it is still possible to calculate it using the instantaneous change of variables formula (Chen et al., 2018) if we look at DPMs from the “continuous” point of view. However, it is necessary to use Hutchinson’s trace estimator to make computations feasible, so in Table 2 log-likelihood for Grad-TTS comes with a 95%95\% confidence interval.

We randomly chose 5050 sentences from the test set and calculated their average log-likelihood under two probabilistic models we consider – Glow-TTS and Grad-TTS. Interestingly, Grad-TTS achieves better log-likelihood than Glow-TTS even though the latter has a decoder with 33x larger capacity and was trained to maximize exact data likelihood. Similar phenomena were observed by Song et al. (2021) in the image generation task.

3 Efficiency estimation

We assess the efficiency of the proposed model in terms of Real-Time Factor (RTF is how many seconds it takes to generate one second of audio) computed on GPU and the number of parameters. Table 2 contains efficiency information for all models under comparison. Additional information regarding absolute inference speed dependency on the input text length is given in Figure 5.

Due to its flexibility at inference, Grad-TTS is capable of real-time synthesis on GPU: if the number of decoder steps is less than 100100, it reaches RTF <0.37<0.37. Moreover, although it cannot compete with Glow-TTS and FastSpeech in terms of inference speed, it still can be approximately twice faster than Tacotron2 if we use 1010 decoder iterations sufficient for getting high-fidelity mel-spectrograms. Besides, Grad-TTS has around 15m15m parameters, thus being significantly smaller than other feature generators we compare.

4 End-to-end TTS

The results of our preliminary experiments show that it is also possible to train an end-to-end TTS model as a DPM. In brief, we moved from U-Net to WaveGrad (Chen et al., 2021) in Grad-TTS decoder: the overall architecture resembles WaveGrad conditioned on the aligned encoder output μ\mu instead of ground truth mel-spectrograms yy as in original WaveGrad. Although synthesized speech quality is fair enough, it cannot compete with the results reported above, so we do not include our end-to-end model in the listening test but provide demo samples at https://grad-tts.github.io/. Encoder and duration predictor parameters are calculated together.

Future work

End-to-end speech synthesis results reported above show that it is a promising future research direction for text-to-speech applications. However, there is also much room for investigating general issues regarding DPMs.

In the analysis in Section 2, we always assume that both forward and reverse diffusion processes exist, i.e., SDEs (2) and (8) have strong solutions. It applies some Lipschitz-type constraints (Liptser & Shiryaev, 1978) on noise schedule βt\beta_{t} and, what is more important, on the neural network sθs_{\theta}. Wasserstein GANs offer an encouraging example of incorporating Lipschitz constraints into neural networks training (Gulrajani et al., 2017), suggesting that similar techniques may improve DPMs.

Little attention has been paid so far to the choice of the noise schedule βt\beta_{t} – most researchers use a simple linear schedule. Also, it is mostly unclear how to choose weights for losses (12) at time tt in the global loss function optimally. A thorough investigation of such practical questions is crucial as it can facilitate applying DPMs to new machine learning problems.

Conclusion

We have presented Grad-TTS, the first acoustic feature generator utilizing the concept of diffusion probabilistic modelling. The main generative engine of Grad-TTS is the diffusion-based decoder that transforms Gaussian noise parameterized with the encoder output into mel-spectrogram while alignment is performed with Monotonic Alignment Search. The model we propose allows to vary the number of decoder steps at inference, thus providing a tool to control the trade-off between inference speed and synthesized speech quality. Despite its iterative decoding, Grad-TTS is capable of real-time synthesis. Moreover, it can generate mel-spectrograms twice faster than Tacotron2 while keeping synthesis quality competitive with common TTS baselines.

References

Appendix

We include an appendix with detailed derivations, proofs and additional information. Our proposed diffusion probabilistic framework employs generalized terminal distribution N(μ,Σ)\mathcal{N}(\mu,\Sigma) instead of N(0,I)\mathcal{N}(0,I) as proposed by Song et al. (2021). The derivation for the solution (3) of SDE (2) that transforms the original data distribution to the terminal distribution is described in Appendix A. In Appendix B we also derive the distribution which the solution (3) for the diffused data XtX_{t} follows. Then, the goal of diffusion probabilistic modelling is to reconstruct the reverse-time trajectories of the forward diffusion process, and Song et al. (2021) showed that these dynamics can follow two different differential equations: either SDE (8) proposed by Anderson (1982) or ODE (9). So, Appendix C contains these differential equations for N(μ,Σ)\mathcal{N}(\mu,\Sigma) serving as terminal distribution. They depend on time-dependent gradient field ∇log⁡p0t(Xt∣X0)\nabla\log{p_{0t}(X_{t}|X_{0})} supposed to be modelled using neural network. In order to train it, we show how to compute the gradient in Appendix D.

Exponential of a diagonal matrix is just element-wise exponential, so we can rewrite it in multidimensional form as

or writing this down in terms of XtX_{t}:

where II is n×nn\times n identity matrix.

Let A(s)=βse−12Σ−1∫stβuduA(s)=\sqrt{\beta_{s}}e^{-\frac{1}{2}\Sigma^{-1}\int_{s}^{t}{\beta_{u}du}}. It is a diagonal matrix and its ii-th diagonal element aii(s)a_{ii}(s) equals βse−12σii2∫stβudu\sqrt{\beta_{s}}e^{-\frac{1}{2\sigma^{2}_{ii}}\int_{s}^{t}{\beta_{u}du}}. Assume aii(s)∈L2[0,T]a_{ii}(s)\in L_{2}[0,T] for each ii. Itô’s integral ∫0taii(s)dWsi\int_{0}^{t}{a_{ii}(s)dW_{s}^{i}} is defined as the limit of integral sums when mesh of partition Δ\Delta tends to zero:

where the first equality in distribution holds due to the properties of Brownian motion and the fact that aii(sk)a_{ii}(s_{k}) are deterministic (implying that aii(sk)ΔWski=aii(sk)(Wsk+1i−Wski)a_{ii}(s_{k})\Delta W_{s_{k}}^{i}=a_{ii}(s_{k})(W_{s_{k+1}}^{i}-W_{s_{k}}^{i}) are independent normal random variables with mean and variance aii2(sk)(sk+1−sk)=aii2(sk)Δska^{2}_{ii}(s_{k})(s_{k+1}-s_{k})=a^{2}_{ii}(s_{k})\Delta s_{k}) and the second equality in distribution follows from Lévy’s continuity theorem (it is easy to check that the sequence of characteristic functions of random variables on the left-hand side converges point-wise to the characteristic function of the random variable on the right-hand side). Then, simple integration gives

It implies that in multidimensional case we have:

C Reverse dynamics

The result by Anderson (1982) implies that if nn-dimensional process of the diffusion type XtX_{t} satisfies

where pt(⋅)p_{t}(\cdot) is the probability density function of random variable XtX_{t} and W~t\widetilde{W}_{t} is the reverse-time standard Brownian motion such that XtX_{t} is independent of its past increments W~s−W~t\widetilde{W}_{s}-\widetilde{W}_{t} for s<ts<t. Reverse-time dynamics means that all the integrals associated with reverse-time differentials have tt as their lower limit (e.g. dXtdX_{t} relates to ∫tTdXs=XT−Xt\int_{t}^{T}{dX_{s}}=X_{T}-X_{t}). Anderson’s result is obtained under the assumption that Kolmogorov equations (for probability density functions) associated with all considered processes have unique smooth solutions. On the other hand, Song et al. (2021) argued that SDE (28) has the same forward Kolmogorov equation as the following ODE:

which means that processes following (28) and (30) are equal in distribution if they start from the same initial distribution Law(X0)Law(X_{0}). In our case f(Xt,t)=12Σ−1(Xt−μ)βtf(X_{t},t)=\frac{1}{2}\Sigma^{-1}(X_{t}-\mu)\beta_{t} and g(t)=βtg(t)=\sqrt{\beta_{t}}, so we have two equivalent reverse diffusion dynamics:

where both differential equations are to be solved backwards.

D Score estimation

If X0X_{0} is known, then (27) implies that

where p0t(⋅∣X0)p_{0t}(\cdot|X_{0}) is the probability density function of conditional distribution Law(Xt∣X0)Law(X_{t}|X_{0}). So, if we sample XtX_{t} by the formula Xt=ρ(X0,Σ,μ,t)+ϵtX_{t}=\rho(X_{0},\Sigma,\mu,t)+\epsilon_{t} where ϵt∼N(0,λ(Σ,t))\epsilon_{t}\sim\mathcal{N}(0,\lambda(\Sigma,t)), then ∇log⁡p0t(Xt∣X0)=−λ(Σ,t)−1ϵt\nabla\log{p_{0t}(X_{t}|X_{0})}=-\lambda(\Sigma,t)^{-1}\epsilon_{t}. In the simplified case when Σ=I\Sigma=I we have λ(I,t)=λtI\lambda(I,t)=\lambda_{t}I where λt=1−e−∫0tβsds\lambda_{t}=1-e^{-\int_{0}^{t}{\beta_{s}ds}}. In this case gradient of noisy data log-density reduces to ∇log⁡p0t(Xt∣X0)=−ϵt/λt\nabla\log{p_{0t}(X_{t}|X_{0})}=-\epsilon_{t}/\lambda_{t}. If ϵt=λtξt\epsilon_{t}=\sqrt{\lambda_{t}}\xi_{t}, then we have