ProDiff: Progressive Fast Diffusion Model For High-Quality Text-to-Speech
Rongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu, Chenye Cui, Yi Ren
Introduction
Text-to-speech (TTS) (Shen et al. 2018; Ren et al. 2020; Huang et al. 2022b) aims to generate almost human-like audios using text, which attracts broad interest in the machine learning community. Previous neural TTS models (Wang et al. 2017; Shen et al. 2018; Li et al. 2019) first generate mel-spectrograms autoregressively from text and then synthesize speech from the generated mel-spectrograms using a separately trained vocoder (Kong et al. 2020a; Yamamoto et al. 2020; You et al. 2021; Huang et al. 2021b). They have been demonstrated to generate high-fidelity audio samples yet suffer from expensive computational costs. In recent years, non-autoregressive approaches (Kim et al. 2020; Popov et al. 2021; Cui et al. 2021) are proposed to generate speech audios with satisfactory speed. However, these models are criticized for other problems, e.g., the limited sample quality (Ren et al. 2022) or sample diversity (Ren et al. 2021).
In text-to-speech synthesis, our goal is mainly three-fold:
High-quality: to improve the naturalness of synthesized speech, the model should capture the details (frequency bins between two adjacent harmonics, unvoiced frames, and high-frequency parts) in natural speech.
Fast: high generation speed is essential when considering real-time speech synthesis. This poses a challenge for all high-quality neural synthesizers.
Diverse: to prevent the synthesized speech from being too dull and tedious when generating long speech, the model should be able to reduce mode collapse and avoid unimodal predictions.
As a blossoming class of generative models, denoising diffusion probabilistic models (DDPMs) (Ho et al. 2020; Lam et al. 2021) have emerged to prove the capability to achieve leading performances in both image and audio synthesis (Huang et al. 2022a; Dhariwal and Nichol 2021). However, the current development of DDPMs in speech synthesis is hampered by two major challenges:
While DDPMs inherently are gradient-based models with score matching objectives, a guarantee of high sample quality typically comes at the cost of hundreds to thousands of denoising steps. This prevents models from real-world deployment.
When reducing refinement iterations, diffusion models show distinct degradation in model convergence due to the complex data distribution, leading to the blurry and over-smooth predictions in mel-spectrograms.
In this work, we start with a preliminary study on diffusion parameterization for text-to-speech. We find that the previous dominant diffusion models, which generate samples via estimating the gradient of the data density (denoted as gradient-based parameterization), require hundreds or thousands of iterations to guarantee high perceptual quality. When reducing the sampling steps, an apparent degradation in quality due to perceivable background noise is observed. On the contrary, the approach to parameterizing the denoising model through directly predicting clean data with a neural network (denoted as generator-based parameterization) has demonstrated its advantages in accelerating sampling from a complex distribution.
Based on these preliminary studies, we design better diffusion models for text-to-speech synthesis. In this paper, we propose ProDiff, on progressive fast diffusion model for high-quality text-to-speech; 1) To avoid significant degradation of perceptual quality when reducing reverse iterations, ProDiff directly predicts clean data and frees from estimating the gradient for score matching; 2) To tackle the model convergence challenge with decreased diffusion iterations, ProDiff reduces the data variance in the target side via knowledge distillation. Specifically, the denoising model uses the generated mel-spectrogram from an N-step DDIM teacher as the training target and distills the behavior into a new model with N/2 steps. As such, it allows the TTS model to make sharp predictions and further accelerates sampling by orders of magnitude.
Experimental results demonstrate that ProDiff achieves outperformed sample quality and diversity. ProDiff enjoys an effective sampling process that needs only 2 iterations to synthesize high-fidelity mel-spectrograms, 24x faster than real-time on a single NVIDIA 2080Ti GPU without engineered kernels. To the best of our knowledge, ProDiff makes diffusion models applicable to interactive, real-world speech synthesis applications at a low computational cost for the first time. The main contributions of this work are summarized as follows:
We analyze and compare different diffusion parameterizations in text-to-speech synthesis. Compared to the traditional gradient-based DDPMs with score matching objectives, the models that directly predict clean data show advantages in accelerating sampling from a complex distribution.
We propose ProDiff, on progressive fast diffusion model for high-quality text-to-speech. Unlike estimating the gradient of data density, ProDiff parameterizes the denoising model by directly predicting clean data. To tackle the model convergence challenge in accelerating refinement, ProDiff uses the generated mel-spectrograms with reduced variance as target and makes sharper predictions. ProDiff is distilled from the behavior of the N-step teacher into a new model with N/2 steps, further decreasing the sampling time by orders of magnitude.
Experimental results demonstrate ProDiff needs only 2 iterations to synthesize high-fidelity mel-spectrograms, while it maintains sample quality and diversity competitive with state-of-the-art models using hundreds of steps. It makes diffusion models practically applicable to text-to-speech deployment for the first time.
Background on Diffusion Models
In this section, we introduce the theory of diffusion probabilistic model (Ho et al. 2020; Lam et al. 2021; Song et al. 2020a; Song et al. 2020b). Diffusion and reverse processes are given by diffusion probabilistic models, which could be used for the denoising neural networks to learn data distribution.
With the predefined fixed noise schedule and diffusion step , we compute the corresponding constants respective to diffusion and reverse process:
Diffusion process Similar as previous work (Ho et al. 2020; Lam et al. 2021; Song et al. 2020a), we define the data distribution as . The diffusion process is defined by a fixed Markov chain from data to the latent variable :
For a small positive constant , a small Gaussian noise is added from to the distribution of under the function of .
The whole process gradually converts data to whitened latents according to the fixed noise schedule .
Reverse process The reverse process aims to recover samples from Gaussian noises, which is a Markov chain from to parameterized by shared :
where each iteration eliminate the Gaussian noise added in the diffusion process:
It has been demonstrated that diffusion models (Dhariwal and Nichol 2021; Xiao et al. 2021) can learn diverse data distribution in multiple domains, such as images and time series. However, the main issue with the proposed neural diffusion process is that it requires up to thousands of iterative steps to reconstruct the target distribution during reverse sampling. In this work, we offer a progressive fast conditional diffusion model to reduce reverse iterations and enjoy computational efficiency.
Diffusion model parameterization
In this section, we discuss how to parameterize the reverse denoising model in a way for which the implied prediction. We classify current diffusion parameterization into two classes: 1) the denoising model learns the gradient of the data log density and predicts samples in the space, which we denote as the Gradient-based method. 2) the denoising model directly predicts clean data and optimizes the sample reconstruction error, which we denote as the Generator-based method.
where is the step size, .
A line of works named score matching neural networks (Song et al. 2020b; Song and Ermon 2020) learn the Stein score function , and use Langevin dynamics for inference. For any step , the denoising score matching objective takes the form:
Simultaneously, another line of works named denoising diffusion probabilistic models (DDPMs) (Ho et al. 2020; Dhariwal and Nichol 2021; Kong et al. 2020b) choose to parameterize the denoising model through directly predicting with a neural network .
Most recently, Ho et al. 2020 observe that denoising diffusion probabilistic models (Sohl-Dickstein et al. 2015; Dhariwal and Nichol 2021; Kong et al. 2020b) and score matching neural networks (Song et al. 2020b; Song and Ermon 2020) are closely related, which we denote as the Gradient-based method. In this case, the training loss is usually defined as mean squared error in the space, and efficient training is optimizing a random term of with stochastic gradient descent:
In a gradient-based diffusion model, a guarantee of high sample quality typically comes at the cost of hundreds to thousands of denoising steps, and thus the huge computational costs hinder its application in real-world text-to-speech deployment.
2. Generator-based method
Different from the aforementioned gradient-based diffusion models that required hundreds of steps with small to estimate the gradient for data density, diffusion models (Salimans and Ho 2022; Liu et al. 2022) can be interpreted as parameterizing the denoising model by directly predicting the clean data, which we denote as the Generator-based method.
It is well-known that has different levels of perturbation, and hence using a single (gradient-based parameterization) network to predict directly at different may be difficult. In contrast, the generator-based diffusion model is free from estimating the gradient for data density. It only needs to predict unperturbed and then add back perturbation using the posterior distribution , which has been given in Appendix A.
Specifically, in the generator-based diffusion models, is the implicit distribution imposed by the neural network that outputs given . And then is sampled using the posterior distribution given and the predicted .
In this case, the training loss is defined as mean squared error in the data space, and efficient training is optimizing a random term of with stochastic gradient descent:
Some recent works (Salimans and Ho 2022; Liu et al. 2022) choose to parameterize the denoising model through directly predicting with a neural network . These generator-based methods advantage in accelerating sampling from a complex distribution.
ProDiff
This section presents our proposed ProDiff, on progressive fast diffusion model for high-quality text-to-speech. We first describe the motivation of each design in ProDiff. Secondly, we introduce how to select a diffusion teacher model and distill from it; Then we describe the model architecture and training loss in ProDiff, following with the illustration of training and inference algorithms.
As a blossoming class of generative models, denoising diffusion probabilistic models (DDPMs) have emerged to prove their capability to achieve leading performances in both image and audio synthesis (Dhariwal and Nichol 2021; San-Roman et al. 2021; Kong et al. 2020b; Chen et al. 2020), while several challenges remain for industrial deployment: 1) The most dominant diffusion TTS models estimate the gradient for data density with score matching objectives, where hundreds of iterations are required to guarantee high-quality synthesis. This prevents models from real-world deployment. 2) When reducing refinement iterations, diffusion models show distinct degradation in model convergence due to the complex data distribution. As such, the original denoising model cannot generate deterministic values, leading to blurry and over-smooth predictions in mel-spectrograms.
In ProDiff, we propose two key techniques to complement the above issues: 1) Utilizing the generator-based parameterization to speed up sampling. Through the preliminary comparisons of diffusion parameterization in Section 3, we conclude that the denoising model that predicts clean data has advantages in accelerating sampling from a complex distribution. 2) Reducing data variance in the target side via knowledge distillation. The denoising model uses the generated mel-spectrogram from an N-step DDIM teacher as the training target and distills the behavior into a new model with N/2 steps. As such, it allows the TTS model to make a sharp prediction and further reduces the sampling time by orders of magnitude.
To conclude, ProDiff maintains the sample quality and frees diffusion models from hundreds of iterative refinements, which is for the first time applicable to interactive, real-world applications.
2. Select a teacher
In this subsection, we describe how to select an expected teacher for knowledge distillation. The teacher model is supposed to achieve the fast, high-quality, and diverse text-to-speech synthesis, and thus the distilled student could inherit its powerful capability. Through the preliminary analyses (see Section 3) and experiments (see Section 6.2) on diffusion parameterization, we empirically find that the 4-step generator-based diffusion model strikes a proper balance between perceptual quality and sampling speed. As such, we pick up a generator-based diffusion model with 4 diffusion steps as the teacher.
3. Distill from teacher
Denoising diffusion implicit model (DDIM) (Song et al. 2020a) formulates a non-Markovian generative process that accelerates the inference while keeping the same training procedure as denoising diffusion probabilistic model (DDPM). Inspired by Salimans and Ho 2022, we utility the sampler to directly predict the coarse-grained mel-spectrogram with reduced variance in the diffusion process.
Specifically, we first initialize ProDiff with a copy of the teacher mentioned in Section 4.2, using both the same parameters and model definition. We sample data from the training set and add noise to it as original. Differently, we get the target value for the denoising model by running 2 DDIM sampling steps using the teacher instead of the original data . Further, we halve the required steps by making a single DDIM step of the student match 2 DDIM steps of the teacher. As illustrated in Algorithm 1, the implementation of training the student model with knowledge distillation stays very close to the original Algorithm 4 in the appendix.
4. Architecture
We build the basic architecture upon FastSpeech 2 (Ren et al. 2020), which is one of the most popular models in non-autoregressive TTS. As illustrated in Figure 1, ProDiff consists of a phoneme encoder, variance adaptor, and spectrogram denoiser. The phoneme encoder converts a sequence of phoneme embedding into a hidden sequence. Then, the variance adaptor predicts the duration of each phoneme to regulate the length of the hidden sequence into the length of speech frames. Furthermore, the variance adaptor indicates different variances in speech, such as pitch and energy. Finally, the spectrogram denoiser iteratively refines the length-regulated hidden sequence into mel-spectrograms. More details have been attached in Appendix C.
Encoder and Variance Adaptor The phoneme encoder is composed of feed-forward transformer blocks (FFT blocks) based on the transformer architecture (Vaswani et al. 2017). The encoder comprises a pre-net, transformer blocks with multi-head self-attention, and the final linear projection layer. In variance adaptor, the duration, pitch, and energy predictors share a similar model structure consisting of a 2-layer 1D-convolutional network with ReLU activation, each followed by the layer normalization and the dropout layer, and an extra linear layer to project the hidden states into the output sequence.
Spectrogram Denoiser Following Liu et al. 2021, we adopt a non-causal WaveNet (Oord et al. 2016) architecture to be our spectrogram denoiser. The decoder comprises a 1x1 convolution layer and convolution blocks with residual connections to project the input hidden sequence with 256 channels. For any step , we use the cosine schedule .
5. Training Loss
Several objectives have been used to optimize the ProDiff model.
Sample Reconstruction Loss Instead of using the original clean data in training ProDiff, we get the target value with reduced variance by running 2 DDIM sampling steps of the teacher:
Structural Similarity Index (SSIM) Loss Structural Similarity Index (SSIM) (Wang et al. 2004) is one of the state-of-the-art perceptual metrics to measure image quality, which can capture structural information and texture. The value of SSIM is between 0 and 1, where 1 indicates perfect perceptual quality relative to the ground truth. Inspired by (Ren et al. 2022), we adopt structural similarity index (SSIM) loss in training TTS models:
Variance Reconstruction Loss To promote the naturalness and expressiveness in generated speech, we provide more acoustic variance information including pitch, duration and energy. Variance reconstruction losses are also added to train the acoustic generator.
where we use , and to denote the target duration, energy and pitch, respectively, and use , and to denote the corresponding predicted values. The loss weights are all set to be 0.1.
6. Training and Inference Procedures
The training and sampling algorithm of ProDiff have been illustrated in Algorithm 1 and Algorithm 2, respectively.
The final loss term in training ProDiff consist of the following parts: 1) sample reconstruction loss : MSE between the predicted and the target mel-spectrogram according to Eq. 10; 2) structural similarity index (SSIM) loss : one minus the SSIM index between the predicted and the target mel-spectrogram according to Eq. 11. 3) variance reconstruction loss, : MSE between the predicted and the target phoneme-level duration, pitch spectrogram, and energy value according to Eq. 12;
Inference
In inference, ProDiff iteratively predicts unperturbed and then adds back perturbation using the posterior distribution, and thus finally it generates high-fidelity mel-spectrograms with increasing details. Specifically, the denoising model first predicts , and then is sampled using the posterior distribution given and the predicted . In final, the generated spectrogram is transformed to waveforms using a pre-trained vocoder. The posterior distribution has been given in Appendix A.
Related works
Text-to-speech (TTS) models convert input text or phoneme sequence into mel-spectrogram (e.g., Tacotron (Wang et al. 2017), FastSpeech (Ren et al. 2019)), which is then transformed to waveform using a separately trained vocoder (Kong et al. 2020a; Huang et al. 2021a), or directly generate waveform from text (e.g., EATS (Donahue et al. 2020) and VITS (Kim et al. 2021)). Early autoregressive models (Wang et al. 2017; Li et al. 2019) sequential generate a sample and suffer from slow inference speed. Several works (Ren et al. 2019; Kim et al. 2020) have been proposed to generate mel-spectrogram frames in parallel, which speed up mel-spectrogram generation over autoregressive TTS models, while preserving the quality of synthesized speech. Recently proposed Diff-TTS (Jeong et al. 2021), Grad-TTS (Popov et al. 2021), and DiffSpeech (Liu et al. 2021) inherently are gradient-based models with score matching objectives to generate high-quality samples, while the iterative sampling process costs hinder their applications to text-to-speech deployment. Unlike the dominant gradient-based TTS models mentioned above, ProDiff directly predicts clean data and frees from estimating the gradient for score matching. The proposed model avoids significant perceptual quality degradation when reducing reverse iterations, and thus it maintains sample quality competitive with state-of-the-art models using hundreds of steps.
2. Diffusion Probabilistic Models
Denoising diffusion probabilistic models (DDPMs) (Ho et al. 2020; Song et al. 2020a; Lam et al. 2021) are likelihood-based generative models that have recently succeeded to advance the state-of-the-art results in several important domains including image synthesis (Dhariwal and Nichol 2021; Song et al. 2020a), audio synthesis (Huang et al. 2022a), and 3D point cloud generation (Luo and Hu 2021), and have proved its capability to produce high-quality samples. One major drawback of diffusion or score-based models is the slow sampling speed due to many iterative steps. To alleviate this issue, multiple methods have been proposed, including learning an adaptive noise schedule (Lam et al. 2022), introducing non-Markovian diffusion processes (Song et al. 2020a), and using adversarial learning for reducing iterations (Xiao et al. 2021). To conclude, all these methods focus on the image domain, while audio data is different for its long-term dependencies and strong condition. Liu et al. 2022 proposes a denoising diffusion generative adversarial networks (GANs) to achieve high-fidelity and efficient text-to-spectrogram synthesis. Differently, our work focuses on designing progressive fast diffusion models for text-to-speech synthesis without unstable adversarial learning procedure, which has been relatively overlooked.
3. Knowledge distillation
Knowledge distillation has been demonstrated for its efficiency in simplifying the data distribution. In non-autoregressive machine translation (Gu et al. 2017), sequence-level knowledge distillation (Kim and Rush 2016) has achieved good performance in transferring the knowledge from the teacher model to the student. In non-autoregressive text-to-speech synthesis (Ren et al. 2019), researchers alleviate the one-to-many mapping problem by using the generated mel-spectrogram from an autoregressive teacher model as the training target. Recently, Salimans and Ho 2022 utilize knowledge distillation and iteratively halve the diffusion steps of the continuous diffusion model, accelerating iterative refinement to a large extent. In contrast, ProDiff with discrete schedules adopts knowledge distillation to reduce data variance and promote training convergence.
Experiments
Dataset For a fair and reproducible comparison against other competing methods, we use the benchmark LJSpeech dataset (Ito 2017). LJSpeech consists of 13,100 audio clips of 22050 Hz from a female speaker for about 24 hours in total. We convert the text sequence into the phoneme sequence with an open-source grapheme-to-phoneme conversion tool (Sun et al. 2019) https://github.com/Kyubyong/g2p. Following the common practice (Chen et al. 2021; Min et al. 2021), we conduct preprocessing on the speech and text data: 1) extract the spectrogram with the FFT size of 1024, hop size of 256, and window size of 1024 samples; 2) convert it to a mel-spectrogram with 80 frequency bins; and 3) extract F0 (fundamental frequency) from the raw waveform using Parselmouth tool https://github.com/YannickJadoul/Parselmouth.
Model Configurations ProDiff consists of 4 feed-forward transformer blocks for the phoneme encoder. In the pitch encoder, the size of the lookup table and encoded pitch embedding are set to 300 and 256. The hidden channel is set to 256. In the denoiser, we set to stack 20 layers of convolution with the kernel size 3, and we set the dilated factor to 1 (without dilation) at each layer. We have attached more detailed information on the model configuration in Appendix C.
Training and Evaluation We train the ProDiff teacher with diffusion steps and take the converged teacher to train ProDiff with diffusion steps via knowledge distillation. The diffusion probabilistic models have been trained for 200,000 steps using 1 NVIDIA 2080Ti GPU with a batch size of 64 sentences. The adam optimizer is used with . We utilize HiFi-GAN(Kong et al. 2020a) (V1) as the vocoder to synthesize waveform from the generated mel-spectrogram in all our experiments. To evaluate the perceptual quality, we conduct crowd-sourced human evaluations with MOS (mean opinion score) on the testing set via Amazon Mechanical Turk, which is rated from 1 to 5 and reported with the 95% confidence intervals (CI). We further include objective evaluation metrics, such as MCD (Kubichek 1993), STOI (Taal et al. 2010), and PESQ (Rix et al. 2001) to evaluate the compatibility between the spectra of two audio sequences. To evaluate the sampling speed, we implement real-time factor (RTF) assessment on a single NVIDIA 2080Ti GPU. In addition, we employ two metrics NDB and JS (Richardson and Weiss 2018) to explore the diversity of generated mel-spectrograms. More information on evaluation has been attached in Appendix G.
2. Preliminary Analyses on Diffusion Parameterization
We compare and examine both choices (i.e., gradient-based or generator-based diffusion parameterization) with varying diffusion steps in Section 3, and conduct evaluations in terms of quality, diversity, and sampling speed. The results have been shown in Table 1, and we also visualize the mel-spectrograms generated in Appendix H. We have some observations from the results: 1) With a large distribution of noise schedule, gradient-based or generator-based diffusion models could synthesize high-fidelity speech samples with similar results. 2) When reducing iterative steps (), a distinct degradation due to the perceivable background noise is observed in gradient-based TTS models. In contrast, the generator-based TTS models which parameterize the denoising model by directly predicting clean data, still maintain sample quality and avoid significant degradation, which is expected for the following reasons:
It is well-known that noisy data has different levels of perturbation in diffusion models, and hence using a neural network to predict directly at different may be difficult. As such, the gradient-based model requires hundreds of diffusion steps with small to estimate the gradient of the data density, and the distinct degradation is observed when accelerating refinement iterations.
In contrast, the generator-based diffusion model is free from estimating the gradient for data density, which only needs to predict unperturbed and then add back perturbation using the posterior distribution . Therefore, the generator-based diffusion parameterization advantages in avoiding significant perceptual quality degradation when accelerating sampling from a complex distribution.
3. Performances
We compare the quality of generated audio samples, inference latency, and sample diversity with other systems, including 1) GT, the ground-truth audio; 2) GT (voc.), where we first convert the ground-truth audio into mel-spectrograms and then convert them back to audio using HiFi-GAN (V1) (Kong et al. 2020a); 3) Tacotron 2 (Shen et al. 2018): the traditional autoregressive TTS model; 4) FastSpeech 2 (Ren et al. 2020): one of the most popular non-autoregressive TTS models; 5) Glow-TTS (Kim et al. 2020): the TTS model with the normalizing flow and monotonic alignment search; 6) GANSpeech (Yang et al. 2021): the TTS model with generative adversarial networks; 7) Grad-TTS (Popov et al. 2021) and DiffSpeech (Popov et al. 2021): two denoising diffusion probabilistic models. The results are compiled and presented in Table 2, and we have the following observations:
Audio Quality In terms of audio quality, ProDiff achieves high perceptual quality with a gap of compared to the ground truth audio. It matches the SOTA DDPMs using hundreds of steps and outperforms other non-autoregressive baselines. For objective evaluation, ProDiff also demonstrates the outperformed performance in MCD, PESQ, and STOI, superior to all baseline models.
Sampling Speed ProDiff enjoys an effective sampling process that needs only 2 iterations to synthesize high-fidelity spectrograms, 24x faster than real-time on a single NVIDIA 2080Ti GPU. ProDiff significantly reduces the inference time compared with the competing diffusion models (Grad-TTS and DiffSpeech).
We visualize the relationship between the inference latency and the length of phoneme sequence. Figure 3(a) shows that the inference latency of ProDiff barely increases with the length of input phoneme, which is almost constant at 20ms. Unlike Tacotron 2, Grad-TTS, and DiffSpeech which linearly increase with the length due to autoregressive sampling or a large number of iterations, ProDiff has a similar scaling performance to FastSpeech 2, making diffusion models practically applicable to text-to-speech deployment for the first time.
Sample Diversity Previous diffusion models (Dhariwal and Nichol 2021; Xiao et al. 2021) in image generation have demonstrated the outperformed sample diversity, while the comparison in the speech domain is relatively overlooked. Similarly, we can intuitively infer that diffusion probabilistic models are good at generating diverse samples. To verify our hypothesis, we employ two metrics NDB and JS (Richardson and Weiss 2018) to explore the diversity of generated mel-spectrograms. As shown in Table 2, we can see that ProDiff achieves higher scores than several one-shot methods, while the GAN-based method generates samples with minimal diversity, which is expected for the following reasons:
1) It is well-known that mode collapse (Creswell et al. 2018) easily appears in the strongly conditional generation task in one-shot generative models, leading to very similar output samples from a single or few modes of the distribution.
2) In contrast, diffusion models (Grad-TTS, DiffSpeech, and ProDiff) are meant to reduce mode collapse. They break the generation process into several conditional diffusion steps where each step is relatively simple to model. Thus, we expect our model to exhibit better training stability and mode coverage.
Besides, we conduct a quality comparison on the multi-speaker dataset and obtain similar conclusions as above. (see Appendix E). We also conduct robustness evaluation on both single-speaker and multi-speaker datasets in Appendix F and find that ProDiff achieves comparable robustness performance with state-of-the-art non-autoregressive TTS models.
4. Visualizations
As illustrated in Figure 2, we then visualize the mel-spectrograms generated by the above systems given the same text sequence and have the following observations: Non-probabilistic models (Tacotron 2, FastSpeech 2) tend to generate less-diverse samples of blurry and over-smooth mel-spectrograms, and the GAN-based model (GANSpeech) suffers from model collapse with very similar output samples. Differently, diffusion probabilistic models (Grad-TTS, DiffSpeech, ProDiff) generate mel-spectrograms with rich frequency details, resulting in the natural and expressive sounds. ProDiff is demonstrated to make a sharp prediction with knowledge distillation, and it maintains sample quality competitive with state-of-the-art models using hundreds of steps.
5. Progressive Diffusion
To evaluate the efficiency of diffusion models in reverse sampling, we report MCD results obtained after each iteration. As shown in Figure 3(b), we compare against the highly optimized stochastic baseline sampler DiffSpeech: 1) DiffSpeech slowly refines the coarse-grained mel-spectrogram from Gaussian noise, and a guarantee of high sample quality comes at the cost of hundreds of denoising steps; 2) In contrast, ProDiff Teacher and ProDiff predict unperturbed and then add back perturbation using the posterior distribution, which produces near-optimal results more faster and effective, with attractive solutions for computational budgets that allow fewer reverse iterations.
6. Ablation Studies
We conduct ablation studies to demonstrate the effectiveness of several key techniques in ProDiff, including the generator-based diffusion parameterization and knowledge distillation. The results of both subjective and objective evaluations have been presented in Table 3, and we have the following observations: 1) With limited diffusion iterations, replacing the generator-based diffusion parameterization by gradient-based parameterization causes a distinct degradation in perceptual quality. ProDiff directly predicts clean data to avoid significant degradation when reducing reverse iterations. 2) Removing the knowledge distillation mechanism and using clean data as training target results in blurry and over-smooth predictions, which demonstrates the effectiveness and efficiency of the proposed distillation in reducing data variance and promoting model convergence.
Furthermore, we compare to distill the behavior of a teacher with varying diffusion steps (i.e., 16, 8, and 4) into a 2-step student. Take the 16-step teacher as an example, we need 3 distillation procedures to obtain the student (16->8->4->2), where the student model iteratively becomes the teacher in the following distillation. However, the perceptual quality of generated samples slightly drops after a series of distillations, indicating that the quality gain from sharper prediction could hardly cover the degradation in reducing iterations. In summary, distilling knowledge from the 4-step teacher could be an optimal choice to strike a balance between computational cost and sample quality.
Conclusion
In this work, we proposed ProDiff, on progressive fast diffusion model for high-quality text-to-speech. The preliminary study on diffusion model parameterization found that previous gradient-based TTS models required hundreds of iterations to guarantee high sample quality, which posed a challenge for accelerating sampling. In contrast, ProDiff parameterized the denoising model by directly predicting the clean data to avoid significant degradation of perceptual quality when reducing reverse iterations. To tackle the model convergence challenge in accelerating refinement, ProDiff adopted the synthesized mel-spectrogram from teacher as target to reduce the data variance and make a sharp prediction. As such, ProDiff was distilled from the behavior of an N-step teacher into a new model with N/2 steps, further reducing the sampling time by orders of magnitude.
Experimental results demonstrated that ProDiff needed only 2 iterations to synthesize high-fidelity mel-spectrograms, while it maintained sample quality and diversity competitive with state-of-the-art models using hundreds of steps. To the best of our knowledge, ProDiff made diffusion models for the first time applicable to interactive, real-world text-to-speech with a low computational cost. Our extensive ablation studies demonstrated that each design in ProDiff was effective, and we showed that our model could be easily extended to a multi-speaker setting. We envisage that our work could serve as a basis for future text-to-speech studies.
Acknowledgements
The author would like to thank Luping Liu and Max W.Y. Lam for the helpful discussions over the initial ideas. This work was supported in part by the Zhejiang Natural Science Foundation LR19F020006 and National Key R&D Program of China under Grant No.2020YFC0832505.
References
Appendix A Diffusion Posterior Distribution
Firstly we compute the corresponding constants respective to diffusion and reverse process:
The Gaussian posterior in diffusion process is defined through the Markov chain, where each iteration adds Gaussian noise. Consider the forward diffusion process in Eq. 2, which we repeat here:
We emphasize the property observed by (Ho et al. 2020), the diffusion process can be computed in a closed form:
Applying Bayes’ rule, we can obtain the forward process posterior when conditioned on
Appendix B Distill Target
Inspired by (Salimans and Ho 2022), to promote the perceptual quality of generated sample, ProDiff reduces the data variance in the target side via knowledge distillation. The denoising model use the generated mel-spectrogram from an N-step DDIM teacher as the training target, and distill the behavior into a new model with N/2 steps.
In this section, the student model needs to predict the specific target in order to match the teacher during sampling. Here we derive what this target needs to be: When training the N-step student, we have that the teacher model samples the next set of noisy data given the current noisy data by taking two steps of DDIM. The student tries to sample the same value in only one step of DDIM. Denoting the student denoising prediction by , and its one-step sample by , we gives:
Appendix C Architecture
We list the model hyper-parameters of ProDiff in Table 4.
Appendix D Training algorithm
Appendix E Results on Multi-Speaker Dataset
For the multi-speaker setting, the train-clean-100 subset of the LibriTTS corpus (Zen et al. 2019) is used, which consists of audio recordings of 247 speakers with a total duration of about 54 hours. We add speaker embedding after the encoder and before the spectrogram denoiser, following the common practice (Kim et al. 2020; Min et al. 2021). The results have been compiled in Table 5. Similar to the experimental results on LJSpeech, we can draw conclusions that ProDiff achieve outperformed quality in terms of both subjective and objective evaluation, even in more complicated (multi-speaker) scenarios.
Appendix F Robustness Evaluation
We conduct the robustness evaluation on LJSpeech and LibriTTS datasets. We select 50 sentences that are particularly hard for TTS systems following FastSpeech (Ren et al. 2019). The results are shown in Tables 6 and 7. We can see that ProDiff achieve comparable robustness performance with state-of-the-art non-autoregressive TTS models.
Appendix G Evaluation Matrix
All our Mean Opinion Score (MOS) tests are crowdsourced and conducted by native speakers. We refer to the rubric for MOS scores in (Protasio Ribeiro et al. [n.d.]), and the scoring criteria has been included in Table 8 for completeness. The samples are presented and rated one at a time by the testers.
G.2. Objective Evaluation
Mel-cepstral distortion (MCD) (Kubichek 1993) measures the spectral distance between the synthesized and reference mel-spectrum features.
Perceptual evaluation of speech quality (PESQ) (Rix et al. 2001) and The short-time objective intelligibility (STOI) (Taal et al. 2010) assesses the denoising quality for speech enhancement.
Number of Statistically-Different Bins (NDB) and Jensen-Shannon divergence (JSD) (Richardson and Weiss 2018). They measure diversity by 1) clustering the training data into several clusters, and 2) measuring how well the generated samples fit into those clusters.