FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

Rongjie Huang, Max W. Y. Lam, Jun Wang, Dan Su, Dong Yu, Yi Ren, Zhou Zhao

Introduction

With the recent development of deep generative models, speech synthesis has seen an extraordinary progress. Among the conventional speech synthesis methods, WaveNets Oord et al. 2016 were demonstrated to generate high-fidelity audio samples in an autoregressive manner yet suffering from prohibitively expensive computational costs. In contrast, non-autoregressive approaches such as flow-based and GAN-based models Prenger et al. 2019; Jang et al. 2021; Kong et al. 2020a; Huang et al. 2021 were also proposed to generate speech audios with satisfactory speed. However, these models were still criticized for other problems, e.g., the limited sample quality or sample diversity Xiao et al. 2021.

In speech synthesis, our goal is mainly two-fold:

High-quality: generating high-quality speech is a challenging problem especially when the sampling rate of an audio is high. It is vital to reconstruct details at different timescales for waveforms of highly variable patterns.

Fast: high generation speed is essential when considering real-time speech synthesis. This poses a challenge for all high-quality neural synthesizers.

As a blossoming class of generative models, denoising diffusion probabilistic models (DDPMs) Ho et al. 2020; Song et al. 2020a; Lam et al. 2022; Liu et al. 2022 has emerged to prove its capability to achieve leading performances in both image and audio syntheses Dhariwal and Nichol 2021; San-Roman et al. 2021; Kong et al. 2020b; Chen et al. 2020; Lam et al. 2022. However, current development of DDPMs in speech synthesis was hampered by two major challenges:

Different from other existing generative models, diffusion models are not trained to directly minimize the difference between the generated audio and the reference audio, but to de-noise a noisy sample given an optimal gradient. This in practice could lead to overly de-noised speech after a large number of sampling steps, in which natural voice characteristics including breathiness and vocal fold closure are removed.

While DDPMs inherently are gradient-based models, a guarantee of high sample quality typically comes at a cost of hundreds to thousands of de-noising steps. When reducing the sampling steps, an apparent degradation in quality due to perceivable background noise is observed.

In this work, we propose FastDiff, a fast conditional diffusion model for high-quality speech synthesis. To improve audio quality, FastDiff adopts a stack of time-aware location-variable convolutions of diverse receptive field patterns to efficiently model long-term time dependencies with adaptive conditions. To accelerate the inference procedure, FastDiff also includes a noise schedule predictor, which derives a short and effective noise schedule and significantly reduces the de-noising steps. Based on FastDiff, we also introduce an end-to-end phoneme-to-waveform synthesizer FastDiff-TTS, which simplifies the text-to-speech generation pipeline and does not require intermediate features or specialized loss functions to enjoy low inference latency.

Experimental results demonstrated that FastDiff achieved a higher MOS score than the best publicly available models and outperformed the strong WaveNet vocoder (MOS: 4.28 vs. 4.20). FastDiff further enjoys an effective sampling process and only needs 4 iterations to synthesize high-fidelity speech, 58x faster than real-time on a V100 GPU without engineered kernels. To the best of our knowledge, FastDiff is the first diffusion model with a sampling speed comparable to previous for the first time applicable to interactive, real-world speech synthesis applications at a low computational cost. FastDiff-TTS successfully simplify the text-to-speech generation pipeline and outperform competing architectures.

Background: Denoising Diffusion Probabilistic Models

In data distribution as q(x0)q({\bm{x}}_{0}), the diffusion process is defined by a fixed Markov chain from data x0{\bm{x}}_{0} to the latent variable xT{\bm{x}}_{T}. For a small positive constant βt\beta_{t}, a small Gaussian noise is added from xt{\bm{x}}_{t} to the distribution of xt−1{\bm{x}}_{t-1} under the function of q(xt∣xt−1)q({\bm{x}}_{t}|{\bm{x}}_{t-1}). The whole process gradually converts data x0{\bm{x}}_{0} to whitened latents xT{\bm{x}}_{T} according to the fixed noise schedule β1,⋯ ,βT\beta_{1},\cdots,\beta_{T}. The reverse process is a Markov chain from xT{\bm{x}}_{T} to x0{\bm{x}}_{0} parameterized by a shared θ\theta, which aims to recover samples from Gaussian noises though eliminating the Gaussian noise added in the diffusion process in each iteration.

It has been demonstrated that diffusion probabilistic models Dhariwal and Nichol 2021; Xiao et al. 2021 can learn diverse data distribution in multiple domains, such as images and time series. While the main issue with the proposed neural diffusion process is that it requires up to thousands of iterative steps to reconstruct the target distribution during reverse sampling. In this work, we offer a fast conditional diffusion model to reduce reverse iterations and improve computational efficiency.

FastDiff

This section presents our proposed FastDiff, a fast conditional diffusion model for high-quality speech synthesis. We first describe the motivation of the design in FastDiff. Secondly, we introduce the iterative refinement model θ\theta for high-quality speech synthesis and the noise predictor ϕ\phi for accelerated sampling. Furthermore, we describe the training and inference procedures in detail. At last, we extend FastDiff to FastDiff-TTS for fully end-to-end text-to-speech syntheses.

While denoising diffusion probabilistic models have shown high potential in synthesizing high-quality speech samples Chen et al. 2020; Kong et al. 2020b; Liu et al. 2021, several challenges remain for industrial deployment: 1) Different from the traditional generative models, diffusion models catch dynamic dependencies from noisy audio instead of clean ones, which introduce more variation information (i.e, noise levels) in addition to the spectrogram fluctuation. 2) With limited receptive field patterns, a distinct degradation could emerge when reducing the reverse iterations, making diffusion models difficult to get accelerated. As a result, hundred or thousand orders of iterations prevent existing diffusion models from real-world deployment.

In FastDiff, we propose two key components to complement the above issues: 1) FastDiff adopts a time-aware location-variable convolution to catch the details of noisy samples at dynamic dependencies. The convolution operations are conditioned on dynamic variations in speech including diffusion steps and spectrogram fluctuations, equipping the model with diverse receptive field patterns and promoting the robustness of diffusion models during reverse acceleration. 2) To accelerate the inference procedure, FastDiff adopts a noise schedule predictor to reduce the number of reverse iterations, frees diffusion models from hundreds or thousands of refinement iterations. This makes FastDiff for the first time applicable to interactive, real-world applications at a low computational cost.

2 Time-Aware Location-Variable Convolution

In comparison with traditional convolution networks, location-variable convolution Zeng et al. 2021 shows efficiency in modeling the long-term dependency of audio and gets neural network free from a significant number of dilated convolution layers. Inspired by this, we introduce the Time-Aware Location-Variable Convolution, which is sensitive to time steps in diffusion probabilistic models. At time step tt, we follow Vaswani et al. 2017 to embed the step index into an 128-dimensional positional encoding (PE) vector et{\bm{e}}_{t}:

In time-aware location-variable convolution, FastDiff requires multiple predicted variation-sensitive kernels to perform convolutional operations on the associated intervals of input sequence. These kernels should be time-aware and sensitive to variations of noisy audio including diffusion steps and acoustic features (i.e., Mel-spectrogram). Therefore, we propose a time-aware location-variable convolution (LVC) module, which is coupled with a kernel predictor as shown in Figure 1(b) and Figure 1(c). We describe the overall calculations below.

Next, we perform convolutional operations on the associated intervals of input sequence using the kernels generated by a kernel predictor α\alpha:

where Ft,Gt\bm{F}_{t},\bm{G}_{t} denote the filter and the gate kernels for xti{\bm{x}}_{t}^{i}, respectively, ∗* denotes the 1d convolution, ⊙\odot denotes the element-wise product and concat⁡(⋅)\operatorname{concat}(\cdot) denotes the concatenation between vectors. Since the time-aware kernels are adaptive to the noise-level and dependent to the acoustic features, FastDiff is capable of precisely estimating de-noising gradient with a superior speed given a noisy signal input.

3 Accelerated Sampling

3.2 Schedule Alignment

In FastDiff, similar to DDPMs, during training we use T=1000T=1000 discrete time steps. Therefore, when needed to condition on tt during sampling, we also need to approximate TmT_{m} discrete time indices by aligning the TmT_{m}-step sampling noise schedule β^\hat{{\bm{\beta}}} to the TT-step training noise schedule β{\bm{\beta}}, with N<<TN<<T. We have attached the detailed algorithms in Appendix C.

4 Training, Noise Scheduling and Sampling

All illustrated in Algorithm 1, we separately parameterize FastDiff by two modules: 1) a iterative refinement model θ\theta that minimizes a variational bound of the score function, and 2) a noise predictor ϕ\phi that optimizes the noise schedule for a tighter evidence lower bound. For inference, we first derive the tighter and more efficient noise schedules β^\hat{\beta} via an one-shot noise scheduling procedure, which makes FastDiff achieve orders of magnitude faster at sampling. It has been demonstrated Lam et al. 2022 that the noise schedule searched for as few as 1 sample could be robust enough to maintain a high-quality generation among all samples in testing set. Secondly, we map the continuous noise schedules to discrete time indexes TmT_{m} using schedule alignment. Finally, FastDiff iteratively refines gaussian noise to generate high-quality samples with computational efficiency. The detailed information on training, noise scheduling and inference procedures has been presented in Appendix C.3.

5 FastDiff-TTS

Existing text-to-speech methods usually adopt a two-stage pipeline: 1) A text-to-spectrogram generation module (a.k.a. acoustic model) aims to generate prosodic attributes according to variance prediction; 2) A conditional waveform generation module (a.k.a. vocoder) adds the phase information and synthesizes a detailed waveform. To further simplify the text-to-speech synthesis pipeline, we propose a fully end-to-end model FastDiff-TTS, which does not require intermediate features or specialized loss functions. FastDiff-TTS is designed to be a fully differentiable and efficient architecture that directly produces waveforms from contexts (e.g. phonemes) without needing to generate acoustic features (e.g., Mel-spectrograms) explicitly.

The architecture design of FastDiff-TTS refers to a convectional non-autoregressive text-to-speech model – FastSpeech 2 Ren et al. 2020 as the backbone. The architecture of FastDiff-TTS is illustrated in Figure 1(d). In FastDiff-TTS, the encoder first converts the phoneme embedding sequence into the phoneme hidden sequence. Then, the duration predictor expands the encoder output to match the length of the desired waveform output. Given the aligned sequence, the variance adaptor adds pitch information into the hidden sequence. Note that it is difficult to use the full audio corresponding to the full text sequence for training due to the typically high sampling rate for high-fidelity waveform (i.e., 24,000 samples per second) and the limited GPU memory. Therefore, we sample a small segment to synthesize the waveform before passing to the FastDiff model. Finally, the FastDiff model decodes the adapted hidden sequence into a speech waveform as in the vocoder task.

5.2 Training Loss

FastDiff-TTS does not require specialized loss functions and adversarial training to improve sample quality as suggested by the previous works Ren et al. 2020; Donahue et al. 2020; Kim et al. 2021. This, to a large extend, simplifies the text-to-speech generation. The final training loss consists of the following terms: 1) a duration prediction loss LdurL_{\text{dur}}: the mean squared error between the predicted and the ground-truth word-level duration in log-scale, 2) a diffusion loss LdiffL_{\text{diff}}: the mean squared error between the estimated and gaussian noise, and 3) a pitch reconstruction loss LpitchL_{\text{pitch}}: the mean squared error between the predicted and the ground-truth pitch sequences. We empirically found that the pitch reconstruction loss LpitchL_{\text{pitch}} is helpful for handling the one-to-many mapping issue in text-to-speech generation.

Related Works

Text-to-speech (TTS) systems aim to synthesize raw speech waveforms from given text. In recent years, Neural network based TTS Ren et al. 2020; Kim et al. 2020; Liu et al. 2021 has made huge progress and attracted a lot of attention in the machine learning and speech community.

Neural vocoder plays the most important role in the recent success of speech synthesis, which require diverse receptive field patterns to catch audio dependencies: 1) autoregressive model WaveNet Oord et al. 2016 requires causal convolutions layers and large filters to increase the receptive field while suffering from slow inference speed. 2) Flow-based generative models Prenger et al. 2019 fully utilize modern parallel computing processors to broaden corresponding receptive fields and speed-up sampling, while they usually achieve a limited sample quality. 3) Generative adversarial networks (GANs) Jang et al. 2021; Kong et al. 2020a are one of the most dominant deep generative models in audio generation. UnivNet Jang et al. 2021 has demonstrated its success in using local-variable convolution on different waveform intervals, and HIFI-GAN Kong et al. 2020a proposes multi-receptive field fusion (MRF) to model the periodic patterns matters. However, GAN-based models are often difficult to train, collapsing Creswell et al. 2018 without carefully selected hyperparameters and regularizers, and showing less sample diversity. 4) Recently proposed diffusion models Diffwave Kong et al. 2020b and WaveGrad Chen et al. 2020 could generate high-quality speech samples, while suffering from a distinct degradation when reducing reverse iterations, making diffusion models difficult to get accelerated. Different from vocoders mentioned above, FastDiff improves the robustness of conditional diffusion model by catching the details of noisy samples at dynamic dependencies, and reduces reverse iterations with predicted noise schedule. The proposed conditional diffuion model allows the high-quality speech synthesis with computational efficiency.

Another important line of work covers directly waveform generation from text: FastSpeech 2s Ren et al. 2020 and VITS Kim et al. 2021 adopt adversarial training process and spectral losses for improving audio quality, while they do not take full advantage of end-to-end training. Recently proposed WaveGrad 2 Chen et al. 2021 estimates the gradient of the log conditional density of the waveform given a phoneme sequence, but suffers from a large model footprint and slow inference. Unlike all of the aforementioned methods, as highlighted in section 3.5, FastDiff-TTS is a fully differentiable and efficient architecture that produces waveforms directly without generating middle features (e.g., spectrograms) explicitly. In additional, our diffuion probabilistic model gets free from hundred or thousands of iterations and enjoy computational efficiency.

Experiments

For a fair and reproducible comparison against other competing methods, we used the benchmark LJSpeech dataset Ito 2017. LJSpeech consists of 13,100 audio clips of 22050 Hz from a Female speaker with about 24 hours in total. To evaluate the generalization ability of our model over unseen speakers in multi-speaker scenarios, we also used the VCTK dataset Yamagishi et al. 2019, which was downsampled to 22050 Hz to match the sampling rate with the LJSpeech datset. VCTK consists of approximately 44,200 audio clips uttered by 109 native English speakers with various accents. For both datasets, we used 80-band Mel-spectrograms as the condition for the vocoding task. The FFT size, window size, and hop size were, respectively, set to 1024, 1024, and 256.

1.2 Model Configurations

FastDiff mainly consists of the refinement model θ\theta and noise schedule predictor ϕ\phi. The refinement model θ\theta comprises three Diffusion-UBlock and DBlock with the upsample or downsample rate of $,respectively.WeadoptalightweightGALRnetworkeffectiveinseparatingtheaddedgaussiannoisefromaudioasthenoiseschedulepredictor, respectively. We adopt a lightweight GALR network effective in separating the added gaussian noise from audio as the noise schedule predictor\phi$. For end-to-end text-to-speech generation, FastDiff-TTS follows the basic structure in FastSpeech 2 Ren et al. 2020, which consists of 4 feed-forward transformer blocks in the phoneme encoder. More details have been attached in the Appendix B.

1.3 Training and Evaluation

The complete training pipeline has been illustrated in Algorithm 1: FastDiff was trained with constant learning rate lr=2×10−4lr=2\times 10^{-4}. The refinement model θ\theta and noise predictor ϕ\phi were trained for 1M and 10K steps until convergence, respectively. FastDiff-TTS was trained until 500k steps using the AdamW optimizer with β1=0.9,β2=0.98,ϵ=10−9\beta_{1}=0.9,\beta_{2}=0.98,\epsilon=10^{-9}. Both models were trained on 4 NVIDIA V100 GPUs using random short audio clips of 16,000 samples from each utterance with a batch size of 16 each GPU. More details have been attached in the Appendix C.

We crowd-sourced 5-scale MOS tests via Amazon Mechanical Turk to evaluate the audio quality. The MOS scores were recorded with 95% confidence intervals (CI). Raters listened to the test samples randomly, where they were allowed to evaluate each audio sample once. We further include additional objective evaluation metrics including STOI and PESQ to test sample quality. To evaluate the sampling speed, we implemented real-time factor (RTF) accessment on a single NVIDIA V100 GPU. In addition, we employed two metrics NDB and JSD to explore the diversity of generated mel-spectrograms. More information about both objective and subjective evaluation has been attached in Appendix D.

2 Comparsion with other models

We compared our FastDiff in audio quality, diversity, and sampling speed with competing models, including 1) WaveNet Oord et al. 2016, the autoregressive generative model for raw audio. 2) WaveGlow Prenger et al. 2019, non-autoregressive flow-based model. 3) HIFI-GAN V1 Kong et al. 2020a and UnivNet Jang et al. 2021, the most dominant and popular GAN-based models. 4) Diffwave Kong et al. 2020b and WaveGrad Chen et al. 2020, recently proposed diffusion probabilistic models which achieve state-of-the-art in speech synthesis. For easy comparison, the results are compiled and presented in Table 2, and we have the following observations:

In terms of audio quality, FastDiff achieved the highest MOS with a gap of 0.240.24 compared to the ground truth audio, and it matched the performance of the autoregressive WaveNet baseline and outperformed the non-autoregressive baselines. For objective evaluation, FastDiff also demonstrated a large improvement in PESQ and STOI. For inference speed, with the efficient noise schedules searched by noise predictor, FastDiff could generate high-quality speech samples within as few as 44 reverse steps, significantly reducing the inference time compared with competing diffusion architectures. To the best of our knowledge, FastDiff makes diffusion models for the first time applicable to interactive, high-quality real-world speech synthesis at a low computational cost. In terms of sample diversity, we can see that FastDiff still witnessed a gap from autoregressive WaveNet, but it achieve a higher variety for generated speeches than non-autoregressive baselines. More detailed evaluation on sample diversity has been attached in Appendix F.

3 Ablation study

We conducted ablation studies to demonstrate the effectiveness of several designs in FastDiff, including the time-aware location variable convolution and noise predictor in neural vocoding. The results of both subjective and objective evaluations have been presented in Table 4, and we have the following observations: 1) Replacing time-aware location-variable convolution by traditional convolutional operations causes a distinct degradation in sampling speed and perceptual quality. 2) Using grid search instead of the noise predictor to search schedules had witnessed the decreased audio quality, demonstrating that the noise schedule prediction process provides more efficient reverse sampling without sacrificing quality.

Further, we compare two variants of FastDiff to test the modality differences of diffusion condition (i.e., continuous noise-level or discrete time-step). Note that the former model does not require the schedule alignment process mentioned in Section 3.3.2. We empirically find that the FastDiff model conditioned on discrete time steps could synthesize samples with higher quality, demonstrating that learning proposed FastDiff with discrete diffusion times could be a better choice. More information on the variant of FastDiff extended to continuous noise schedules has been attached in Appendix E

4 Generalization to unseen speakers

We used 5050 randomly selected utterances of 55 unseen speakers in the VCTK dataset that were excluded from the training set for the MOS test. Table 2 shows the experimental results for the mel-spectrogram inversion of the unseen speakers: In summary, we noticed that FastDiff achieved state-of-the-art in terms of audio quality for out-of-domain generalization, indicating that FastDiff could universally generate high-fidelity audio from entirely new (unseen) speakers outside the train set.

5 End-to-End Text-to-Speech

To demonstrate the robustness of the proposed model in end-to-end text-to-speech synthesis, we compare FastDiff-TTS with other neural TTS systems, including 1) GT, the ground truth audio; 2) GT (voc.), where we first convert the ground truth audio into mel-spectrograms, and then convert the mel-spectrograms back to audio using FastDiff; 3) PortaSpeech Ren et al. 2021 + FastDiff: vocoder cascaded with mel-spectrogram generation using the most popular non-autoregressive TTS models; 4) FastSpeech 2s Ren et al. 2020: the extension of FastSpeech 2 to fully end-to-end text-to-waveform generation with multi-task learning; 5) WaveGrad 2 Chen et al. 2021: a diffusion probabilistic model to generate waveforms via gradient estimation. The results are shown in Table 4: FastDiff-TTS could surpass competing end-to-end speech synthesis models and match the voice quality of the state-of-the-art cascaded TTS systems, demonstrating that FastDiff-TTS is efficient in simplifying the overall text-to-speech synthesis pipeline.

Conclusion

In this work, we proposed FastDiff, a fast conditional diffusion model for high-quality speech synthesis. FastDiff employed a stack of time-aware location-variable convolutions with diverse receptive field patterns to model long-term time dependencies with adaptive conditions. A noise predictor was further adopted to derive tighter schedules for reducing reverse iterations without distinct quality degradation. The extension model FastDiff-TTS discarded intermediate features (e.g., spectrograms) and simplified the end-to-end text-to-waveform syntheses pipeline. Experimental results demonstrated that our proposed model outperformed the best publicly available models in terms of synthesis quality, even comparable to the human level. Moreover, FastDiff showed a significant improvement in synthesis speed, which required as few as 44 iterations to generate high-quality samples. To the best of our knowledge, FastDiff made diffusion models for the first time applicable to interactive, real-world speech generation with a low computational cost. In addition, FastDiff performed strong robustness and enjoyed high-quality synthesis in out-of-domain generalization to unseen speakers. We will release our code and pre-trained models in the future, and we envisage that our work could serve as a basis for future speech synthesis studies.

References

Appendix A Diffusion Probabilistic models

Similar as previous work Ho et al. 2020; Lam et al. 2022; Song et al. 2020a, we define the data distribution as q(x0)q(\mathbf{x}_{0}). The diffusion process is defined by a fixed Markov chain from data x0x_{0} to the latent variable xTx_{T}:

For a small positive constant βt\beta_{t}, a small Gaussian noise is added from xtx_{t} to the distribution of xt−1x_{t-1} under the function of q(xt∣xt−1)q(x_{t}|x_{t-1}).

The whole process gradually converts data x0x_{0} to whitened latents xTx_{T} according to the fixed noise schedule β1,⋯ ,βT\beta_{1},\cdots,\beta_{T}.

Efficient training is optimizing a random term of tt with stochastic gradient descent:

Unlike the diffusion process, reverse process is to recover samples from Gaussian noises. The reverse process is a Markov chain from xTx_{T} to x0x_{0} parameterized by shared θ\theta:

where each iteration eliminate the Gaussian noise added in the diffusion process:

Recently, Bilateral denoising diffusion models (BDDMs) Lam et al. 2022 demonstrates its tighter evidence lower bound (ELBO) for noise schedule prediction. Given a leaned diffusion network θ\theta, a scheduling network ϕ\phi could be applied in reducing the gap between the proposed surrogate objective. To be more specific, instead of using the fixed one in diffusion process, a much more efficient N-step noise schedule (i.e., β^\hat{\beta}) could be derived by the well-leaned noise scheduling network ϕ\phi. The noise schedule could be applied in reverse process, making it possible to explicitly trade-off between inference computation and output quality in one model.

For learning the noise schedule predictor ϕ\phi, we apply the loss function as a KL divergence term between the forward and the reverse distribution:

where Ct=14log⁡1−αt2βt+D2(βt1−αt2−1)C_{t}=\frac{1}{4}\log\frac{1-\alpha_{t}^{2}}{\beta_{t}}+\frac{D}{2}(\frac{\beta_{t}}{1-\alpha_{t}^{2}}-1) is a constant that can be ignored during training.

Appendix B Model Architectures

As illustrated in Table 5, we list the hyper-parameters of FastDiff. We further visualize the detailed architectures of the noise predictor and DBlock in the refinement model in Figure 3.

B.2 FastDiff-TTS

In this section, we list the model hyper-parameters of FastDiff-TTS in Table 5.

Appendix C Training, Noise scheduling and Inference details

We list the diffusion hyper-parameters of FastDiff and FastDiff/FastDiff-TTS in Table 7.

C.2 Noise Scheduling

Our noise scheduling algorithm mainly follows the bilateral denoising diffusion models Lam et al. 2022:

C.3 Schedule Alignment

Here we search and interpolate αs\alpha_{s} between two training noise constants ltl_{t} and lt+1l_{t+1}, enforcing αs\alpha_{s} to get closed to ltl_{t}. In the end, we gain the well-mapped diffusion step tmt_{m}:

Firstly we compute the corresponding constants respective to diffusion and reverse process:

Here we search and interpolate αs\alpha_{s} between two training noise constants ltl_{t} and lt+1l_{t+1}, enforcing αs\alpha_{s} to get closed to ltl_{t}. In the end, we gain the well-mapped diffusion step tmt_{m}:

Where integer tt represents a single pre-defined diffusion step, and ss presents a single step of noise schedule obtained through the scheduling process. Given these two schedules mentioned above, we could conduct schedule alignment and derive the floating-point tmt_{m} for much more efficient reverse sampling.

Appendix D Evaluation Matrix

Perceptual evaluation of speech quality (PESQ) Rix et al. 2001 and The short-time objective intelligibility (STOI) Taal et al. 2010 assesses the denoising quality for speech enhancement.

D.2 NDB and JSD

Number of Statistically-Different Bins (NDB) and Jensen-Shannon divergence (JSD). They measure diversity by 1) clustering the training data into several clusters, and 2) measuring how well the generated samples fit into those clusters.

D.3 Details in MOS Evaluation

All our Mean Opinion Score (MOS) tests are crowd-sourced and conducted by native speakers. The scoring criteria has been included in Table 8 for completeness. The samples are presented and rated one at a time by the testers, each tester is asked to evaluate the subjective naturalness of a sentence on a 1-5 Likert scale. The screenshots of instructions for testers are shown in Figure 4. We paid 8toparticipantshourlyandtotallyspentabout8 to participants hourly and totally spent about750 on participant compensation.

Appendix E Extension to Continuous Condition

Our ablation study extends FastDiff to be conditioned on continuous noise levels and compares it to the basic model with the discrete condition. To be more specific, the FastDiff model conditioned on continuous noise levels does not require an additional schedule alignment process, which has a separated training and sampling procedure:

Appendix F Sample Diversity

Previous works Dhariwal and Nichol 2021; Xiao et al. 2021 in the image generation task has demonstrated that diffusion probabilistic model outperforms GAN in sample diversity, while the comparison in the speech domain is relatively overlooked. Similarly, we can intuitively infer that diffusion probabilistic models are good at generating high-fidelity diverse speech samples. To verify our hypothesis, we employed two metrics NDB and JSD to explore the diversity of generated mel-spectrograms. As shown in Table 2, we can see that diffusion probabilistic model achieve a higher JSD and matching NDB score for generated speeches compare to GAN-based model, which is expected for the following reasons:

1) It is well-known that the mode collapse problem Creswell et al. 2018 appears in the dominated GAN-based generative models, which leads to very similar output samples from a single or few modes of the distribution, especially in the strongly conditional generation task. 2) In contrast, diffusion probabilistic model is meant to reduce mode collapse compared to one-shot generation. It breaks the generation process into several conditional denoising diffusion steps in which each step is relatively simple to model. Thus, we expect our model to exhibit better training stability and mode coverage.