DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs

Songxiang Liu, Dan Su, Dong Yu

Introduction

Text-to-speech (TTS) synthesis is a typical multimodal generation task, where there could be various speech outputs (e.g., with different speaker identities, emotions, speaking styles, etc.) for a given text input. Typical modern neural TTS systems consist of three key components: text analysis frontend, acoustic model, and vocoder. The text analysis frontend normalizes input text and transforms it into linguistic representations. The acoustic model then converts the linguistic representation into time-frequency domain acoustic features, such as mel spectrograms. Finally, the vocoder module generates time-domain waveforms from acoustic features. Different types of generative models have been used as acoustic models to model acoustic variation information in this one-to-many mapping problem to improve the expressiveness and fidelity of synthetic speech.

Neural network-based autoregressive (AR) models have been adopted for TTS and have shown the capability of generating highly natural speech by producing acoustic features frame by frame (Wang et al. 2017; Gibiansky et al. 2017; Sotelo et al. 2017; Li et al. 2019). Nonetheless, AR TTS models often suffer from pronunciation issues, e.g., word skipping and repeating, due to accumulated prediction errors at inference. Moreover, their sequential generative process limits the synthesis speed. To address these limitations, various non-AR TTS models have been proposed. They either leverage an external text-to-acoustic alignment module (Ren et al. 2019; Ren et al. 2021a; Peng et al. 2020; Elias et al. 2021) or jointly train one within the TTS model (Zeng et al. 2020; Miao et al. 2021; Badlani et al. 2021). Other generative models have also been studied for TTS, such as Flow-based models(Kim et al. 2020; Miao et al. 2020), variational auto-encoder (VAE)-based models(Lee et al. 2021; Liu et al. 2021b), and generative adversarial network (GAN)-based models (Donahue et al. 2021; Yang et al. 2021). TTS models combining different generative modeling techniques are also investigated, such as Flow with VAE(Ren et al. 2021b), Flow with VAE and GAN (Kim et al. 2021).

Another class of generative models called denoising diffusion probabilistic models (DDPMs), or abbreviated as diffusion models, has shown an impressive capability to model complex data distributions. Diffusion models have obtained state-of-the-art results in several important domains, including image synthesis(Ho et al. 2020), audio synthesis(Kong et al. 2021; Chen et al. 2021; Popov et al. 2021), graphs(Niu et al. 2020), and symbolic music generation (Mittal et al. 2021). Typical diffusion models comprise a parameter-free TT-step Markov chain called the diffusion process, which gradually adds small random noise into the data, and a parameterized TT-step Markov chain called the denoising process (also known as the reverse process), which removes the added noise as a denoising function. In spite of wide and successful applications of diffusion models in different tasks, sampling from them is not efficient and often requires hundreds or even thousands of denoising steps, making them unsuited for real-time applications. In (Xiao et al. 2021), this slow sampling issue of diffusion models is attributed to the fact that they commonly assume that the denoising distribution can be approximated by Gaussian distributions. This assumption imposes a constraint on typical diffusion models that the denoising step size is sufficiently small and the number of diffusion steps large enough. To use larger denoising step sizes, a conditional GAN is adopted as a non-Gaussian multimodal function to model the denoising distribution (Xiao et al. 2021), leading to much more efficient sampling and, at the same time, competitive sample diversity and fidelity.

This paper introduces DiffGAN-TTS, which is a novel non-AR TTS model based on diffusion models and achieves high-fidelity and efficient TTS. DiffGAN-TTS leverages the powerful modeling capability of diffusion models to address the challenging one-to-many text-to-spectrogram mapping problem. Partially inspired by the denoising diffusion GAN model (Xiao et al. 2021), we model the denoising distribution with an expressive acoustic generator, which is adversarially trained to match the true denoising distribution. DiffGAN-TTS allows large denoising steps at inference, which greatly reduces the number of denoising steps and accelerates sampling. We introduce an active shallow diffusion mechanism into DiffGAN-TTS to further accelerate its sampling process. A two-stage training scheme is designed, where a basic acoustic model trained in stage 1 provides strong prior information for a denoising model trained in stage 2.

Diffusion Models

Diffusion models usually consist of a parameter-free T{T}-step Markov chain named the diffusion process and a parameterized TT-step Markov chain called the reverse process or the denoising process. The diffusion process gradually adds small Gaussian noises into the data until the data structure is totally destroyed at step TT, while the reverse process learns a denoising function to remove the added noise to restore the data structure.

where q(xt∣xt−1):=N(xt;1−βtxt−1,βtI).q(\mathbf{x}_{t}|\mathbf{x}_{t-1}):=\mathcal{N}(\mathbf{x}_{t};\sqrt{1-\beta_{t}}\mathbf{x}_{t-1},\beta_{t}\mathbf{I}). The reverse or denoising process parameterzied with θ\theta is defined by:

The denoising distribution pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) is often modeled as a conditional Gaussian distribution as pθ(xt−1∣xt):=N(xt−1;μθ(xt),σt2I)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}):=\mathcal{N}(\mathbf{x}_{t-1};\mathbf{\mu}_{\theta}(\mathbf{x}_{t}),\sigma_{t}^{2}\mathbf{I}), where μθ(xt,t)\mu_{\theta}(\mathbf{x}_{t},t) and σt2I\sigma_{t}^{2}\mathbf{I} are the mean and variance for the denoising model. Given the parameterized reverse process with the well-learned parameter θ\theta, the sampling process (i.e., the generative process) is to first sample a Gaussian noise xT∼N(0,I)\mathbf{x}_{T}\sim\mathcal{N}(\mathbf{0},\mathbf{I}), and then iteratively sample xt−1∼pθ(xt−1∣xt)\mathbf{x}_{t-1}\sim p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) for t=T,T−1,⋯ ,1t=T,T-1,\cdots,1 along the reverse process, according to the so-called Langevin dynamics. The x0\mathbf{x}_{0} is the generated data.

The likelihood pθ(x0)=∫pθ(x0:T)dx1:Tp_{\theta}(\mathbf{x}_{0})=\int p_{\theta}(\mathbf{x}_{0:T})\text{d}\mathbf{x}_{1:T} is intractable. Hence, the goal of training is to maximize its evidence lower bound (ELBO≤log⁡pθ(x0)\leq\log p_{\theta}(\mathbf{x}_{0})), which can be optimized to match the true denoising distribution q(xt−1∣xt)q(\mathbf{x}_{t-1}|\mathbf{x}_{t}) with the parameterized denoising model pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) with:

where DKLD_{KL} denotes the Kullback-Leibler (KL) divergence and cc contains constant terms that does not dependent on θ\theta. The KL divergence terms in Eq. 3 are generally intractable due to the unknown true denoising distritbuion q(xt−1∣xt)q(\mathbf{x}_{t-1}|\mathbf{x}_{t}). Therefore, in (Ho et al. 2020) the step size βt\beta_{t} in the variance schedule is assumed to be small and the number of denoising steps TT to be large enough such that both the diffusion and denoising processes have the same functional form (i.e., conditional Gaussian). Based on this, (Ho et al. 2020) shows a certain parameterization, transforming the ELBO optimization problem into a simple regression problem.

DiffGAN-TTS

Although DDPMs have demonstrated their capability in modeling complex data distributions, their slow inference speed prevents them from being used in real-time applications. This is due to the two key commonly-made assumptions in DDPMs (Xiao et al. 2021), as alluded to in Section 1: First, the denoising distribution pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) is modeled with a Gaussian distribution. Second, the number of denoising steps TT is often assumed to be large enough such that βt\beta_{t} is small. When the denoising step gets larger and the data distribution is non-Gaussian, the true denoising distribution becomes more complex and multimodal, and in such a case, adopting a parameterized Gaussian transformation to approximate the denoising distribution is insufficient for high-quality generation. Conditional GANs have been adopted to model the multimodal denoising distribution in image generation tasks (Xiao et al. 2021). In this section, we show in great detail how we apply this idea to efficient and high-fidelity multi-speaker TTS.

In this work, we focus on the more challenging multi-speaker TTS tasks, which is much harder than single-speaker TTS since acoustic variations are enlarged by data from different speakers. As illustrated in Figure 2(a)(a), DiffGAN-TTS takes phoneme sequence (denoted as y\mathbf{y}) obtained from a text analysis tool as input to generate intermediate mel-spectrogram features x0\mathbf{x}_{0} with a multi-speaker acoustic generator and then uses a HiFi-GAN-based neural vocoder (Kong et al. 2020) to produce time-domain waveforms. DDPMs are introduced into the acoustic generator to solve the ill-posed multimodal phoneme-to-mel-spectrogram mapping problem.

The training process of DiffGAN-TTS is illustrated in Figure 1. Our goal is to reduce the number of denoising steps TT (e.g., T≤4T\leq 4) of DiffGAN-TTS such that its inference process is efficient, adequate for real-time speech processing applications without degrading the quality of the generated speech. We focus on discrete-time diffusion models, where denoising steps are large, and use a conditional GAN to model the denoising distribution. DiffTTS-GAN trains a conditional GAN-based acoustic generator pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) to approximate the true denoising distribution q(xt−1∣xt)q(\mathbf{x}_{t-1}|\mathbf{x}_{t}) with an adversarial loss that minimizes a divergence DadvD_{\text{adv}} per denoising step:

where we adopt the least-squares GAN (LS-GAN) training formulation (Mao et al. 2017) to minimize DadvD_{\text{adv}} because of its various successful practices in audio generation domain (Kumar et al. 2019; Kong et al. 2020; Yang et al. 2021; Kim et al. 2021).

Let us denote the speaker-ID as ss. The discriminator is designed to be diffusion-step-dependent and speaker-aware to aid the generator to achieve high-fidelity multi-speaker speech generation, as illustrated in Figure 2(cc). The discriminator, denoted as Dϕ(xt−1,xt,t,s)D_{\phi}(\mathbf{x}_{t-1},\mathbf{x}_{t},t,s) with learnable parameters ϕ\phi, is modeled as joint conditional and unconditional (JCU)(Yang et al. 2020; Yang et al. 2021). It not only outputs unconditional logits, but also conditional logits, where diffusion step embedding and speaker embedding are regarded as conditions.

We follow the same scheme in (Xiao et al. 2021) to parameterize the denoising function as an implicit denoising model. Specifically, instead of directly modeling pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) by predicting xt−1\mathbf{x}_{t-1} from xt\mathbf{x}_{t}, the denoising function is modeled as pθ(xt−1∣xt):=q(xt−1∣xt,x0=fθ(xt,t))p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}):=q(\mathbf{x}_{t-1}|\mathbf{x}_{t},\mathbf{x}_{0}=f_{\theta}(\mathbf{x}_{t},t)), where x0\mathbf{x}_{0} is predicted from diffused sample xt\mathbf{x}_{t} using a diffusion function fθ(xt,t)f_{\theta}(\mathbf{x}_{t},t) parameterized with θ\theta. During training, xt−1′\mathbf{x}^{\prime}_{t-1} is sampled using the posterior distribution q(xt−1′∣x0′,xt)q(\mathbf{x}^{\prime}_{t-1}|\mathbf{x}^{\prime}_{0},\mathbf{x}_{t}), where x0′\mathbf{x}^{\prime}_{0} is a predicted version of x0\mathbf{x}_{0}. The predicted tuple (xt−1′,xt)(\mathbf{x}^{\prime}_{t-1},\mathbf{x}_{t}) is then fed into the JCU discriminator to compute the divergence DadvD_{\text{adv}} to the corresponding bonafide counterpart (xt−1,xt)(\mathbf{x}_{t-1},\mathbf{x}_{t}). Different from (Xiao et al. 2021), we do not use the latent variable z∼N(0,I)z\sim\mathcal{N}(\mathbf{0},\mathbf{I}) as an input to the acoustic generator, since the diffusion decoder takes variance-adapted text encodings (happened in variance adaptor of the acoustic generator) and speaker-ID as auxiliary input. To summarize, the implicit distribution fθ(xt,t)f_{\theta}(\mathbf{x}_{t},t) is modeled with the acoustic generator, denoted as Gθ(xt,y,t,s)G_{\theta}(\mathbf{x}_{t},\mathbf{y},t,s), which predicts x0\mathbf{x}_{0} from xt\mathbf{x}_{t} conditioned on phoneme input y\mathbf{y}, diffusion step index tt and speaker ID ss.

2 Training loss

The discriminator is trained to minimize the loss

To train the acoustic generator, we also use the feature matching loss Lfm\mathcal{L}_{fm}, which learns similarity metric to discriminate real and fake data in the feature space (Larsen et al. 2016). Lfm\mathcal{L}_{fm} is computed by summing l1l_{1} distances between every discriminator feature maps of real and generated samples:

where NN is the total number of hidden layers in the discriminator. Acoustic reconstruction loss is also used as additional loss to train the acoustic generator following FastSpeech2 (Ren et al. 2021a), as:

where d\mathbf{d}, p\mathbf{p} and e\mathbf{e} are target duration, pitch and energy, respectively, and d^\mathbf{\hat{d}}, p^\mathbf{\hat{p}} and e^\mathbf{\hat{e}} are their corresponding predicted values. λd\lambda_{d}, λp\lambda_{p} and λe\lambda_{e} are loss weights, which are all set to be 0.1. Lmel\mathcal{L}_{mel} uses MAE loss, while Lduration\mathcal{L}_{duration}, Lpitch\mathcal{L}_{pitch} and Lenergy\mathcal{L}_{energy} use MSE loss. In total, the acoustic generator is trained by minimizing:

and λfm\lambda_{fm} is a dynamically scaled scalar computed as λfm=Lrecon/Lfm\lambda_{fm}=\mathcal{L}_{recon}/\mathcal{L}_{fm} following (Yang et al. 2021). Detailed training procedure as well as inference procedure is presented in Appendix B.

3 Active shallow diffusion mechanism

In the TTS literature, many acoustic models are trained with a simple loss, such as mean squared error (MSE) loss or mean absolute error (MAE) loss. Acoustic features generated from these acoustic models often suffer from over-smoothing issues due to incorrect uni-modal distribution assumptions on data, leading to non-desired synthesis performance. Nonetheless, the blurry acoustic predictions are not without use. It has been shown that output from acoustic models trained with MSE or MAE loss could provide strong prior knowledge of acoustic features (e.g., coarse harmonic structure), which could be further leveraged by a DDPM to generate refined features, leading to improved synthesis performance(Liu et al. 2021a).

To further accelerate inference of DiffGAN-TTS, we introduce an active shallow diffusion mechanism. As illustrated in Figure 3, a two-stage training scheme is designed. At training stage 1, a basic acoustic model, denoted as Gψbase(y,s)G_{\psi}^{\text{base}}(\mathbf{y},s) parameterized with ψ\psi, is trained by:

where Div(⋅,⋅)\text{Div}(\cdot,\cdot) is a distance function to measure the divergence between the predicted and ground-truth, and qdifft(⋅)q_{\text{diff}}^{t}(\cdot) is the diffusion sampling function at step tt, e.g., xt=qdifft(x0)\mathbf{x}_{t}=q_{\text{diff}}^{t}(\mathbf{x}_{0}). It’s worthy to note that qdiff0(⋅)q_{\text{diff}}^{0}(\cdot) is an identity function. The training objective forces the base acoustic model to actively learn to make the diffused samples from ground-truth acoustic features and those from the predicted indistinguishable. At training stage 2, pre-trained weights of the basic acoustic model are copied to initialize the corresponding weights of the acoustic generator of DiffGAN-TTS and then freeze, as illustrated in Figure 3. The base acoustic model generates coarse mel spectrogram x^0\hat{\mathbf{x}}_{0}, which is taken as conditioning by the diffusion decoder. The divergence Dadv(q(xt−1∣xt)∣∣pθ(xt−1∣xt))D_{\text{adv}}(q(\mathbf{x}_{t-1}|\mathbf{x}_{t})||p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x_{t}})) in Eq. 10 is approximated by Dadv(q(xt−1∣xt)∣∣pθ(xt−1∣xt,x^0))D_{\text{adv}}(q(\mathbf{x}_{t-1}|\mathbf{x}_{t})||p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t},\hat{\mathbf{x}}_{0})). Empirically, the diffusion decoder can be regarded as a post filter that conducts “super-resolution” on the coarse prediction produced by the basic acoustic model. We focus on reducing the number of denosing steps to one. During inference, the basic acoustic model first generates a coarse mel spectrogram x^0\hat{\mathbf{x}}_{0}, on which a diffused sample at diffusion step 1 is computed (i.e., x^1\hat{\mathbf{x}}_{1}). Then DiffGAN-TTS takes x^1\hat{\mathbf{x}}_{1} as prior and runs one denoising step to get the final output. We term this variant as DiffGAN-TTS (two-stage), whose detailed training and inference procedures are presented in Appendix B.

4 Model architecture

The transformer encoder in the acoustic generator of DiffGAN-TTS uses the same architecture as that in FastSpeech 2, which consists of 4 feed-forward transformer (FFT) blocks. The hidden size, number of attention heads, kernel size and filter size of the one-dimensional convolution in the FFT block are set as 256, 2, 9 and 1024, respectively. The variance adaptor has the same network structure and hyper-parameters as that in FastSpeech2, which consists of a duration predictor, a pitch predictor and an energy predictor. Differently, the pitch predictor and the energy predictor output phoneme-level fundamental frequency (F0F_{0}) contour and energy contour, respectively, whose labels are obtained by averaging frame-level F0F_{0} and energy values according to phoneme-audio alignment information obtained from a hidden-Markov-model (HMM)-based forced aligner.

The diffusion decoder in DiffGAN-TTS uses a non-causal WaveNet architecture (Oord et al. 2016) with a slight modification. We use a dilation rate of 1 because we are working with mel spectrograms rather than raw waveforms. With a dilation rate of 1, the diffusion decoder’s receptive field is large enough. The diffusion decoder first performs an one-dimensional (1D) convolutional operation with kernel-size 1 (Conv1x1) on the noisy mel spectrogram xt\mathbf{x}_{t} before applying the ReLU activation to the output. Diffusion step tt is encoded using the same sinusoidal positional encoding as in (Vaswani et al. 2017). The mel spectrogram feature maps are added with the diffusion step embedding, which is then fed into 20 WaveNet residual blocks with a hidden dimension of 256. The transformer encoder’s output is imported into each residual block via separate Conv1x1 layers, whose output is added to the hidden feature maps. Then the gated mechanism introduced in (Oord et al. 2016) is used to further process the feature maps. We add the skip connections from all WaveNet blocks and then process them with two Conv1x1 layers interleaved with a ReLU activation to get the diffusion decoder output. All WaveNet residual blocks have their speaker-IDs transformed into embedding vectors. Appendix A contains additional information.

The JCU discriminator, as depicted in Figure 2(cc), adopts purely convolution networks. The Conv1D block consists of 3 one-dimensional convolutional layers with LeakyReLU (slope=0.20.2) as the activation function. The diffusion step embedding layer is the same as that in the diffusion decoder introduced above. The unconditional block and the conditional block have the same network structure, consisting of two 1D convolutional layers. The channels of convolution are 64, 128, 512, 128, and 1. The kernel sizes are 3, 5, 5, 5, and 3 and the strides are 1, 2, 2, 1, and 1.

The transformer encoder and variance adaptor in the basic acoustic model introduced in Section 3.3 are the same as those in DiffGAN-TTS acoustic generator. The mel decoder is composed of 4 FFT blocks.

Experiments

We conduct experiments on an internal gender-balanced corpus comprising transcribed speech data from 228 Mandarin Chinese speakers. In total, the corpus has 200 hours of speech data. We randomly split 1024 utterances for validation and another 1024 utterances for testing. Speech samples are all sampled at 24,000 Hz with 16-bit quantization. Mel spectrograms with 80 frequency bins are computed through a short-time Fourier transform (STFT) using a 1024-point window size and a 10 ms frame-shift. We use the PyWorld toolkit https://github.com/JeremyCCHsu/Python-Wrapper-for-World-Vocoder to compute F0F_{0} values from speech signals. Energy features are computed by taking the l2l_{2}-norm of frequency bins in STFT magnitudes.

2 Training

We train DiffGAN-TTS models with TT=1, 2, and 4 using the Adam optimizer (Kingma & Ba 2015), with β1=0.5\beta_{1}=0.5 and β2=0.9\beta_{2}=0.9, for both the generator and the discriminator. The detailed computation of the variance schedule is presented in Appendix B. We use an exponential learning rate decay with rate 0.999 for training both the generator and discriminator. Initial learning rate for the generator is 10−410^{-4}, while that for the discriminator is 2×10−42\times 10^{-4}. Models are trained using one NVIDIA V100 GPU. The batch size is set to 64, and models are trained for at least 300k steps until losses converge. In the two-stage training scheme, the basic acoustic model is trained 200k steps with the Adam optimizer with β1=0.9\beta_{1}=0.9 and β2=0.98\beta_{2}=0.98 and follows the same learning rate schedule in (Vaswani et al. 2017). As the distance function, we employ the simple MAE loss.

3 Experimental Setup for Comparison

We compare the proposed DiffGAN-TTS model with three strong counterpart models. The first counterpart is the representative non-AR TTS model FastSpeech 2 (Ren et al. 2021a). The second model is the GANSpeech model introduced in (Yang et al. 2021). The third model is the DiffSpeech model presented in (Liu et al. 2021a). We use 60 denoising steps for the best performance on our data.

All compared models use the same acoustic feature setting and the same HiFi-GAN vocoder to ensure fairness. The official implementation https://github.com/jik876/hifi-gan (the “config_v1.json” configuration) is used with slight changes. Since we use 24 kHz audio samples and the hop-size for computing the mel spectrogram is 240, we factorize the upsample rate as 240=8×5×3×2240=8\times 5\times 3\times 2. Moreover, we use the temporal nearest interpolation layer followed by a 1D convolutional layer as the upsampling operation to avoid possible checkerboard artifacts caused by the ”ConvTranspose1d” upsampling layer (Odena et al. 2016).

Results

The quality and voice similarity of generated speech are measured. We conducted an objective evaluation using the structural similarity index measure (SSIM) (Wang et al. 2004), mel-cepstral distortion (MCD)(Kubichek 1993), F0F_{0} root mean squared error (RMSE), and voice Cosine similarity. The computation of MCD and F0F_{0} adopts dynamic time warping (DTW) (dtw 2007) to align the generated speech and the corresponding ground-truth recording. We use the first 24 coefficients when computing mel cepstrums. For computing voice Cosine similarity, we use a pre-trained speaker classifier to extract embedding vectors (the so-called d-vectors) of a generated sample and its ground-truth counterpart, and then compute the Cosine distance between the two vectors. The results are shown in Table 1. We have the following observations: 1) The four DiffGAN-TTS models achieve the best SSIM, MCD, F0F_{0} RMSE and voice Cosine similarity. Specifically, the DiffGAN-TTS (TT=4) model outperforms other compared models in terms of SSIM, MCD, and F0F_{0} RMSE. This indicates that adopting adversarial training in diffusion models well addresses the one-to-many mapping problem in TTS and is able to achieve high-quality speech synthesis, even when modeling multiple speakers within one model. 2) The best voice Cosine similarity is achieved by the proposed DiffGAN-TTS (TT=2) model, and meanwhile, DiffGAN-TTS (TT=1) has the second-highest voice Cosine similarity score. Although the voice Cosine similarity of DiffGAN-TTS (TT=4) is not the best, it reaches a value of up to 0.806, which is still better than previous TTS models (e.g., FastSpeech 2, GANSpeech and DiffSpeech). Comparing the voice Cosine similarity of DiffGAN-TTS (TT=1, 2, 4 and two-stage) with that of DiffSpeech, which takes up to 60 diffusion steps at inference, we conjecture that taking more diffusion steps should degrade voice similarity. 3) The DiffGAN-TTS (two-stage) model, which uses the active shallow diffusion mechanism, achieves overall good performance in terms of the four objective metrics. It is noteworthy that DiffGAN-TTS (two-stage) obtains better SSIM and F0F_{0} RMSE results than DiffGAN-TTS (TT=1 and TT=2) and an on-par MCD score, indicating the effectiveness of the proposed active shallow diffusion mechanism.

2 Subjective Evaluation

We conduct crowd-sourced mean opinion score (MOS) tests to evaluate the quality of the generated speech perceptually. We keep the text content consistent among different models to exclude other interference factors, only examining the audio quality. Each audio sample is rated by at least 20 testers, who are asked to estimate the quality of synthesized speech on a nine-point Likert scale, with the lowest and highest scores being 1 (‘‘Bad”) and 5 (‘‘Excellent”), and with an increment step of 0.5. A subset of the audio samples is available online https://anonym-demo.github.io/diffgan-tts/. The evaluation results are shown in the last column of Table 1. We observe that the DiffSpeech model (TT=60) outperforms other models in spite of its sub-optimal performance in objective measurement. This indicates that diffusion models may take a trade-off between model capacity and inference speed and be able to obtain higher quality with a larger number of denoising steps. Overall, the DiffGAN-TTS family achieves good MOS results, with DiffGAN-TTS (TT=4) obtaining the second-highest MOS (i.e., 4.22). We can see that DiffGAN-TTS (two-stage) achieves a MOS of 4.17, which is significantly better than DiffGAN-TTS (TT=1 and 2). This verifies the two-stage training scheme and the active shallow diffusion mechanism introduced in Section 3.3. A visualization of taking one diffusion step on a predicted mel spectrogram by DiffGAN-TTS (two-stage) at training stage 1 (denoted as x^1\hat{\mathbf{x}}_{1}) and its corresponding ground truth (denoted as x1\mathbf{x}_{1}) is shown in Figure 4. We see that x^1\hat{\mathbf{x}}_{1} has a similar harmonic structure to x1\mathbf{x}_{1}, which inspires us to use the former as a strong prior for the latter.

3 Synthesis Speed

We assess the synthesis speed of mel spectrograms in terms of the Real-Time factor (RTF), indicating how many seconds it takes to generate one second of audio, on an NVIDIA T4 GPU and the number of model parameters. The efficiency information for all models is presented in Table 1. We also evaluate the scaling performance (inference time v.s. text length), as illustrated in Figure 5, where DiffSpeech is not depicted since its inference is too slow, making the plots of other models indistinguishable. It shows that DiffGAN-TTS (TT=1) has similar scaling performance to FastSpeech 2 and other models have satisfactory scaling performance.

4 Ablation studies

We conducted ablation studies in DiffGAN-TTS (TT=4) to demonstrate the effectiveness of the use of mel loss Lmel\mathcal{L}_{mel} and feature matching loss Lfm\mathcal{L}_{fm} introduced in Section 3.2, as well as the exclusion of modeling a latent variable z\mathbf{z} in the diffusion decoder. Details of how to use the latent variable are presented in Appendix A. Objective metrics are computed for these ablations. The results are shown in Table 2. The model without using Lmel\mathcal{L}_{mel} and Lfm\mathcal{L}_{fm} does not train at all, which demonstrates that an adversarial loss alone is not sufficient for training a multi-speaker DiffGAN-TTS model. When comparing the model without Lfm\mathcal{L}_{fm} to that without Lmel\mathcal{L}_{mel}, we observe that mel loss is more important than feature matching loss to successfully train a DiffGAN-TTS model. Adding a latent variable into the model makes all four metrics degrade more or less, indicating that the variance adaptor and speaker conditioning manage to model acoustic variations.

5 Synthesis Variation

Unlike the FastSpeech 2 and GANSpeech models, whose output is uniquely determined by the input text and speaker conditioning at inference, DiffGAN-TTS takes sampling processes at denoising steps and can inject some variations into the generated speech. To demonstrate this, we run a DiffGAN-TTS (TT=4) model 10 times for a particular input text and speaker and compute the F0F_{0} contours of the generated speech samples. We visualize in Figure 6(aa) and observe that DiffGAN-TTS generates speech with diverse pitches. We also generate speech samples for 10 speakers with the same text input, whose F0F_{0} contours are visualized in Figure 6(bb), demonstrating that DiffGAN-TTS expresses very different prosody patterns for each speaker.

Related Work

The Diffusion model is a family of generative models with the capacity to model complex data distribution and has attracted a lot of research attention in recent years. Two streams of efforts have pushed research on diffusion models forward. One research stream is on score matching models (Hyvärinen 2005), where the problem of data density estimation is simplified into a score matching problem, i.e., estimating the gradient of the data distribution probabilistic density. The other stream is research on denoising diffusion probabilistic models (DDPMs) (Sohl-Dickstein et al. 2015), where a diffusion Markov chain is adopted to add noise to the data structure and another Markov chain-based reverse process is used to generate data from noise. Lately, progress has been made in these two streams. The diffusion models have been shown to be able to generate high-quality images (Ho et al. 2020). Afterward, DDPMs achieve high-quality audio generation (Kong et al. 2021; Chen et al. 2021) using the same parameterization introduced in (Ho et al. 2020). On the other hand, score matching models have also been proven to generate high-resolution images by adopting a neural network to estimate the log gradient of data density. Slightly later, (Song et al. 2021) generalizes and improves previous work in score matching models through the lens of stochastic differential equations (SDEs) and shows that DDPMs and score matching models are special cases under this unified framework. In spite of wide and successful applications of diffusion models in different tasks, sampling from them is not efficient and often requires hundreds or even thousands of denoising steps, making them unsuited for real-time applications.

This work is inspired partially by (Xiao et al. 2021), which applies diffusion denoising GANs to accelerate inference in DDPMs for image synthesis, whereas we focus on high-fidelity and efficient text-to-speech (TTS) synthesis. Furthermore, we develop a new model architecture and employ new losses to make the model suitable for TTS tasks that cannot be accomplished naturally. The two-stage version of DiffGAN-TTS is closely related to DiffSpeech (Liu et al. 2021a), which introduces a shallow diffusion scheme by training a boundary prediction to find the starting diffusion step at inference. Our work is much different: 1) We train the model to actively fuse the diffusion processes starting from the ground-truth mel spectrogram and the coarse mel spectrogram produced by a basic acoustic model. 2) Our model is adversarial trained and has a much faster inference speed than DiffSpeech. 3) We concentrate on the more difficult multi-speaker TTS tasks, whereas DiffSpeech concentrates on single-speaker TTS tasks. 4) Our model uses coarse model predictions as conditioning in the diffusion module, whereas DiffSpeech uses text encodings.

Conclusion

In this work, we have presented DiffGAN-TTS, a novel diffusion model-based non-AR TTS model able to achieve high-fidelity and efficient speech synthesis. DiffGAN-TTS adopts an expressive model as a denoising function to approximate the true denoising distribution with adversarial training. Large denoising steps are allowed in DiffGAN-TTS, leading to faster inference. We show with challenging multi-speaker TTS experiments that DiffGAN-TTS is able to generate high-fidelity speech samples with only four denoising steps. To further accelerate its inference process, we propose an active shallow diffusion mechanism and devise a two-stage training scheme to fully leverage prior knowledge from a basic acoustic model. Experiments show that DiffGAN-TTS is able to achieve high TTS performance with only one denoising step. We hope that DiffGAN-TTS will be used in many speech processing applications, especially those requiring real-time speech generation, to enjoy the powerful modeling capacity of diffusion models. We also want to point out that even though the proposed model achieves decent TTS, it is still within the cascade of an acoustic model with a neural vocoder paradigm. Extending the model to support end-to-end text-to-waveform generation could be a possible direction for a shorter TTS pipeline and even better synthesis performance.

References

Appendix A Model details

In this section, we present model details of the diffusion decoder in DiffGAN-TTS and the one in the ablation model taking an extra latent variable z\mathbf{z} as input (introduced in Section 5.4). The diffusion decoder has structure based on the residual block introduced in WaveNet (Oord et al. 2016). Differently, we make the model non-causal. Figure 7(aa) shows the diffusion decoder used in DiffGAN-TTS (TT=1, 2 and 4) and as well as DiffGAN-TTS (two-stage). In the figure, FC represents fully-connected layer, and Swish represents the swish activation function (Ramachandran et al. 2018). Figure 7(bb) shows the diffusion decoder where we inject a latent variable z∼N(0,I)\mathbf{z}\sim\mathcal{N}(\mathbf{0},\mathbf{I}) as input. Inspired by StyleGAN (Karras et al. 2019) and following (Xiao et al. 2021), the latent variable z\mathbf{z} is first transformed by a mapping network, which is simply a fully-connected layer, into an embedding vector, denoted as w\mathbf{w}. In each residual block, we incorporated the latent variable with adaptive layer normalization (AdaLN) layers. Specifically, a fully-connected layer is applied to convert w\mathbf{w} into scale γ\mathbf{\gamma} and shift β\mathbf{\beta} features in each residual block. And then in each AdaLN layer, mean-variance normalization is first conducted on the residual output, and then we use γ\mathbf{\gamma} and β\mathbf{\beta} to modulate the normalized feature hnorm\mathbf{h}_{norm} as hnorm∗γ+β\mathbf{h}_{norm}*\mathbf{\gamma}+\mathbf{\beta}.

Appendix B Training and inference algorithms

In this section, we include a derivation for the Gaussian posterior distribution presented in (Ho et al. 2020) for completeness. Given a data point sampled from a real distribution x0∼q(x)\mathbf{x}_{0}\sim q(\mathbf{x}), consider the diffusion process as a Markov chain as:

Letting αt=1−βt\alpha_{t}=1-\beta_{t} and αˉt=∏i=1tαi\bar{\alpha}_{t}=\prod_{i=1}^{t}\alpha_{i}, a nice property of the diffusion process is that we can sample xt\mathbf{x}_{t} at any diffusion step tt in a closed form (a detailed derivation is given in Appendix A in (Kong et al. 2021)), as:

By Bayes’ rule and Markov chain property, the reverse conditional probability is tractable when conditioned on x0\mathbf{x}_{0}:

where C(xt,x0)C(\mathbf{x}_{t},\mathbf{x}_{0}) is not involving xt−1\mathbf{x}_{t-1}. From the functional form in Eq. 13 and some derivations, we observe that the posterior q(xt−1∣xt,x0)q(\mathbf{x}_{t-1}|\mathbf{x}_{t},\mathbf{x}_{0}) is a Gaussian distribution, which can be written as:

B.2 Variance schedule

We use the discretization of the continuous-time extension of the diffusion process in Eq. 12 with the variance preserving (VP) SDE (Song et al. 2021) to compute the variance schedule β1,⋯ ,βT\beta_{1},\cdots,\beta_{T}. Specifically, for t∈{1,⋯ ,T}t\in\{1,\cdots,T\}, we compute βt\beta_{t} as:

where we set the constants as βmin=0.1\beta_{\text{min}}=0.1 and βmin=40\beta_{\text{min}}=40 in all experiments.

B.3 Training algorithm of DiffGAN-TTS

The training procedures of the DiffGAN-TTS models with TT=1, 2 and 4 are the same across our experiments. We summarize the algorithm in Algorithm 1.

B.4 Inference algorithm of DiffGAN-TTS

The inference procedures of the DiffGAN-TTS models with TT=1, 2 and 4 are the same across our experiments. We summarize the algorithm in Algorithm 2. We visualize the denoising steps of the DiffGAN-TTS (TT=4) model in Figure 8.

B.5 Training algorithm of DiffGAN-TTS with active shallow diffusion

The training procedures of the DiffGAN-TTS model with active shallow diffusion, i.e., DiffGAN-TTS (two-stage), is summarized in Algorithm 3. We visualize the denoising steps of the DiffGAN-TTS (two-stage) model in Figure 9.

B.6 Inference algorithm of DiffGAN-TTS with active shallow diffusion

The inference procedures of the DiffGAN-TTS model with active shallow diffusion, i.e., DiffGAN-TTS (two-stage), is summarized in Algorithm 4.