DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training

Hyung-Seok Oh, Sang-Hoon Lee, Seong-Whan Lee

I Introduction

Recent advancements in neural text-to-speech (TTS) models have significantly enhanced the naturalness of synthetic speech. In several studies, prosody modeling has been leveraged to synthesize speech that closely resembles human expression. Prosody, which encompasses various speech properties, such as pitch, energy, and duration, plays a crucial role in the synthesis of expressive speech.

In some studies, reference encoders have been used to extract prosody vectors for prosody modeling. A global style token (GST) is an unsupervised style modeling method that uses learnable tokens to model and control various styles. Meta-StyleSpeech proposes the application of style vectors extracted using a reference encoder through a style-adaptive layer norm. Progressive variational autoencoder TTS presents a method for gradual style adaptation. A zero-shot method for speech synthesis that comprises the use of a normalization architecture, speaker encoder, and feedforward transformer-based architecture was proposed. Despite the intuitive and effective nature of using a reference encoder, these methods cannot reflect the details of prosody in their synthetic speech without ground-truth (GT) prosody information.

Recently, methods for inferring prosody from text in the absence of reference audio have been developed . FastPitch, for instance, synthesizes speech under text and fundamental frequency conditions. FastSpeech 2 aims to generate natural, human-like speech by using extracted prosody features, such as pitch, energy, and duration, through an external tool and introduce a variance adaptor module that predicts these features. Some studies have proposed hierarchical models through the design of prosody features at both coarse and fine-grained levels. However, the separate modeling of prosodic features may yield unnatural results owing to their inherent correlation.

Some studies have predicted a unified prosody vector, thus enhancing the representation of prosody, given the interdependence of prosody features. Text-predicted GST is a method for modeling prosody without reference audio by predicting the weight of the style token from the input text. proposed a method for paragraph-based prosody modeling by introducing a paragraph encoder. Gaussian-mixture-model-based phone-level prosody modelling is a method for sampling reference prosody from Gaussian components. proposed a method for modeling prosody using style, perception, and frame-level reconstruction loss. There are also studies in which prosody is modeled using pre-trained language models. ProsoSpeech models prosody with vector quantization (VQ) using large amounts of data and predicts the index of the codebook using an autoregressive (AR) prosody predictor. However, when predicting prosody vectors, an AR prosody predictor encounters challenges related to long-term dependencies.

To address these issues, we propose DiffProsody, a novel approach that generates expressive speech by employing a diffusion-based latent prosody generator (DLPG) and prosody conditional adversarial training.

The primary contributions of this work are as follows:

We propose a diffusion-based latent prosody modeling method that can generate high-quality latent prosody representations, thereby enhancing the expressiveness of synthetic speech. Furthermore, we adopted denoising diffusion generative adversarial networks (DDGANs) to reduce the number of timesteps, resulting in speeds that were 2.48× and 16× faster than those of the AR model and denoising diffusion probabilistic model (DDPM) , respectively.

We propose prosody conditional adversarial training to ensure an accurate reflection of prosody using the TTS module. A significant improvement in smoothness, attributable to vector quantization, was observed in the generated speech.

Objective and subjective evaluations demonstrated that the proposed method outperforms comparative models.

The implementationhttps://github.com/hsoh0306/DiffProsody of proposed method and audio sampleshttps://prml-lab-speech-team.github.io/demo/DiffProsody/ for various datasets, such as VCTKhttps://datashare.ed.ac.uk/handle/10283/2651 and LibriTTS, are available online.

II Related works

Traditional TTS models function autoregressively. This implies that the spectrogram generates one frame at a time, with each frame conditioned on the preceding frames. Despite the high-quality speech that this approach can produce, it has a drawback in terms of speed owing to the sequential nature of the generation process. To address this issue, non-autoregressive TTS (NAR-TTS) models have been proposed as an alternative for parallel generation. These models have the advantage of simultaneously generating the entire spectrogram, thus resulting in a significant acceleration of the speech synthesis. FastSpeech and FastSpeech 2 serve as examples of NAR-TTS models that can synthesize speech at a much faster rate than their AR counterparts while maintaining a comparable level of quality. For parallel generation, these models require phoneme-level durations. FastSpeech uses an AR teacher model to obtain durations through knowledge distillation. The phoneme-level input is scaled to the frame-level using a length regulator, and a transformer-based network is used to generate the entire utterance at once. FastSpeech 2 addresses some of the disadvantages of FastSpeech by extracting the phoneme duration from forced alignment as the training target instead of relying on the attention map of the AR teacher model and introducing more variation information in speech as conditional inputs. In contrast to these models that use an external aligner, are parallel TTS models that use an internal aligner to model duration. These parallel-generation models exhibit faster and more robust generation than the AR models. In this study, we adopted a transformer-based NAR model with a simple structure to focus on prosody modeling.

II-B Generative adversarial networks

Generative adversarial networks (GANs) are generative models in which generative and discriminative networks compete against each other. The objective of the generative network is to create samples that closely resemble the true data distribution, whereas the discriminative network strives to differentiate between the data sampled from the true and generated distributions.

These two networks play a minimax game with the following value function V(D,G)V(D,G):

where GG represents the generative network, and DD represents the discriminative network. The training process involves minimizing D(1−G(z))D(1-G(z)) to recognize DD as real for GG, and maximizing log(D(x))log(D(x)) for DD to learn the likelihood of real data. The conditional generation model was obtained by introducing condition cc into both GG and DD. In this study, we incorporated adversarial training conditions into prosody features to generate expressive speech.

II-C Denoising diffusion models

The denoising diffusion model is a generative model that gradually collapses data into noise and generates data from noise. The processes of collapsing and denoising data are called the forward and reverse processes, respectively. The forward process gradually collapses data x0\mathbf{x}_{0} into noise over the TT-step, with a predefined variance schedule βt\beta_{t}.

where q(xt∣xt−1):=N(xt;1−βtxt−1,βtI)q(\mathbf{x}_{t}|\mathbf{x}_{t-1}):=\mathcal{N}(\mathbf{x}_{t};\sqrt{1-\beta_{t}}\mathbf{x}_{t-1},\beta_{t}I). The reverse process is defined as follows:

The reverse process is driven by the θ\theta denoising model parameterized by a. The denoising model was optimized for a variational bound on the negative log-likelihood.

In the DDPM, the denoising distribution pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) is assumed to comprise a Gaussian distribution. Moreover, it has been demonstrated that the diffusion model can generate a diverse range of complex distributions, provided that a sufficient number of iterations are performed. Proceeding with only a small step at a time is possible by setting the denoising distribution to a Gaussian distribution, which implies that a considerable number of timesteps are required. DDGAN was proposed by modeling the denoising distribution with a non-Gaussian multimodal distribution to reduce the sampling step. It predicts x0\mathbf{x}_{0} with the generator GθG_{\theta}, which models the implicit distribution, in contrast to the denoising network of the DDPM, which predicts noise. In the DDGAN, the conditional probability pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) is defined as follows:

where pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) represents the implicit distribution imposed by the generator Gθ(x,z,t)G_{\theta}(\mathbf{x},\mathbf{z},t), which outputs x0\mathbf{x}_{0} given xt\mathbf{x}_{t} and the latent variable z∼p(z):=N(z;0,I)\mathbf{z}\sim p(\mathbf{z}):=\mathcal{N}(\mathbf{z};0,\mathbf{I}). In the DDGAN, the denoising distribution pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) is modeled as a complex multimodal distribution, in contrast to the unimodal distribution in the DDPM. Sampling x0\mathbf{x}_{0} with small timesteps is possible by leveraging the complicated pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}). We introduced this model to reduce the sampling timesteps while maintaining its ability to generate diffusion models.

II-D Prosody modeling

Although the pronunciation capabilities of TTS models have seen significant advancements, they still fail to replicate the naturalness inherent in human speech. Various studies have proposed methods for prosody modeling to address this limitation. One such method is the reference-based approach, which is a type of expressive speech synthesis that extracts styles from the reference audio. This approach is particularly beneficial for style transfers, but can result in unnatural results when the text and reference audio do not align well. However, there are methods that directly model prosody properties such as pitch, energy, and duration. These methods offer the advantages of explainability and controllability because they directly use and model prosody features. In the context of ProsoSpeech, the authors argue that prosodic features are interdependent and that modeling them separately can result in unnatural outcomes. To address this, they proposed the modeling of prosody as a latent prosody vector (LPV) and introduced an AR prosody predictor to obtain the LPV. In this study, we adopt a similar latent vector approach to model prosody. The use of pre-trained language models was also suggested . These models comprise the use of models that have been pre-trained on large datasets, such as BERT or GPT-3.

III DiffProsody

The proposed method, called DiffProsody, aims to enhance speech synthesis by incorporating a diffusion-based latent prosody generator (DLPG) and prosody conditional adversarial training. The overall structure and process of DiffProsody are presented in Figure 1. In the first stage, we trained a TTS module and a prosody encoder using a text sequence and a reference Mel-spectrogram as inputs. The prosody conditional discriminator evaluates the prosody vector from the prosody encoder and the Mel-spectrogram from the TTS module to provide feedback on their quality. In the second stage, we train a DLPG to sample a prosody vector that corresponds to the input text and speaker. During inference, the TTS module synthesizes speech without relying on a reference Mel-spectrogram. Instead, it uses the output of a DLPG. This facilitates the generation of expressive speech that accurately reflects the desired prosody.

The TTS module is designed to transform text into Mel-spectrograms using speaker and prosody vectors as conditions. The overall structure of the model is presented in Figure 1a. The TTS module comprises a text encoder and a decoder. The text encoder processes the text at both the phoneme and word levels, as illustrated in Figure 1c. The input text, denoted as xtxt\mathbf{x}_{txt}, is converted into a text hidden representation htxt\mathbf{h}_{txt}, by the phoneme encoder EpE_{p} and word encoder EwE_{w}. The EpE_{p} takes the phoneme-level text xph\mathbf{x}_{ph} and the EwE_{w} takes as input the word-level text xwd\mathbf{x}_{wd}. The htxt\mathbf{h}_{txt} is then obtained as the element-wise sum of the outputs of Ep(xph)E_{p}(\mathbf{x}_{ph}) and Ew(xwd)E_{w}(\mathbf{x}_{wd}) expanded to the phoneme-level.

where expandexpand is an operation that expands the word-level features to the phoneme-level. Obtaining the quantized prosody vector zpros\mathbf{z}_{pros} involves using htxt\mathbf{h}_{txt} and speaker hidden representation hspk\mathbf{h}_{spk} as inputs for the prosody module. In addition, hspk\mathbf{h}_{spk} is acquired using a pre-trained speaker encoder. We use Resemblyzerhttps://github.com/resemble-ai/Resemblyzer, an open-source model trained with generalized end-to-end loss (GE2E), to extract hspk\mathbf{h}_{spk}.

During the first stage of training, a prosody encoder is employed, which receives the target Mel-spectrogram. In the inference, zpros′\mathbf{z}_{pros}^{\prime} is obtained by inputting htxt\mathbf{h}_{txt} and hspk\mathbf{h}_{spk} into a DLPG, and this is performed without a reference Mel-spectrogram. Finally, the information related to the text, speaker, and prosody is combined by expanding the latent vectors htxt\mathbf{h}_{txt}, hspk\mathbf{h}_{spk}, and zpros\mathbf{z}_{pros} to the phoneme-level and then performing an element-wise summation.

The phoneme duration is modeled using the duration predictor DPDP. The goal of the DPDP is to predict the phoneme duration at the frame-level based on the input variable htotal\mathbf{h}_{total}.

In addition, there is a length regulator LRLR that expands the input variable to the frame-level using the phoneme duration durdur. The expanded htotal\mathbf{h}_{total} is then transformed to Mel-spectrogram y′\mathbf{y}^{\prime} by DmelD_{mel}.

For TTS modeling, we use two types of losses: the mean square error (MSE) and structural similarity index (SSIM) loss. These losses aid in accurately modeling the TTS. For the duration modeling, we use the MSE loss.

III-B Prosody module

During the inference stage, the prosody vector hpros′\mathbf{h}_{pros}^{\prime} is obtained using the prosody generator trained in the second stage.

where zprosi\mathbf{z}_{pros}^{i} is ii-th element of zpros\mathbf{z}_{pros}. In the first stage, the TTS module is trained jointly with the codebook ZZ and prosody encoder EprosE_{pros}.

where sg[⋅]sg[\cdot] denotes the stop-gradient operation. Moreover, we employ an exponential moving average (EMA) to enhance the learning efficiency by applying it to codebook updates.

III-C Prosody conditional adversarial training

For our prosody adversarial training, we develop prosody conditional discriminators (PCDs) that handle inputs of varying lengths. This design is inspired by multi-length window discriminators. The PCD structure is presented in Figure 1e. The PCD is designed to accept a Mel-spectrogram yy and a quantized prosody vector zpros\mathbf{z}_{pros} as inputs, and its role is to determine whether these input features are original or generated. The PCD comprises two lightweight convolutional neural networks (CNNs) and fully connected layers. One of the CNNs is designed to receive only the Mel-spectrogram, whereas the other is designed to receive a combination of zpros\mathbf{z}_{pros} and yy. To match the corresponding PCD, the length of each Mel-spectrogram and the extended zpros\mathbf{z}_{pros} are randomly clipped. For our objective function, we adopt the least square GAN loss:

where LD\mathcal{L}_{D} denotes the training goal of the discriminators and LG\mathcal{L}_{G} represents the feedback on the TTS module. The final object LTTS\mathcal{L}_{TTS} of the TTS module is as follows:

where λ1\lambda_{1} corresponds to the weight of the adversarial loss.

III-D Diffusion-based latent prosody generator

We propose a new module called the DLPG, which leverages the powerful generative capabilities of diffusion models. In addition, we introduce a DDGAN framework, which enables faster sampling by reducing the number of required timesteps. Figure 2 presents the training process for the DLPG. During training, the DLPG aims to generate target hpros\mathbf{h}_{pros}, which is extracted from the prosody encoder trained in the first stage. The DLPG is trained to produce hpros′\mathbf{h}_{pros}^{\prime} based on the hspk\mathbf{h}_{spk} and htxt\mathbf{h}_{txt}. In the diffusion model, we set x0\mathbf{x}_{0} as the target hpros\mathbf{h}_{pros}. The DLPG generator GθG_{\theta} directly generates x0′\mathbf{x}^{\prime}_{0}.

where tt is timestep of diffusion process. To ensure adversarial training, xt−1′\mathbf{x}^{\prime}_{t-1} is derived from xt\mathbf{x}_{t} and x0′\mathbf{x}^{\prime}_{0} using posterior sampling q(xt−1′∣xt,x0′)q(\mathbf{x}^{\prime}_{t-1}|\mathbf{x}_{t},\mathbf{x}^{\prime}_{0}). Subsequently, a time-dependent discriminator DϕD_{\phi} determines the compatibility of xt−1\mathbf{x}_{t-1} (obtained from the forward processing of x0\mathbf{x}_{0}) and xt−1′\mathbf{x}^{\prime}_{t-1} (generated through posterior sampling of x0′\mathbf{x}_{0}^{\prime}) with respect to tt and xt\mathbf{x}_{t}, conditioned on hspk\mathbf{h}_{spk} and htxt\mathbf{h}_{txt}. The objective function of GθG_{\theta} is then defined as follows:

where LGθadv\mathcal{L}_{G_{\theta}}^{adv} corresponds to the adversarial loss, and LGθrec\mathcal{L}_{G_{\theta}}^{rec} denotes the reconstruction loss of GθG_{\theta}. The total generator loss LGθ\mathcal{L}_{G_{\theta}} is expressed as follows:

where λ2\lambda_{2} is the weight of adversarial loss LGθadv\mathcal{L}_{G_{\theta}adv}. The objective function of the DϕD_{\phi} is as follows:

The DLPG leverages the DDGAN framework to achieve stable and high-quality results in only a few timesteps. This process involves the GθG_{\theta}, which iteratively generates x0′\mathbf{x}^{\prime}_{0} TT times during the inference. We set xT\mathbf{x}_{T} to follow a normal distribution. The hpros′\mathbf{h}_{pros}^{\prime} is obtained as the final x0′\mathbf{x}^{\prime}_{0} of the reverse process. The final prosody vector zpros\mathbf{z}_{pros} is then derived through the vector quantization of hpros′\mathbf{h}_{pros}^{\prime}.

III-E Inference

Here is the step-by-step process of generating a Mel-spectrogram using the trained TTS module and DLPG.

The htxt\mathbf{h}_{txt} is extracted from the text encoder, and hspk\mathbf{h}_{spk} is extracted from the pre-trained speaker encoder.

The DLPG generates a hpros′\mathbf{h}_{pros}^{\prime} with htxt\mathbf{h}_{txt} and hspk\mathbf{h}_{spk} as inputs.

hpros′\mathbf{h}_{pros}^{\prime} is mapped to a codebook, denoted as ZZ, in the VQ layer to obtain the prosody vector, which is denoted as zpros\mathbf{z}_{pros}.

The decoder DmelD_{mel} generates a Mel-spectrogram y′\mathbf{y}^{\prime} using htxt\mathbf{h}_{txt}, hspk\mathbf{h}_{spk} and zpros\mathbf{z}_{pros}. This process involves the expansion htxt\mathbf{h}_{txt}, hspk\mathbf{h}_{spk}, and zpros\mathbf{z}_{pros} to the frame-level. The phoneme-duration is predicted by the duration predictor.

The Mel-spectrogram y′y^{\prime} was converted to a raw waveform using a pre-trained vocoder.

IV Experimental result and discussion

We conducted experiments using the VCTK dataset, a multispeaker English dataset consisting of audio clips of 44,200 sentences recorded by 109 speakers for approximately 400 sentences. A total of 2,180 audio clips were randomly selected with 20 sentences per speaker, which constituted the test set. Furthermore, 545 audio clips randomly selected from five sentences per speaker were used as the verification sets, and the remainder were used as the training sets. We sampled audio at 22,050 Hz and then transformed it to an 80-bin Mel-spectrogram using an STFT with a window length of 1,024 and hop size of 256. The text was converted into phoneme sequences for text input using a grapheme-to-phoneme toolhttps://github.com/Kyubyong/g2p. We extracted the phoneme duration using the Montreal Forced Aligner tool. Furthermore, we used the AdamW optimizer with β1=0.9\beta_{1}=0.9 and β2=0.98\beta_{2}=0.98. The learning rates for the TTS and latent prosody generator training were set as 5×10−45\times 10^{-4} and 2×10−42\times 10^{-4}, respectively. Throughout the training process, the batch size used was 48. The TTS and prosody encoder were updated in 160k steps, and the latent prosody generator was updated in 320k steps. In the experiment, all the audio was synthesized using the official implementation of the HiFi-GANhttps://github.com/jik876/hifi-gan and a pre-trained model. We trained the TTS module and prosody encoder for approximately 16 h and the DLPG for 7 h using a single NVIDIA RTX A6000 GPU.

IV-B Implementation details

The number of layers, hidden size, filter size, and kernel size of the feed-forward transformer blocks of the phoneme encoder, word encoder, and decoder were set as 4, 192, 384, and 5, respectively. The extracted speaker embedding was projected onto 192 dimensions, and the number of dimensions of the prosody vector was 192. The structure of the prosody encoder follows that of ProsoSpeech. For VQ, we set the size of codebook and dimension of code as 128 and 192, respectively. and updated it using EMA with a decay rate of 0.998. The codebook was initialized as the center of the k-means clustering after 20k steps of TTS training. The number of Mel-spectrogram bins used in the prosody encoder NN was set as 20. We set multiple PCDs to receive different input sizes, such as. The PCD comprises two 2D convolution stacks and three fully connected layers. The convolution stacks in the PCD consist of three 2D convolutions: LeakyReLU, BatchNorm, and a linear layer with a multilength window size of . The latent prosody generator consists of 20 residual blocks with a hidden size of 384. The prosody discriminator is configured with four convolution layers and hidden dimensions of 384. We investigated λ1\lambda_{1} and λ2\lambda_{2} between 0.001 and 1.0 and set λ1\lambda_{1} and λ2\lambda_{2} as 0.01 and 0.05, respectively. The number of timesteps used in the training and inference was set as four.

IV-C Comparative studies

We developed a prosody model from previous works, our proposed model, and the model of the ablation study for performing comparative experiments. All the models produced an 80-bin Mel-spectrogram and synthesized speech using the same vocoder.

IV-C2 GT (vocoded)

The generated audio was obtained by converting the GT Mel-spectrogram using HiFi-GAN V1.

IV-C3 FastSpeech 2

FastSpeech 2 is a speech-synthesis model that directly predicts prosody features (pitch and energy). We trained FastSpeech 2 using an open-source implementationhttps://github.com/NATSpeech/NATSpeech. For the purpose of a fair comparison, we used a text encoder at the phoneme and word-levels.

IV-C4 ProsoSpeech

ProsoSpeech is a speech synthesis model that models prosody features as latent prosody vectors and predicts them using an AR predictor. We implemented the model by following the hyperparameters in the study and used the model that provided the best performance.

IV-C5 DiffProsody

The proposed model was implemented using a text encoder and decoder with the same structure as FastSpeech 2 and ProsoSpeech. We used a prosody encoder with the same structure as that of ProsoSpeech. In the first stage, a prosody conditional discriminator was added. In the second stage, a DLPG was used instead of an AR predictor.

IV-C6 DiffProsody (AR)

DiffProsody (AR) is a model in which a DiffProsody trained with an AR prosody predictor to estimate prosody vectors. It has the same structure as the ProsoSpeech prosody predictor.

IV-C7 DiffProsody (DDPM)

DiffProsody (DDPM) is a model in which a DiffProsody trained with a DDPM framework to estimate a prosody vector. It has the same structure as DiffProsody’s diffusion-based latent generator.

IV-C8 DiffProsody (w/o PCD)

DiffProsody (w/o PCD) is a model in which a DiffProsody trained without the assistance of a prosody conditional discriminator.

IV-C9 DiffProsody (w/o VQ)

DiffProsody (w/o VQ) is a model in which a DiffProsody trained without a VQ layer in the prosody encoder.

IV-C10 DiffProsody N

DiffProsody NN is a model in which a DiffProsody trained with the lowest NN Mel-bins in the prosody encoder.

IV-D Subjective metrics

A subjective assessment is conducted to confirm the effectiveness of the proposed method. To measure the level of naturalness, we used the mean opinion score (MOS). For this evaluation, we employed Amazon Mechanical Turk (MTurk), a crowdsourcing service, to gather feedback from 20 native Americans. MOS was assessed using a 5-point scale, and confidence intervals were calculated at a 95% level. For the evaluation, 100 samples were randomly selected from the test set.

IV-E Objective metrics

For realizing an objective evaluation, we calculated the equal error rate (EER) using a pre-trained speaker verification modelhttps://github.com/clovaai/voxceleb_trainer . We used the pre-trained wav2vec 2.0 to the compute character error rate (CER) and word error rate (WER).

For the prosodic evaluation, we computed the average differences in utterance duration (DDUR), pitch error (RMSEf0{}_{f_{0}}) (in cents), periodicity error (RMSEperiod) and F1 score of the voiced/unvoiced classification (F1v/uv). We used torchcrepehttps://github.com/maxrmorrison/torchcrepe to extract the pitch and periodicity features for evaluation. In addition, we measured the Kullback–Leibler (KL) divergence of log f0 and log energy to compare the distributions of the prosody features in the generated audio. Finally, we calculated the real-time factor (RTF) to compare the generation speeds.

We obtained DDURDDUR by calculating the mean absolute error of the difference in the duration of each utterance.

where duridur_{i} denotes the duration of the ii-th GT utterance and duri′dur_{i}^{\prime} denotes the duration of the ii-th generated utterance.

IV-E2 Pitch error

To measure the pitch error RMSEf0{}_{f_{0}}, we aligned the pitches extracted in hertz with dynamic time warping (DTW) and calculated the root mean square (RMSE) in cents, defined as 1200log⁡2(y/y^)1200\log_{2}(y/\hat{y}) for pitch of the GT speech yy and pitch of the generated speech y^\hat{y}. We measured the portion wherein both the GT speech and generated speech were voiced.

IV-E3 Periodicity error

We measured the periodicity error RMSEperiod by root mean square between the periodicities ψ\psi aligned with the DTW.

where ψi\psi_{i} means the ii-th periodicity value.

IV-E4 F1 score of voiced/unvoiced

We obtained voiced/unvoiced flags from the aligned pitches using DTW and calculated the F1 score (F1v/uv) between them. We defined a match between the GT voiced flag uu and generated voiced flag u′u^{\prime} as a true positive (TP), a match between the GT unvoiced flag uvuv and generated voiced flag v′v^{\prime} as a false positive (FP), and a match between the GT voiced flag vv and generated unvoiced flag uv′uv^{\prime} as a false negative (FN).

where nn is the length of the sequence, and [ai=bi][a_{i}=b_{i}] is a function that returns 1 if the ii-th element has the same value, and 0 otherwise. The precision and recall are defined as follows:

We then calculated the F1 score as follows:

IV-E5 KL divergence of log f0 / log energy

We also measured the KL D to analyze the pitch and energy distribution. We first extracted the pitch and energy in log-scale. We then binned the entire range into 100 bins and applied a kernel density estimation to each bin to calculate the KL D for a smoothed distribution. The KL divergence for feature xx is defined as follows:

where NN is the number of bins, xix_{i} is the probability in the ii-th bin of the distribution of xx, KDEKDE is the kernel density estimator function, and KLDKLD is the KL divergence calculation.

IV-F Evaluation results

Table I lists the MOS results and objective evaluations. The results demonstrate that the proposed model DiffProsody surpasses the other models in terms of both subjective and objective metrics. In particular, our model achieved a superior MOS for the subjective metrics, with a p-value of less than 0.05. It also exhibited outstanding performance in terms of the DDUR, EER, RMSEf0{}_{f_{0}}, RMSEperiod, and F1v/uv, which suggests that the speech generated by DiffProsody closely resembles the target prosody. Furthermore, our model displayed a lower WER and CER, thus indicating its capability for synthesizing speech with a more accurate pronunciation.

Figure 3 presents the Mel-spectrogram and pitch contour of speech from each model. The red box in the figure indicates that the DiffProsody model was more similar to the GT. To further evaluate the prosody of the generated speech, we examined pitch and energy distributions.

Figure 5 presents the histogram distribution for log f0 (pitch), and Figure 5 presents the histogram distribution for log energy. In both the figures, the blue bars represent the distribution of the GT features, and the orange bars represent the distribution of the generated features. The results indicate that the proposed model aligns more closely with the GT distribution than the other models. Table II lists the KL divergence values for comparison. ProsoSpeech exhibited a better performance in f0 than FastSpeech 2, but not in the case of energy. However, DiffProsody outperformed the comparison model in terms of f0 and energy.

IV-G Latent prosody generation method

In Table III, DiffProsody (AR) is a model comprising the use of an AR prosody predictor. DiffProsody (DDPM) is a model comprising the use of DDPM (100 timesteps). DiffProsody is a model comprising the use of DDGAN (four timesteps). The AR predictor has the same structure as the prosody predictor in ProsoSpeech, but it does not use context encoders, and the denoising network in the DDPM has the same structure as generator gθg_{\theta} in the DDGAN. We also conducted a 7-point comparative mean opinion score (CMOS) evaluation to compare the latent prosody generation methods and measured the RTF to compare the generation speeds. DDGAN achieved −-0.172 and ++0.015 CMOS compared with AR and DDPM, respectively. The objective metrics showed that the DDPM and DDGAN outperformed the AR. The DDGAN achieved results nearly identical to those of the DDPM for all the metrics. The experimental results showed that the diffusion models performed better than the AR models, as reported in. Furthermore, DDGAN can generate high-quality prosody vectors such as DDPM in only four timesteps. According to the RTF results, the DDGAN produces a 2.7×\times and 16×\times faster prosody than the AR and DDPM, respectively.

Figure 6 presents the traces of the WER, CER, and EER results for the speech generated by the DLPG with the prosody vector generated by each denoising iteration. The red and blue lines represent the results obtained using the DDPM and DDGAN, respectively. We observed that the DDGAN has a larger error rate than the DDPM in the early stages and that the midpoint and final stages of each model have almost the same value. This implies that the DDGAN compresses the timesteps of the DDPM by modeling the denoising process as a multimodal non-Gaussian distribution.

IV-H Prosody conditional adversarial training

We conducted CMOS to assess the effectiveness of PCD. The CMOS values were used to measure the degree of PCD preference in comparison with the reference model. Our evaluation involved the analysis of three different models: DiffProsody using PCD and a diffusion-based model, DiffProsody without PCD (referred to as DiffProsody (w/o PCD)) using a diffusion-based model, and ProsoSpeech using an AR prosody predictor without PCD. The results presented in Table IV clearly indicate that models comprising the use of the PCD are preferred over those trained without the PCD. Furthermore, by comparing the CMOS results of DiffProsody (w/o PCD) and ProsoSpeech, we can infer that the diffusion-based method is preferable to the method employed by ProsoSpeech. To provide a more comprehensive analysis, we also incorporated the objective metrics for DiffProsody (w/o PCD) in Table III. These objective metric results demonstrate that DiffProsody outperformed DiffProsody (w/o PCD) in all aspects. However, DiffProsody (w/o PCD) still received a better objective evaluation than ProsoSpeech. In addition, when comparing DiffProsody (AR) with ProsoSpeech, DiffProsody (AR) consistently achieved higher scores on the majority of the objective evaluations. These findings validate that both PCD and DLPG significantly improve the model performance.

IV-I Prosody encoder evaluation

In this section, the effectiveness of the prosody encoder is evaluated. We focus on two main aspects: the impact of the number of Mel-bins used in the reference Mel-spectrogram and the role of the vector quantization layer. Detailed results of these evaluations are presented in the following subsections.

We conducted an experiment to compare the performance of DiffProsody when trained using various numbers of Mel-bins. The results, including the EER, CER, and WER scores, are presented in Table V. We examined the performance of 10, 20, 30, 40, 60, and 80 (full-band) Mel-bins. The findings indicated that as the number of bins exceeded 20 (baseline), the EER tended to increase, while no significant difference was observed in terms of the CER and WER. It should be noted that the model with 10 bins outperformed the baseline (20 bins) in terms of the WER but yielded higher EER results. Figure 7 illustrates the Mel-spectrogram at the initial iteration (noise) during the diffusion timestep of the DLPG, which was trained using various numbers of Mel-bins. As the number of Mel-bins used for the training was increased, the results of the Mel-spectrogram progressively smoothened and eventually collapsed. This phenomenon occurs because the large amount of information in the reference forces the model to reconstruct the Mel-spectrogram by leveraging the prosody vectors. Through this experiment, we found that, as NN increases, linguistic information becomes increasingly entangled. Consequently, it is reasonable to employ 20 Mel-bins as the input to the prosody encoder for realizing effective prosody modeling.

IV-I2 Vector quantization layer analysis

Figure 8 presents a Mel-spectrogram trace synthesized using the prosody vector generated for each diffusion step of the DLPG. We compared two versions of DiffProsody: one trained without the vector quantization layer (Figure 8a; defined as DiffProsody (w/o VQ)) and the other trained with the vector quantization layer (Figure 8b; defined as DiffProsody). In the case of DiffProsody (w/o VQ), the early steps exhibited a completely distorted Mel-spectrogram, but there was a significant recovery in the middle steps. Conversely, the initial step of DiffProsody exhibits a smoothed but slightly distorted Mel-spectrogram that gradually returns to its original state over the subsequent steps. This phenomenon is also related to prosody disentangling, and we confirmed that the prosody disentangling failed in DiffProsody (w/o VQ). This experiment demonstrated that vector quantization plays a crucial role in effective prosody disentangling. The last row of Table III presents the objective evaluation results for DiffProsody (w/o VQ). DiffProsody (w/o VQ) performed worse for all the objective measurements. These results provide an objective assessment of how the failure to properly disentangle prosody affects the overall performance of the system.

V Conclusion

In this study, a novel technique called DiffProsody is proposed, the aims of which is to synthesize high-quality expressive speech. Through prosody conditional adversarial training, we observed significant improvements in speech quality with a more pronounced display of expressive prosody. In addition, our DLPG successfully generated expressive prosody. Our proposed method outperformed comparative models in terms of producing accurate and expressive prosody, as evidenced by the prosody evaluation metrics. Moreover, our method demonstrated superior accuracy in pronunciation, as indicated by the CER and WER evaluations. The KL divergence and histogram analysis further support the claim that DiffProsody yields a more accurate prosody distribution than the other models. Furthermore, we successfully reduced the sampling speed while maintaining the expected performance by introducing DDGAN.

VI Future works

Despite the importance of vector quantization for disentangling, it has been observed that this approach can negatively affect model performance. This problem is expected to be addressed with the introduction of methods such as residual vector quantizers. Moreover, we acknowledge the limitations of attempting to model prosody using the TTS dataset. It has been suggested that using a language model pre-trained on a large dataset, such as HuBERT, could result in significant improvements. In the future, we plan to extend the latent diffusion method to include controllable emotional prosody modeling.

References