DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
Hyung-Seok Oh, Sang-Hoon Lee, Seong-Whan Lee
I Introduction
Recent advancements in neural text-to-speech (TTS) models have significantly enhanced the naturalness of synthetic speech. In several studies, prosody modeling has been leveraged to synthesize speech that closely resembles human expression. Prosody, which encompasses various speech properties, such as pitch, energy, and duration, plays a crucial role in the synthesis of expressive speech.
In some studies, reference encoders have been used to extract prosody vectors for prosody modeling. A global style token (GST) is an unsupervised style modeling method that uses learnable tokens to model and control various styles. Meta-StyleSpeech proposes the application of style vectors extracted using a reference encoder through a style-adaptive layer norm. Progressive variational autoencoder TTS presents a method for gradual style adaptation. A zero-shot method for speech synthesis that comprises the use of a normalization architecture, speaker encoder, and feedforward transformer-based architecture was proposed. Despite the intuitive and effective nature of using a reference encoder, these methods cannot reflect the details of prosody in their synthetic speech without ground-truth (GT) prosody information.
Recently, methods for inferring prosody from text in the absence of reference audio have been developed . FastPitch, for instance, synthesizes speech under text and fundamental frequency conditions. FastSpeech 2 aims to generate natural, human-like speech by using extracted prosody features, such as pitch, energy, and duration, through an external tool and introduce a variance adaptor module that predicts these features. Some studies have proposed hierarchical models through the design of prosody features at both coarse and fine-grained levels. However, the separate modeling of prosodic features may yield unnatural results owing to their inherent correlation.
Some studies have predicted a unified prosody vector, thus enhancing the representation of prosody, given the interdependence of prosody features. Text-predicted GST is a method for modeling prosody without reference audio by predicting the weight of the style token from the input text. proposed a method for paragraph-based prosody modeling by introducing a paragraph encoder. Gaussian-mixture-model-based phone-level prosody modelling is a method for sampling reference prosody from Gaussian components. proposed a method for modeling prosody using style, perception, and frame-level reconstruction loss. There are also studies in which prosody is modeled using pre-trained language models. ProsoSpeech models prosody with vector quantization (VQ) using large amounts of data and predicts the index of the codebook using an autoregressive (AR) prosody predictor. However, when predicting prosody vectors, an AR prosody predictor encounters challenges related to long-term dependencies.
To address these issues, we propose DiffProsody, a novel approach that generates expressive speech by employing a diffusion-based latent prosody generator (DLPG) and prosody conditional adversarial training.
The primary contributions of this work are as follows:
We propose a diffusion-based latent prosody modeling method that can generate high-quality latent prosody representations, thereby enhancing the expressiveness of synthetic speech. Furthermore, we adopted denoising diffusion generative adversarial networks (DDGANs) to reduce the number of timesteps, resulting in speeds that were 2.48× and 16× faster than those of the AR model and denoising diffusion probabilistic model (DDPM) , respectively.
We propose prosody conditional adversarial training to ensure an accurate reflection of prosody using the TTS module. A significant improvement in smoothness, attributable to vector quantization, was observed in the generated speech.
Objective and subjective evaluations demonstrated that the proposed method outperforms comparative models.
The implementationhttps://github.com/hsoh0306/DiffProsody of proposed method and audio sampleshttps://prml-lab-speech-team.github.io/demo/DiffProsody/ for various datasets, such as VCTKhttps://datashare.ed.ac.uk/handle/10283/2651 and LibriTTS, are available online.
II Related works
Traditional TTS models function autoregressively. This implies that the spectrogram generates one frame at a time, with each frame conditioned on the preceding frames. Despite the high-quality speech that this approach can produce, it has a drawback in terms of speed owing to the sequential nature of the generation process. To address this issue, non-autoregressive TTS (NAR-TTS) models have been proposed as an alternative for parallel generation. These models have the advantage of simultaneously generating the entire spectrogram, thus resulting in a significant acceleration of the speech synthesis. FastSpeech and FastSpeech 2 serve as examples of NAR-TTS models that can synthesize speech at a much faster rate than their AR counterparts while maintaining a comparable level of quality. For parallel generation, these models require phoneme-level durations. FastSpeech uses an AR teacher model to obtain durations through knowledge distillation. The phoneme-level input is scaled to the frame-level using a length regulator, and a transformer-based network is used to generate the entire utterance at once. FastSpeech 2 addresses some of the disadvantages of FastSpeech by extracting the phoneme duration from forced alignment as the training target instead of relying on the attention map of the AR teacher model and introducing more variation information in speech as conditional inputs. In contrast to these models that use an external aligner, are parallel TTS models that use an internal aligner to model duration. These parallel-generation models exhibit faster and more robust generation than the AR models. In this study, we adopted a transformer-based NAR model with a simple structure to focus on prosody modeling.
II-B Generative adversarial networks
Generative adversarial networks (GANs) are generative models in which generative and discriminative networks compete against each other. The objective of the generative network is to create samples that closely resemble the true data distribution, whereas the discriminative network strives to differentiate between the data sampled from the true and generated distributions.
These two networks play a minimax game with the following value function :
where represents the generative network, and represents the discriminative network. The training process involves minimizing to recognize as real for , and maximizing for to learn the likelihood of real data. The conditional generation model was obtained by introducing condition into both and . In this study, we incorporated adversarial training conditions into prosody features to generate expressive speech.
II-C Denoising diffusion models
The denoising diffusion model is a generative model that gradually collapses data into noise and generates data from noise. The processes of collapsing and denoising data are called the forward and reverse processes, respectively. The forward process gradually collapses data into noise over the -step, with a predefined variance schedule .
where . The reverse process is defined as follows:
The reverse process is driven by the denoising model parameterized by a. The denoising model was optimized for a variational bound on the negative log-likelihood.
In the DDPM, the denoising distribution is assumed to comprise a Gaussian distribution. Moreover, it has been demonstrated that the diffusion model can generate a diverse range of complex distributions, provided that a sufficient number of iterations are performed. Proceeding with only a small step at a time is possible by setting the denoising distribution to a Gaussian distribution, which implies that a considerable number of timesteps are required. DDGAN was proposed by modeling the denoising distribution with a non-Gaussian multimodal distribution to reduce the sampling step. It predicts with the generator , which models the implicit distribution, in contrast to the denoising network of the DDPM, which predicts noise. In the DDGAN, the conditional probability is defined as follows:
where represents the implicit distribution imposed by the generator , which outputs given and the latent variable . In the DDGAN, the denoising distribution is modeled as a complex multimodal distribution, in contrast to the unimodal distribution in the DDPM. Sampling with small timesteps is possible by leveraging the complicated . We introduced this model to reduce the sampling timesteps while maintaining its ability to generate diffusion models.
II-D Prosody modeling
Although the pronunciation capabilities of TTS models have seen significant advancements, they still fail to replicate the naturalness inherent in human speech. Various studies have proposed methods for prosody modeling to address this limitation. One such method is the reference-based approach, which is a type of expressive speech synthesis that extracts styles from the reference audio. This approach is particularly beneficial for style transfers, but can result in unnatural results when the text and reference audio do not align well. However, there are methods that directly model prosody properties such as pitch, energy, and duration. These methods offer the advantages of explainability and controllability because they directly use and model prosody features. In the context of ProsoSpeech, the authors argue that prosodic features are interdependent and that modeling them separately can result in unnatural outcomes. To address this, they proposed the modeling of prosody as a latent prosody vector (LPV) and introduced an AR prosody predictor to obtain the LPV. In this study, we adopt a similar latent vector approach to model prosody. The use of pre-trained language models was also suggested . These models comprise the use of models that have been pre-trained on large datasets, such as BERT or GPT-3.
III DiffProsody
The proposed method, called DiffProsody, aims to enhance speech synthesis by incorporating a diffusion-based latent prosody generator (DLPG) and prosody conditional adversarial training. The overall structure and process of DiffProsody are presented in Figure 1. In the first stage, we trained a TTS module and a prosody encoder using a text sequence and a reference Mel-spectrogram as inputs. The prosody conditional discriminator evaluates the prosody vector from the prosody encoder and the Mel-spectrogram from the TTS module to provide feedback on their quality. In the second stage, we train a DLPG to sample a prosody vector that corresponds to the input text and speaker. During inference, the TTS module synthesizes speech without relying on a reference Mel-spectrogram. Instead, it uses the output of a DLPG. This facilitates the generation of expressive speech that accurately reflects the desired prosody.
The TTS module is designed to transform text into Mel-spectrograms using speaker and prosody vectors as conditions. The overall structure of the model is presented in Figure 1a. The TTS module comprises a text encoder and a decoder. The text encoder processes the text at both the phoneme and word levels, as illustrated in Figure 1c. The input text, denoted as , is converted into a text hidden representation , by the phoneme encoder and word encoder . The takes the phoneme-level text and the takes as input the word-level text . The is then obtained as the element-wise sum of the outputs of and expanded to the phoneme-level.
where is an operation that expands the word-level features to the phoneme-level. Obtaining the quantized prosody vector involves using and speaker hidden representation as inputs for the prosody module. In addition, is acquired using a pre-trained speaker encoder. We use Resemblyzerhttps://github.com/resemble-ai/Resemblyzer, an open-source model trained with generalized end-to-end loss (GE2E), to extract .
During the first stage of training, a prosody encoder is employed, which receives the target Mel-spectrogram. In the inference, is obtained by inputting and into a DLPG, and this is performed without a reference Mel-spectrogram. Finally, the information related to the text, speaker, and prosody is combined by expanding the latent vectors , , and to the phoneme-level and then performing an element-wise summation.
The phoneme duration is modeled using the duration predictor . The goal of the is to predict the phoneme duration at the frame-level based on the input variable .
In addition, there is a length regulator that expands the input variable to the frame-level using the phoneme duration . The expanded is then transformed to Mel-spectrogram by .
For TTS modeling, we use two types of losses: the mean square error (MSE) and structural similarity index (SSIM) loss. These losses aid in accurately modeling the TTS. For the duration modeling, we use the MSE loss.
III-B Prosody module
During the inference stage, the prosody vector is obtained using the prosody generator trained in the second stage.
where is -th element of . In the first stage, the TTS module is trained jointly with the codebook and prosody encoder .
where denotes the stop-gradient operation. Moreover, we employ an exponential moving average (EMA) to enhance the learning efficiency by applying it to codebook updates.
III-C Prosody conditional adversarial training
For our prosody adversarial training, we develop prosody conditional discriminators (PCDs) that handle inputs of varying lengths. This design is inspired by multi-length window discriminators. The PCD structure is presented in Figure 1e. The PCD is designed to accept a Mel-spectrogram and a quantized prosody vector as inputs, and its role is to determine whether these input features are original or generated. The PCD comprises two lightweight convolutional neural networks (CNNs) and fully connected layers. One of the CNNs is designed to receive only the Mel-spectrogram, whereas the other is designed to receive a combination of and . To match the corresponding PCD, the length of each Mel-spectrogram and the extended are randomly clipped. For our objective function, we adopt the least square GAN loss:
where denotes the training goal of the discriminators and represents the feedback on the TTS module. The final object of the TTS module is as follows:
where corresponds to the weight of the adversarial loss.
III-D Diffusion-based latent prosody generator
We propose a new module called the DLPG, which leverages the powerful generative capabilities of diffusion models. In addition, we introduce a DDGAN framework, which enables faster sampling by reducing the number of required timesteps. Figure 2 presents the training process for the DLPG. During training, the DLPG aims to generate target , which is extracted from the prosody encoder trained in the first stage. The DLPG is trained to produce based on the and . In the diffusion model, we set as the target . The DLPG generator directly generates .
where is timestep of diffusion process. To ensure adversarial training, is derived from and using posterior sampling . Subsequently, a time-dependent discriminator determines the compatibility of (obtained from the forward processing of ) and (generated through posterior sampling of ) with respect to and , conditioned on and . The objective function of is then defined as follows:
where corresponds to the adversarial loss, and denotes the reconstruction loss of . The total generator loss is expressed as follows:
where is the weight of adversarial loss . The objective function of the is as follows:
The DLPG leverages the DDGAN framework to achieve stable and high-quality results in only a few timesteps. This process involves the , which iteratively generates times during the inference. We set to follow a normal distribution. The is obtained as the final of the reverse process. The final prosody vector is then derived through the vector quantization of .
III-E Inference
Here is the step-by-step process of generating a Mel-spectrogram using the trained TTS module and DLPG.
The is extracted from the text encoder, and is extracted from the pre-trained speaker encoder.
The DLPG generates a with and as inputs.
is mapped to a codebook, denoted as , in the VQ layer to obtain the prosody vector, which is denoted as .
The decoder generates a Mel-spectrogram using , and . This process involves the expansion , , and to the frame-level. The phoneme-duration is predicted by the duration predictor.
The Mel-spectrogram was converted to a raw waveform using a pre-trained vocoder.
IV Experimental result and discussion
We conducted experiments using the VCTK dataset, a multispeaker English dataset consisting of audio clips of 44,200 sentences recorded by 109 speakers for approximately 400 sentences. A total of 2,180 audio clips were randomly selected with 20 sentences per speaker, which constituted the test set. Furthermore, 545 audio clips randomly selected from five sentences per speaker were used as the verification sets, and the remainder were used as the training sets. We sampled audio at 22,050 Hz and then transformed it to an 80-bin Mel-spectrogram using an STFT with a window length of 1,024 and hop size of 256. The text was converted into phoneme sequences for text input using a grapheme-to-phoneme toolhttps://github.com/Kyubyong/g2p. We extracted the phoneme duration using the Montreal Forced Aligner tool. Furthermore, we used the AdamW optimizer with and . The learning rates for the TTS and latent prosody generator training were set as and , respectively. Throughout the training process, the batch size used was 48. The TTS and prosody encoder were updated in 160k steps, and the latent prosody generator was updated in 320k steps. In the experiment, all the audio was synthesized using the official implementation of the HiFi-GANhttps://github.com/jik876/hifi-gan and a pre-trained model. We trained the TTS module and prosody encoder for approximately 16 h and the DLPG for 7 h using a single NVIDIA RTX A6000 GPU.
IV-B Implementation details
The number of layers, hidden size, filter size, and kernel size of the feed-forward transformer blocks of the phoneme encoder, word encoder, and decoder were set as 4, 192, 384, and 5, respectively. The extracted speaker embedding was projected onto 192 dimensions, and the number of dimensions of the prosody vector was 192. The structure of the prosody encoder follows that of ProsoSpeech. For VQ, we set the size of codebook and dimension of code as 128 and 192, respectively. and updated it using EMA with a decay rate of 0.998. The codebook was initialized as the center of the k-means clustering after 20k steps of TTS training. The number of Mel-spectrogram bins used in the prosody encoder was set as 20. We set multiple PCDs to receive different input sizes, such as. The PCD comprises two 2D convolution stacks and three fully connected layers. The convolution stacks in the PCD consist of three 2D convolutions: LeakyReLU, BatchNorm, and a linear layer with a multilength window size of . The latent prosody generator consists of 20 residual blocks with a hidden size of 384. The prosody discriminator is configured with four convolution layers and hidden dimensions of 384. We investigated and between 0.001 and 1.0 and set and as 0.01 and 0.05, respectively. The number of timesteps used in the training and inference was set as four.
IV-C Comparative studies
We developed a prosody model from previous works, our proposed model, and the model of the ablation study for performing comparative experiments. All the models produced an 80-bin Mel-spectrogram and synthesized speech using the same vocoder.
IV-C2 GT (vocoded)
The generated audio was obtained by converting the GT Mel-spectrogram using HiFi-GAN V1.
IV-C3 FastSpeech 2
FastSpeech 2 is a speech-synthesis model that directly predicts prosody features (pitch and energy). We trained FastSpeech 2 using an open-source implementationhttps://github.com/NATSpeech/NATSpeech. For the purpose of a fair comparison, we used a text encoder at the phoneme and word-levels.
IV-C4 ProsoSpeech
ProsoSpeech is a speech synthesis model that models prosody features as latent prosody vectors and predicts them using an AR predictor. We implemented the model by following the hyperparameters in the study and used the model that provided the best performance.
IV-C5 DiffProsody
The proposed model was implemented using a text encoder and decoder with the same structure as FastSpeech 2 and ProsoSpeech. We used a prosody encoder with the same structure as that of ProsoSpeech. In the first stage, a prosody conditional discriminator was added. In the second stage, a DLPG was used instead of an AR predictor.
IV-C6 DiffProsody (AR)
DiffProsody (AR) is a model in which a DiffProsody trained with an AR prosody predictor to estimate prosody vectors. It has the same structure as the ProsoSpeech prosody predictor.
IV-C7 DiffProsody (DDPM)
DiffProsody (DDPM) is a model in which a DiffProsody trained with a DDPM framework to estimate a prosody vector. It has the same structure as DiffProsody’s diffusion-based latent generator.
IV-C8 DiffProsody (w/o PCD)
DiffProsody (w/o PCD) is a model in which a DiffProsody trained without the assistance of a prosody conditional discriminator.
IV-C9 DiffProsody (w/o VQ)
DiffProsody (w/o VQ) is a model in which a DiffProsody trained without a VQ layer in the prosody encoder.
IV-C10 DiffProsody N
DiffProsody is a model in which a DiffProsody trained with the lowest Mel-bins in the prosody encoder.
IV-D Subjective metrics
A subjective assessment is conducted to confirm the effectiveness of the proposed method. To measure the level of naturalness, we used the mean opinion score (MOS). For this evaluation, we employed Amazon Mechanical Turk (MTurk), a crowdsourcing service, to gather feedback from 20 native Americans. MOS was assessed using a 5-point scale, and confidence intervals were calculated at a 95% level. For the evaluation, 100 samples were randomly selected from the test set.
IV-E Objective metrics
For realizing an objective evaluation, we calculated the equal error rate (EER) using a pre-trained speaker verification modelhttps://github.com/clovaai/voxceleb_trainer . We used the pre-trained wav2vec 2.0 to the compute character error rate (CER) and word error rate (WER).
For the prosodic evaluation, we computed the average differences in utterance duration (DDUR), pitch error (RMSE) (in cents), periodicity error (RMSEperiod) and F1 score of the voiced/unvoiced classification (F1v/uv). We used torchcrepehttps://github.com/maxrmorrison/torchcrepe to extract the pitch and periodicity features for evaluation. In addition, we measured the Kullback–Leibler (KL) divergence of log f0 and log energy to compare the distributions of the prosody features in the generated audio. Finally, we calculated the real-time factor (RTF) to compare the generation speeds.
We obtained by calculating the mean absolute error of the difference in the duration of each utterance.
where denotes the duration of the -th GT utterance and denotes the duration of the -th generated utterance.
IV-E2 Pitch error
To measure the pitch error RMSE, we aligned the pitches extracted in hertz with dynamic time warping (DTW) and calculated the root mean square (RMSE) in cents, defined as for pitch of the GT speech and pitch of the generated speech . We measured the portion wherein both the GT speech and generated speech were voiced.
IV-E3 Periodicity error
We measured the periodicity error RMSEperiod by root mean square between the periodicities aligned with the DTW.
where means the -th periodicity value.
IV-E4 F1 score of voiced/unvoiced
We obtained voiced/unvoiced flags from the aligned pitches using DTW and calculated the F1 score (F1v/uv) between them. We defined a match between the GT voiced flag and generated voiced flag as a true positive (TP), a match between the GT unvoiced flag and generated voiced flag as a false positive (FP), and a match between the GT voiced flag and generated unvoiced flag as a false negative (FN).
where is the length of the sequence, and is a function that returns 1 if the -th element has the same value, and 0 otherwise. The precision and recall are defined as follows:
We then calculated the F1 score as follows:
IV-E5 KL divergence of log f0 / log energy
We also measured the KL D to analyze the pitch and energy distribution. We first extracted the pitch and energy in log-scale. We then binned the entire range into 100 bins and applied a kernel density estimation to each bin to calculate the KL D for a smoothed distribution. The KL divergence for feature is defined as follows:
where is the number of bins, is the probability in the -th bin of the distribution of , is the kernel density estimator function, and is the KL divergence calculation.
IV-F Evaluation results
Table I lists the MOS results and objective evaluations. The results demonstrate that the proposed model DiffProsody surpasses the other models in terms of both subjective and objective metrics. In particular, our model achieved a superior MOS for the subjective metrics, with a p-value of less than 0.05. It also exhibited outstanding performance in terms of the DDUR, EER, RMSE, RMSEperiod, and F1v/uv, which suggests that the speech generated by DiffProsody closely resembles the target prosody. Furthermore, our model displayed a lower WER and CER, thus indicating its capability for synthesizing speech with a more accurate pronunciation.
Figure 3 presents the Mel-spectrogram and pitch contour of speech from each model. The red box in the figure indicates that the DiffProsody model was more similar to the GT. To further evaluate the prosody of the generated speech, we examined pitch and energy distributions.
Figure 5 presents the histogram distribution for log f0 (pitch), and Figure 5 presents the histogram distribution for log energy. In both the figures, the blue bars represent the distribution of the GT features, and the orange bars represent the distribution of the generated features. The results indicate that the proposed model aligns more closely with the GT distribution than the other models. Table II lists the KL divergence values for comparison. ProsoSpeech exhibited a better performance in f0 than FastSpeech 2, but not in the case of energy. However, DiffProsody outperformed the comparison model in terms of f0 and energy.
IV-G Latent prosody generation method
In Table III, DiffProsody (AR) is a model comprising the use of an AR prosody predictor. DiffProsody (DDPM) is a model comprising the use of DDPM (100 timesteps). DiffProsody is a model comprising the use of DDGAN (four timesteps). The AR predictor has the same structure as the prosody predictor in ProsoSpeech, but it does not use context encoders, and the denoising network in the DDPM has the same structure as generator in the DDGAN. We also conducted a 7-point comparative mean opinion score (CMOS) evaluation to compare the latent prosody generation methods and measured the RTF to compare the generation speeds. DDGAN achieved 0.172 and 0.015 CMOS compared with AR and DDPM, respectively. The objective metrics showed that the DDPM and DDGAN outperformed the AR. The DDGAN achieved results nearly identical to those of the DDPM for all the metrics. The experimental results showed that the diffusion models performed better than the AR models, as reported in. Furthermore, DDGAN can generate high-quality prosody vectors such as DDPM in only four timesteps. According to the RTF results, the DDGAN produces a 2.7 and 16 faster prosody than the AR and DDPM, respectively.
Figure 6 presents the traces of the WER, CER, and EER results for the speech generated by the DLPG with the prosody vector generated by each denoising iteration. The red and blue lines represent the results obtained using the DDPM and DDGAN, respectively. We observed that the DDGAN has a larger error rate than the DDPM in the early stages and that the midpoint and final stages of each model have almost the same value. This implies that the DDGAN compresses the timesteps of the DDPM by modeling the denoising process as a multimodal non-Gaussian distribution.
IV-H Prosody conditional adversarial training
We conducted CMOS to assess the effectiveness of PCD. The CMOS values were used to measure the degree of PCD preference in comparison with the reference model. Our evaluation involved the analysis of three different models: DiffProsody using PCD and a diffusion-based model, DiffProsody without PCD (referred to as DiffProsody (w/o PCD)) using a diffusion-based model, and ProsoSpeech using an AR prosody predictor without PCD. The results presented in Table IV clearly indicate that models comprising the use of the PCD are preferred over those trained without the PCD. Furthermore, by comparing the CMOS results of DiffProsody (w/o PCD) and ProsoSpeech, we can infer that the diffusion-based method is preferable to the method employed by ProsoSpeech. To provide a more comprehensive analysis, we also incorporated the objective metrics for DiffProsody (w/o PCD) in Table III. These objective metric results demonstrate that DiffProsody outperformed DiffProsody (w/o PCD) in all aspects. However, DiffProsody (w/o PCD) still received a better objective evaluation than ProsoSpeech. In addition, when comparing DiffProsody (AR) with ProsoSpeech, DiffProsody (AR) consistently achieved higher scores on the majority of the objective evaluations. These findings validate that both PCD and DLPG significantly improve the model performance.
IV-I Prosody encoder evaluation
In this section, the effectiveness of the prosody encoder is evaluated. We focus on two main aspects: the impact of the number of Mel-bins used in the reference Mel-spectrogram and the role of the vector quantization layer. Detailed results of these evaluations are presented in the following subsections.
We conducted an experiment to compare the performance of DiffProsody when trained using various numbers of Mel-bins. The results, including the EER, CER, and WER scores, are presented in Table V. We examined the performance of 10, 20, 30, 40, 60, and 80 (full-band) Mel-bins. The findings indicated that as the number of bins exceeded 20 (baseline), the EER tended to increase, while no significant difference was observed in terms of the CER and WER. It should be noted that the model with 10 bins outperformed the baseline (20 bins) in terms of the WER but yielded higher EER results. Figure 7 illustrates the Mel-spectrogram at the initial iteration (noise) during the diffusion timestep of the DLPG, which was trained using various numbers of Mel-bins. As the number of Mel-bins used for the training was increased, the results of the Mel-spectrogram progressively smoothened and eventually collapsed. This phenomenon occurs because the large amount of information in the reference forces the model to reconstruct the Mel-spectrogram by leveraging the prosody vectors. Through this experiment, we found that, as increases, linguistic information becomes increasingly entangled. Consequently, it is reasonable to employ 20 Mel-bins as the input to the prosody encoder for realizing effective prosody modeling.
IV-I2 Vector quantization layer analysis
Figure 8 presents a Mel-spectrogram trace synthesized using the prosody vector generated for each diffusion step of the DLPG. We compared two versions of DiffProsody: one trained without the vector quantization layer (Figure 8a; defined as DiffProsody (w/o VQ)) and the other trained with the vector quantization layer (Figure 8b; defined as DiffProsody). In the case of DiffProsody (w/o VQ), the early steps exhibited a completely distorted Mel-spectrogram, but there was a significant recovery in the middle steps. Conversely, the initial step of DiffProsody exhibits a smoothed but slightly distorted Mel-spectrogram that gradually returns to its original state over the subsequent steps. This phenomenon is also related to prosody disentangling, and we confirmed that the prosody disentangling failed in DiffProsody (w/o VQ). This experiment demonstrated that vector quantization plays a crucial role in effective prosody disentangling. The last row of Table III presents the objective evaluation results for DiffProsody (w/o VQ). DiffProsody (w/o VQ) performed worse for all the objective measurements. These results provide an objective assessment of how the failure to properly disentangle prosody affects the overall performance of the system.
V Conclusion
In this study, a novel technique called DiffProsody is proposed, the aims of which is to synthesize high-quality expressive speech. Through prosody conditional adversarial training, we observed significant improvements in speech quality with a more pronounced display of expressive prosody. In addition, our DLPG successfully generated expressive prosody. Our proposed method outperformed comparative models in terms of producing accurate and expressive prosody, as evidenced by the prosody evaluation metrics. Moreover, our method demonstrated superior accuracy in pronunciation, as indicated by the CER and WER evaluations. The KL divergence and histogram analysis further support the claim that DiffProsody yields a more accurate prosody distribution than the other models. Furthermore, we successfully reduced the sampling speed while maintaining the expected performance by introducing DDGAN.
VI Future works
Despite the importance of vector quantization for disentangling, it has been observed that this approach can negatively affect model performance. This problem is expected to be addressed with the introduction of methods such as residual vector quantizers. Moreover, we acknowledge the limitations of attempting to model prosody using the TTS dataset. It has been suggested that using a language model pre-trained on a large dataset, such as HuBERT, could result in significant improvements. In the future, we plan to extend the latent diffusion method to include controllable emotional prosody modeling.