Diff-HierVC: Diffusion-based Hierarchical Voice Conversion with Robust Pitch Generation and Masked Prior for Zero-shot Speaker Adaptation

Ha-Yeong Choi, Sang-Hoon Lee, Seong-Whan Lee

Introduction

Voice conversion (VC) tasks typically convert the voice of a source speaker into the voice of a specific target speaker, and the linguistic information of the converted target speaker must be consistent with the source speech. The primary concept of VC is to disentangle the individual components of speech so that each component can be controlled and transformed to the target speaker voice. Recently, the VC system has significantly advanced with deep learning approaches, allowing clarity and naturalness of the converted voice . Moreover, the expansion of the conversion system has enabled effective applications in a variety of fields, such as cross-lingual and emotional VC . Despite these advancements, converted voice is still perceived as unnatural owing to mispronunciation in converted speech, and low speaker adaptation performance is still a challenge that requires addressing .

Pitch modeling is essential to achieving speech intelligibility and naturalness in VC and text-to-speech (TTS) tasks. Pitch characteristics are crucial in speaker identity and correct pronunciation . train the model using normalized fundamental frequency (F0F_{0}) to obtain the same mean and variance for all speakers. This approach contributes to expressiveness by considering the pitch information. However, F0F_{0} is not entirely separated from the speaker style, so that it still causes perceptual unnaturalness in the conversion. SR uses a vector quantized variational auto-encoder (VQ-VAE) to learn a speaker-irrelevant pitch representation. Although speaker-irrelevant pitch can be extracted, mispronunciation occurs due to the loss of pitch information during vector quantization. In addition, it is difficult to precisely predict the pitch of a voice with a high degree of expressiveness. To address the above problem, we propose Diff-HierVC, a novel diffusion-based hierarchical VC system. Diff-HierVC consists of DiffPitch and DiffVoice, which hierarchically convert the voice style from disentangled speech representations. DiffPitch generates the pitch information of the target speaker during the inference step, and DiffVoice constructs a high-quality Mel-spectrogram utilizing the generated pitch information and the source-filter representation according to the source-filter theory. We found that a hierarchical VC architecture is an effective structure for decoupling speech components and generating the converted speech. Moreover, using the data-driven prior, we improved the conversion performance by regulating the inception of the denoising process of the diffusion model. Furthermore, we introduce a masked prior that allows the diffusion model to consider the context and condition for better generalization ability and robust training. The experimental results denote that pitch and voice modeling have considerable effects. Our main contributions are summarized as follows:

We propose Diff-HierVC, a diffusion-based hierarchical VC system with robust pitch generation and masked prior for expressive zero-shot voice style transfer.

To the best of our knowledge, this is the first study to utilize the diffusion process to generated F0F_{0}. We demonstrated that using the generated F0F_{0} by the denoising diffusion process rather than conventional pitch modeling methods resulted in more accurate pronunciation and natural intonation of the converted voice.

The experimental results reveal that Diff-HierVC achieves a significantly improved zero-shot style transfer in various conversion scenarios with cross-lingual and expressive real-world speech dataset. Our demos are available at https://diff-hiervc.github.io/.

Background: diffusion models

Diffusion models have shown extraordinary performance in generative tasks in various domains, such as images, videos, and audio, and have recently achieved considerable success in multi-modal tasks . Specifically, in speech, the diffusion model is utilized in applications such as audio generation , speech enhancement , and TTS synthesis . The fundamental concept underlying the stochastic differential equation (SDE)-based continuous-time diffusion process is to train an estimator that repeatedly removes noise by estimating log-density gradient of data and generates samples with an iterative denoising process via SDE. The SDE-based diffusion model was also applied to VC task using a maximum likelihood (ML)-SDE solver for fast sampling.

Diff-HierVC

As illustrated in Figure 1, we first analyze speech into representations of content, pitch, and style: (1) Data perturbation is applied to the input waveform to eliminate content-irrelevant information. Subsequently, we extract the content features from the intermediate layer representation of XLS-R , a pre-trained self-supervised model using a large-scale cross-lingual speech dataset. (2) We utilize a style encoder to extract the voice style, which is the speaker style representation from the Mel-spectrogram. The style embedding serves as a guide for both the content encoder and pitch encoder. (3) We extract a fundamental frequency (F0F_{0}) using the YAAPT algorithm with a 4×\times high-resolution higher than Mel-spectrogram for precise pitch extraction. The content encoder receives log⁡(F0+1)\log(F_{0}+1), and the pitch encoder takes the normalized F0F_{0} as the mean and variance of the source speaker's F0F_{0}.

2 Hierarchical VC

For hierarchical VC, we introduce a two-stage diffusion models, DiffPitch and DiffVoice. DiffPitch initially converts the F0F_{0} with the target voice style, and the converted F0F_{0} is fed to DiffVoice to convert the speech with the target voice style hierarchically. The details of each diffusion models are described as follows.

We introduce DiffPitch, a pitch generator based on the diffusion process. To consider continuous pitch information, we adopt a WaveNet based conditional diffusion model , which can iteratively obtain a significant receptive field with a single denoiser. The pitch encoder transforms the normalized F0F_{0} of the source speech to the pitch representation ZpZ_{p}. We regularized the pitch representation by pitch reconstruction loss to utilize ZpZ_{p} as a data-driven prior of DiffPitch as follows:

The diffusion process of DiffPitch uses the log-scale F0F_{0} extracted using the YAAPT algorithm as a target ground-truth XpX_{p}.

The forward process of the DiffPitch is defined as follows:

where t∈t\in, βt\beta_{t} regulates the amount of stochastic noise injected in the process, and wt{\mathbf{w}_{t}} is the forward standard Wiener process. DiffPitch executes denoising to recover the original pitch contour in the reverse process. The reverse process of the pitch denoiser is defined as follows:

where wtˉ\bar{\mathbf{w}_{t}} denote the backward standard Wiener process. According to , in the forward process, a sample of noisy pitch is drawn from the following distribution:

Therefore, DiffPitch approximates the score function with the following denoising objective:

where sθps_{\theta_{p}} is the pitch score estimator and λt=1−e−∫0tβsds\lambda_{t}=1-e^{-\int_{0}^{t}{\beta_{s}ds}}. Furthermore, we derive fast sampling using the ML-SDE solver , which maximizes the log-likelihood of forward diffusion with the reverse SDE solver. During inference, the converted F0F_{0} from the pitch encoder is utilized as a prior of DiffPitch, and DiffPitch generates the refined F0F_{0} with the target voice style ss. Note that we normalize F0F_{0} only with the statistic of a single sentence for the fair zero-shot voice conversion scenario.

2.2 DiffVoice

We present DiffVoice, a conditional diffusion model for high-quality speech synthesis from content, target F0F_{0}, and target voice style. We also utilize a data-driven prior for the diffusion models to guide the inception. According to the source-filter theory , we first disentangle the speech components into a pitch and content representation. For a data-driven prior of DiffVoice, the source-filter encoder which consists of the source encoder EsrcE_{src} and filter encoder EftrE_{ftr} reconstructs the intermediate Mel-spectrogram ZmZ_{m} from the disentangled speech representation as Zm=Zsrc+ZftrZ_{m}=Z_{src}+Z_{ftr}, where Zsrc=Esrc(F0,s)Z_{src}=E_{src}(F_{0},s), Zftr=Eftr(content,s)Z_{ftr}=E_{ftr}(content,s), and ss denotes style embedding. Mel-spectrogram ZmZ_{m} is regularized as follows:

where XmelX_{mel} is the Mel-spectrogram of the ground-truth speech. Subsequently, DiffVoice can utilize the source-filter encoder output ZmZ_{m} as a prior, and use speaker representation ss as condition to maximize speaker adaptation capacity.

The following equation describes the forward process of DiffVoice:

The reverse process of DiffVoice is defined by:

In the forward process, a sample of noisy Mel-spectrogram Xm,tX_{m,t} is taken in the same manner as equation (4). Finally, the objective of training the Mel-spectrogram noise estimation network sθms_{\theta_{m}} is to optimize the score matching loss:

During inference, the source-filter encoder takes a content representation from the source speech, target voice style ss, and the converted F0F_{0} from DiffPitch with the target voice style. The converted Mel-spectrogram ZmZ_{m} from the source-filter encoder is used as a data-driven prior, and DiffVoice generates the converted speech conditioned with the target voice style.

3 Denoising models with masked prior

Although the data-driven prior can significantly improve the conversion performance, DiffVoice may rely on the reconstructed Mel-spectrogram in the source-filter encoder. To improve generalization performance of DiffVoice, we introduce a masked prior to the denoising diffusion models. Before fed to the DiffVoice, the prior ZmZ_{m} is masked, and the diffusion network jointly learns the reconstruction and denoising process. Consequently, the model can reconstruct the masked area from the surrounding context. Specifically, we apply frequency masking by interpreting continuous pitch information from a contextual point of view.

Experiment and result

We train the model with a large-scale publicly available multi-speaker dataset, LibriTTS . We utilize the train-clean-360 and train-clean-100 subsets of LibriTTS, which contain 245 hours of speech from 1,151 speakers. We additionally use dev-clean-other subsets of LibriTTS for validation. Then, we use VCTK dataset to evaluate the zero-shot VC performance. We randomly select sentences from the paired speech of VCTK dataset. We downsample the audio to 16 kHz, and transform the audio into a log-scale Mel-spectrogram with 80 bins using short-time Fourier transform (STFT) and Mel-filters. We use a hop size of 320 and a window size of 1,280 to map the time-resolution of the self-supervised speech representation.

1.2 Training

We train the model using LibriTTS for 2M steps with a batch size of 64 on two NVIDIA A100 GPUs (five days), and use AdamW optimizer with the setting of , and implemented the learning rate schedule with a decay of 0.9991/80.999^{1/8} at an initial learning rate of 5×10−55\times 10^{-5}. We segment the audio clip into 35,840 frame during training. For fine-tuning, we set the initial learning rate to 2×10−52\times 10^{-5}. A non-causal dilated WaveNet with 128 dimensions is used for all encoders, and DiffPitch uses the DiffWave with 64 dimensions and additional conditional layers for the pitch and style representations. DiffVoice consists of a 2D-UNet structure with the initial channel of 64 and three blocks, and the dimension of blocks are . Following , we use the noise schedule parameters of β0\beta_{0} and β1\beta_{1} with 0.05 and 20 respectively. The masking ratio is set to 30% for the masked prior. For vocoder, we train the HiFi-GAN with the same training dataset, and we only replace the discriminators with multi-scale STFT discriminator of EnCodec .

Most previous VC systems utilize normalized or quantized F0F_{0} for speaker-irrelevant pitch modeling. However, we estimate a raw F0F_{0} with a target voice style for better speaker adaptation. We compared three F0F_{0} prediction methods: F0F_{0} transformation with a statistic of F0F_{0}, simple F0F_{0} prediction with WaveNet, a diffusion-based F0F_{0} prediction with DiffPitch. Figure 2 depicts that DiffPitch with 30 iteration steps has a similar F0F_{0} contour with the ground-truth F0F_{0}. Figure 3 also show the diversity of pitch contours with different target voice styles. Hence, we utilize the converted F0F_{0} by DiffPitch during VC with DiffVoice.

3 Zero-shot VC

We conduct various subjective and objective evaluation on the zero-shot VC scenario with three models: (1) autoencoder based VC model, AutoVC , (2) GAN based VC model, VoiceMixer, (3) unit-based end-to-end speech model, Speech Resynthesis (SR) , and (4) diffusion-based VC model, DiffVCTo train with the same settings as our model, we trained the speaker encoder of DiffVC with train-clean-100 and 360 with 1,151 speakers. Also, DiffVC* denotes the VC results using the official checkpoint. The official code implementation uses a speaker encoder trained with a large multi-speaker dataset (voxceleb1, voxceleb2, and LibriTTS-other) containing 8,371 speakers to extract and use the speaker embedding.. Following , we conduct the naturalness and similarity mean opinion score (nMOS and sMOS, respectively). Table 1 depicts that our model has a better nMOS and sMOS than the others. Specifically, our model achieves significantly improved content consistencyWe use an automatic speech recognition model, Whisper-large , and calculate the character error rate (CER) and word error rate (WER) on the 400 converted speeches with the text normalizer and the speaker adaptation performanceFor 400×\times20 = 8,000 paired speeches, we measure the equal error rate (EER) of automatic speaker verification model . We use Resemblyzer to calculate the speaker encoder cosine similarity (SECS).. In addition, we conducted cross-lingual VC to demonstrate zero-shot conversion performance in unseen languages. Figure 4 shows the robust generalization performance of our model in both resynthesis and VC scenarios, even for unseen languages. Furthermore, we fine-tune the model using only one sample per speaker. Fine-tuning with small steps (1,000 steps) can improve the performance of speaker adaptation. However, the model fine-tuned with more steps shows a lower robustness of content consistency by decreasing the CER and WER.

4 Ablation study

We compared three pitch modeling methods: DiffPitch, F0F_{0} transformation with denormalization (Denorm.) , and a simple F0F_{0} prediction with the F0F_{0} Encoder. All methods utilize the same normalized F0F_{0} of the source speech to convert the F0F_{0} with target voice style. Although Denorm. could transform the normalized F0F_{0} with the mean and variance of target speech, inaccurate F0F_{0} extracted from target speech decreases the voice style transfer performance regarding CER and WER with a mispronunciation and inaccurate intonation as indicated in Table 2. In addition, using only the F0F_{0} encoder decreases the voice style transfer performance with a higher EER than the DiffPitch even with the same WaveNet structure.

4.2 Data-driven prior and masked prior

As indicated in Table 2, the data-driven prior significantly improves the voice style transfer performance. However, we found that the diffusion models may rely on the performance of source-filter encoder and the diffusion models slightly reflect the conditional information to generate the converted speech. In addition, the inaccurate ground-truth F0F_{0} extracted by YAPPT is sometimes fed to the models during training and inference. Employing masked prior improves the performance with better generalization on the diffusion models taking advantage of data-driven prior. In addition, an experiment was carried out to determine the suitable masking ratio, and as a result of Table 3, a masking ratio of 30% performed best.

4.3 Source-filter encoder

It is well known that disentangling the speech plays a important role in controlling speech representation. In this work, we adopt the source-filter (SF) encoder to disentangle the speech components and regulate the starting point of diffusion models. To evaluate the effectiveness of source-filter encoder, we replace the source-filter encoder with a single encoder. Using the single encoder decreases the performance of voice conversion of the entire model, which could not appropriately disentangle the speech representation and it results in the converted speech for a prior of DiffVoice having a lower speaker similarity with the target speech as indicated in Table 2.

Conclusion

In this paper, we presented Diff-HierVC, a diffusion-based hierarchical VC system for high-fidelity converted pitch and Mel-spectrogram generation. DiffPitch improves performance in terms of speaker similarity and phonetic intelligibility. Then, DiffVoice restores high-quality speech through a denoising process. Subsequently, for better generalization of the diffusion model, we proposed a masked prior that can be robustly converted by considering the context and diffusion conditions. Consequently, our model outperformed the state-of-the-art in all metrics even with 6.8×\times fewer parameters, and we demonstrated the feasibility of building the zero-shot cross-lingual VC system, which can Break Down Barriers on various low-resource speech and language technologies. However, although our methods can significantly improve speaker adaptation quality, there are cases where the noise of input data is also considered as style. Hence, there is room for improvement towards high-quality and noise-free audio. In future work, we will decouple the noise and speech style with noise augmentation to generate high-fidelity audio even in a noisy environment.

Acknowledgements

This work was partly supported by Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No. 2019-0-00079, Artificial Intelligence Graduate School Program (Korea University) and No. 2021-0-02068, Artificial Intelligence Innovation Hub) and ESTsoft Corp., Seoul, Korea.

References