NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, Zhizheng Wu, Tao Qin, Xiang-Yang Li, Wei Ye, Shikun Zhang, Jiang Bian, Lei He, Jinyu Li, Sheng Zhao
Introduction
In recent years, significant advancements have been achieved in text-to-speech (TTS) synthesis. Traditional TTS systems are typically trained on limited datasets recorded in studios, and thus fail to support high-quality zero-shot speech synthesis. Recent works have made considerable progress for zero-shot TTS by largely scaling up both the corpus and the model sizes. However, the synthesis results of these large-scale TTS systems are not satisfactory in terms of voice quality, similarity, and prosody.
The challenges of inferior results stem from the intricate information embedded in speech, since speech encompasses numerous attributes, such as content, prosody, timbre, and acoustic detail. Previous works using raw waveform and mel-spectrogram as data representations suffer from these intricate complexities during speech generation. A natural idea is to factorize speech into disentangled subspaces representing different attributes and generate them individually. However, achieving this kind of disentangled factorization is non-trivial. Previous works encode speech into multi-level discrete tokens using a neural audio codec based on residual vector quantization (RVQ). Although this approach decomposes speech into different hierarchical representations, it does not effectively disentangle the information of different attributes of speech across different RVQ levels and still suffers from modeling complex coupled information.
To effectively generate speech with better quality, similarity and prosody, we propose a TTS system with novel factorized diffusion models to generate natural speech in a zero-shot way. Specifically, 1) we introduce a novel neural speech codec with factorized vector quantization (FVQ), named FACodec, to decompose speech waveform into distinct subspaces of content, prosody, timbre, and acoustic details and reconstruct speech waveform with these disentangled representations, leveraging information bottleneck , various supervised losses, and adversarial training to enhance disentanglement; 2) we propose a factorized diffusion model, which generates the factorized speech representations of duration, content, prosody, and acoustic detail, based on their corresponding prompts. This design allows us to use different prompts to control different attributes. The overview of our method, referred to NaturalSpeech 3, is shown in Figure 1.
We decompose complex speech into subspaces representing different attributes, thus simplifying the modeling of speech representation. This approach offers several advantages: 1) our factorized diffusion model is able to learn these disentangled representations efficiently, resulting in higher quality speech generation; 2) by disentangling timbre information in our FACodec, we enable our factorized diffusion model to avoid directly modeling timbre. This reduces learning complexity and leads to improved zero-shot speech synthesis; 3) we can use different prompts to control different attributes, enhancing the controllability of NaturalSpeech 3.
Benefiting from these designs, NaturalSpeech 3 has achieved significant improvements in speech quality, similarity, prosody, and intelligibility. Specifically, 1) it achieves comparable or better speech quality than the ground-truth speech on the LibriSpeech test set in terms of CMOS; 2) it achieves a new SOTA on the similarity between the synthesized speech and the prompt speech (0.64 0.67 on Sim-O, 3.69 4.01 on SMOS); 3) it shows a significant improvement in prosody compared to other TTS systems with 0.16 average MCD (lower is better), 0.21 SMOS; 4) it achieves a SOTA on intelligibility (1.94 1.81 on WER); 5) it achieves human-level naturalness on multi-speaker datasets (e.g., LibriSpeech), another breakthrough after NaturalSpeechWhile NaturalSpeech 1 achieved human-level quality on the single-speaker LJSpeech dataset, NaturalSpeech 3 achieved human-level quality on the diverse multi-speaker LibriSpeech dataset for the first time.. Furthermore, we demonstrate the scalability of NaturalSpeech 3 by scaling it to 1B parameters and 200K hours of training data. Audio samples can be found in https://speechresearch.github.io/naturalspeech3.
Background
In this section, we discuss the recent progress in TTS including: 1) zero-shot TTS; 2) speech representations in TTS; 3) generation methods in TTS; 4) speech attribute disentanglement.
Zero-shot TTS. Zero-shot TTS aims to synthesize speech for unseen speakers with speech prompts. We can systematically categorize these systems into four groups based on data representation and modelling methods: 1) Discrete Tokens + Autoregressive ; 2) Discrete Tokens + Non-autoregressive ; 3) Continuous Vectors + Autoregressive ; 4) Continuous Vectors + Non-autoregressive . Discrete tokens are typically derived from neural codec, while continuous vectors are generally obtained from mel-spectrogram or latents from audio autoencoder or codec. In addition to the aforementioned perspectives, we disentangle speech waveforms into subspaces based on attribute disentanglement and propose a factorized diffusion model to generate attributes within each subspace, motivated by the principle of divide-and-conquer. Meanwhile, we can reuse previous methods, employing discrete tokens along with autoregressive models.
Speech Representations in TTS. Traditional works propose using prior-based speech representation such as raw waveform or mel-spectrogram . Recently, large-scale TTS systems leverage data-driven representation, i.e., either discrete tokens or continuous vectors form an auto-encoder . However, these methods ignore that speech contains various complex attributes and encounter intricate complexities during speech generation. In this paper, we factorize speech into individual subspaces representing different attributes which can be effectively and efficiently modeled.
Generation Methods in TTS. Previous works have demonstrated that NAR-based models enjoy better robustness and generation speed than AR-based models, because they explicitly model the duration and predict all features simultaneously. Instead, AR-based models have better diversity, prosody, expressiveness, and flexibility than NAR-based models, due to their implicitly duration modeling and token sampling strategy. In this study, we adopt the NAR modeling approach and propose a factorized diffusion model to support our disentangled speech representations and also extend it to AR modeling approaches. This allows NaturalSpeech 3 to achieve better expressiveness while maintaining stability and generation speed.
Speech Attribute Disentanglement. Prior works utilize disentangled representation for speech generation, such as speech content from self-supervised pre-trained models , fundamental frequency, and timbre, but speech quality is not satisfying. Recently, some works explore attribute disentanglement in neural speech codec. SpeechTokenizer uses HuBERT for semantic distillation, aiming to render the first-layer RVQ representation as semantic information. Disen-TF-Codec proposes the disentanglement with content and timbre representation, and applies them for zero-shot voice conversion. In this paper, we achieve better disentanglement with more speech attributes including content, prosody, acoustic details and timbre while ensuring high-quality reconstruction. We validate such disentanglement can bring about significant improvements in zero-shot TTS task.
NaturalSpeech 3
In this section, we present NaturalSpeech 3, a cutting-edge system for natural and zero-shot text-to-speech synthesis with better speech quality, similarity and controllability. As shown in Figure 1, NaturalSpeech 3 consists of 1) a neural speech codec (i.e., FACodec) for attribute disentanglement; 2) a factorized diffusion model which generates factorized speech attributes. Since the speech waveform is complex and intricately encompasses various attributes, we factorize speech into five attributes including: duration, prosody, content, acoustic details, and timbre. Specifically, although the duration can be regarded as an aspect of prosody, we choose to model it explicitly due to our non-autoregressive speech generation design. We use our internal alignment tool to alignment speech and phoneme and obtain phoneme-level duration. For other attributes, we implicitly utilize the factorized neural speech codec to learn disentangled speech attribute subspaces (i.e., content, prosody, acoustic details, and timbre). Then, we use the factorized diffusion model to generate each speech attribute representation. Finally, we employ the codec decoder to reconstruct the waveform with the generated speech attributes. We introduce the FACodec in Section 3.2 and the factorized diffusion model in Section 3.3.
2 FACodec for Attribute Factorization
We propose a factorized neural speech codec (i.e., FACodec We will release the code and pre-trained checkpoint of FACodec soon.) to convert complex speech waveform into disentangled subspaces representing speech attributes of content, prosody, timbre, and acoustic details and reconstruct high-quality speech waveform from these.
As shown in Figure 2, our FACodec consists of a speech encoder, a timbre extractor, three factorized vector quantizers (FVQ) for content, prosody, acoustic detail, and a speech decoder. Given a speech , 1) following , we adopt several convolutional blocks for the speech encoder with a downsample rate of 200 for 16KHz speech data (i.e., each frame corresponding to a 12.5ms speech segment) to obtain pre-quantization latent ; 2) the timbre extractor is a Transformer encoder which converts the output of the speech encoder into a global vector representing the timbre attributes; 3) for other attribute ( for prosody, content, and acoustic detail, respectively), we use a factorized vector quantizer () to capture fine-grained speech attribute representation and obtain corresponding discrete tokens; 4) the speech decoder mirrors the structure of speech encoder but with much larger parameter amount to ensure high-quality speech reconstruction. We first add the representation of prosody, content, and acoustic details together and then fuse the timbre information by conditional layer normalization to obtain the input for the speech decoder. We discuss how to achieve better speech attribute disentanglement in the next section.
2.2 Attribute Disentanglement
Directly factorizing speech into different subspaces does not guarantee the disentanglement of speech. In this section, we introduce some techniques to achieve better speech attribute disentanglement: 1) information bottleneck, 2) supervision, 3) gradient reverse, and 4) detail dropout. Please refer to Appendix B.1 for more training details.
Information Bottleneck. Inspired by , to force the model to remove unnecessary information (such as prosody in content subspace), we construct the information bottleneck in prosody, content, and acoustic details FVQ by projecting the encoder output into a low-dimensional space (i.e., 8-dimension) and subsequently quantize within this low-dimensional space. This technique ensures that each code embedding contains less information, facilitating information disentanglement . After quantization, we will project the quantized vector back to original dimension.
Supervision. To achieve high-quality speech disentanglement, we introduce supervision as auxiliary task for each attribute. For prosody, since pitch is an important part of prosody , we take the post-quantization latent to predict pitch information. We extract the F0 for each frame and use normalized F0 (z-score) as the target. For content, we directly use the phoneme labels as the target (we use our internal alignment tool to get the frame-level phoneme labels). For timbre, we apply speaker classification on by predicting the speaker ID.
Gradient Reversal. Avoiding the information leak (such as the prosody leak in content) can enhance disentanglement. Inspired by , we adopt adversarial classifier with the gradient reversal layer (GRL) to eliminate undesired information in latent space. Specifically, for prosody, we apply phoneme-GRL (i.e., GRL layer by predicting phoneme labels) to eliminate content information; for content, since the pitch is an important aspect of prosody, we apply F0-GRL to reduce the prosody information for simplicity; for acoustic details, we apply both phoneme-GRL and F0-GRL to eliminate both content and prosody information. In addition, we apply speaker-GRL on the sum of to eliminate timbre.
Detail Dropout. We have the following considerations: 1) empirically, we find that the codec tends to preserve undesired information (e.g., content, prosody) in acoustic details subspace since there is no supervision; 2) intuitively, without acoustic details, the decoder should reconstruct speech only with prosody, content and timbre, although in low-quality. Motivated by them, we design the detail dropout by randomly masking out during the training process with probability . With detail dropout, we achieve the trade-off of disentanglement and reconstruction quality: 1) the codec can fully utilize the prosody, content and timbre information to reconstruct the speech to ensure the decouple ability, although in low-quality; 2) we can obtain high-quality speech when the acoustic details are given.
3 Factorized Diffusion Model
We generate speech with discrete diffusion for better generation quality. We have the following considerations: 1) we factorize speech into the following attributes: duration, prosody, content, and acoustic details, and generate them in sequential with specific conditions. Firstly, as we mentioned in Section 3.1, due to our non-autoregressive generation design, we first generate duration. Secondly, intuitively, the acoustic details should be generated at last; 2) following the speech factorization design, we only provide the generative model with the corresponding attribute prompt and apply discrete diffusion in its subspace; 3) to facilitate in-context learning in diffusion model, we utilize the codec to factorize speech prompt into attribute prompts (i.e., content, prosody and acoustic details prompt) and generate the target speech attribute with partial noising mechanism following . For example, for prosody generation, we directly concatenate prosody prompt (without noise) and target sequence (with noise) and gradually remove noise from target sequence with prosody prompt.
With these thoughts, as shown in Figure 3, we present our factorized diffusion model, which consists of a phoneme encoder and speech attribute (i.e., duration, prosody, content, and acoustic details) diffusion modules with the same discrete diffusion formulation: 1) we generate the speech duration by applying duration diffusion with duration prompt and phoneme-level textural condition encoded by phoneme encoder. Then we apply the length regulator to obtain frame-level phoneme condition ; 2) we generate prosody with prosody prompt and phoneme condition ; 3) we generate content prosody with content prompt and use generated prosody and phoneme as conditions; 4) we generate acoustic details with acoustic details prompt and use generated prosody, content and phoneme as conditions. Specifically, we do not explicitly generate the timbre attribute. Due to the factorization design in our FACodec, we can obtain timbre from the prompt directly and do not need to generate it. Finally, we synthesize the target speech by combining attributes and and decoding it with codec decoder. We discuss the diffusion formulation in Section 3.3.2.
3.2 Diffusion Formulation
Forward Process. Denote the target discrete token sequence, where is the sequence length, is the prompt discrete token sequence, and is the condition. The forward process at time is defined as masking a subset of tokens in with the corresponding binary mask , formulated as , by replacing with [MASK] token if , and otherwise leaving unmasked if . and is a monotonically increasing function. In this paper, . Specially, we denote for the original token sequence and for the fully masked sequence.
Reverse Process. The reverse process gradually restores by sampling from reverse distribution , starting from full masked sequence . Since is unavailable in inference, we use the diffusion model , parameterized by , to predict the masked tokens conditioned on and , denoted as . The parameters are optimized to minimize the negative log-likelihood of the masked tokens:
Then we can get the reverse transition distribution:
Inference. During inference, we progressively replace masked tokens, starting from the fully masked sequence , by iteratively sampling from . Inspire by , we first sample from , and then sample from , which involves remask tokens in with the lowest confidence score, where we define the confidence score of in to if , otherwise, we set confidence score of to , which means that tokens already unmasked in will not be remasked.
Classifier-free Guidance. Moreover, we adapt the classifier-free guidance technique . Specifically, in training, we do not use the prompt with a probability of . In inference, we extrapolate the model output towards the conditional generation guided by the prompt and away from the unconditional generation , i.e., , with a guidance scale selected based on experimental results. We then rescale it through , following .
4 Connections to the NaturalSpeech Series
NaturalSpeech 3 is an advanced TTS system of the NaturalSpeech series. Compared with the previous versions NaturalSpeech and NaturalSpeech 2 , NaturalSpeech 3 has the following connections and distinctions:
Goal. The NaturalSpeech series aims to generate natural speech with high quality and diversity. We approach this goal in several stages: 1) Achieving high-quality speech synthesis in single-speaker scenarios. To this end, NaturalSpeech generates speech with quality on par with human recordings and only tackles single-speaker recording-studio datasets (e.g., LJSpeech). 2) Achieving high-quality and diverse speech synthesis on multi-style, multi-speaker, and multi-lingual scenarios. NaturalSpeech 2 firstly focuses on speech diversity by exploring the zero-shot synthesis ability based on large-scale, multi-speaker, and in-the-wild datasets. Furthermore, NaturalSpeech 3 further achieves human-level naturalness on the multi-speaker dataset (e.g., LibriSpeech).
Architecture. The NaturalSpeech series shares the basic components such as encoder/decoder for waveform reconstruction and duration prediction for non-autoregressive speech generation. Different from NaturalSpeech which utilizes flow-based generative models and NaturalSpeech 2 which leverages latent diffusion models, NaturalSpeech 3 proposes the concept of factorized diffusion models to generate each factorized speech attribute in a divide-and-conquer way.
Speech Representations. Due to the complexity of speech waveform, the NaturalSpeech series uses an encoder/decoder to obtain speech latent for high-quality speech synthesis. NaturalSpeech utilizes naive VAE-based continuous representations, NaturalSpeech 2 leverages the continuous representations from the neural audio codec with residual vector quantizers, while NaturalSpeech 3 proposes a novel FACodec to convert complex speech signal into disentangled subspaces (i.e., prosody, content, acoustic details, and timbre) and reduces the speech modeling complexity.
Experiments and Results
In this subsection, we introduce the training, inference and evaluation for the Factorized Diffusion Model. Please refer to Appendix A.1 for model configuration.
Implementation Details. We use Librilight , which contains K hours of KHz unlabeled speech data and around 7000 distinct speakers from LibriVox audiobooks, as the training set. In duration diffusion, we further improve the performance by conditioning phoneme-level prosody codes. Specifically, we perform phoneme-level pooling according to duration on the pre-quantized vectors, and then feed these phoneme-level representations into the prosody quantizer in our codec to obtain the phoneme-level prosody codes. We employ an additional discrete diffusion to generate these in inference. We perform iterations in each diffusion process. We generate duration without classifier-free guidance and generate others with a classifier-free guidance scale of . This strategy results in for phoneme-level prosody, for duration, for each token sequence of prosody, content, and acoustic details, totaling forward passes due to the double computation with classifier-free guidance. Please refer to Appendix B.1 for details of the FACodec and Appendix A.2 for more details of our factorization diffusion model.
Evaluation Dataset. We employ two benchmark datasets: 1) LibriSpeech test-clean, a widely-used testset for zero-shot TTS task. It contains 40 distinct speakers and 5.4-hour speech. Following , we randomly select one sentence for each speaker for LibriSpeech test-clean benchmark. Specifically, we randomly select -second clips as prompts from the same speaker’s speech. 2) RAVDESS , an emotional TTS dataset featuring 24 professional actors ( female, male) across emotions (neutral, calm, happy, sad, angry, fearful, surprise, and disgust) in emotional intensity (normal and strong). We use strong-intensity samples for RAVDESS benchmark. We adopt this benchmark for prosody evaluation, considering 1) for the same speaker, speech with the same emotion shares similar prosody, while speech with different emotions displays varied prosodies; 2) the benchmark provides speech samples with the same text from the same speaker across eight different emotions.
Evaluation Metrics. Objective Metrics: In the Librispeech test-clean benchmark, we evaluate speaker-similarity (SIM-O and SIM-R), speech quality (UTMOS), and robustness (WER). In specific, 1) for SIM-O and SIM-R, we employ the WavLM-TDCNNhttps://github.com/microsoft/UniSpeech/tree/main/downstreams/speaker_verification speaker embedding model to assess speaker similarity between generated samples and the prompt. Results are reported for both similarity to original prompt (SIM-O) and reconstructed prompt (SIM-R); 2) for speech quality, we employ UTMOS which is a surrogate objective metric of MOS; 3) for Word Error Rate (WER), we use an ASR modelhttps://huggingface.co/facebook/hubert-large-ls960-ft to transcribe generated speech. The model is a CTC-based HuBERT pre-trained on Librilight and fine-tuned on the 960 hours training set of LibriSpeech. We also use an advanced ASR model based on transducer https://huggingface.co/nvidia/stt_en_conformer_transducer_xlarge. In the RAVDESS benchmark, we evaluate the prosody similarity (MCD and MCD-Acc). In specific, 1) following , we adopt Mel-Ceptral Distortion (MCD) for prosody evaluation by measuring the differences between generated samples and ground truth samples. We report the results for eight emotions, along with the average result. 2) for MCD-Acc, we evaluate the top-1 emotion accuracy of the generated speech on the RAVDESS benchmark for prosodic similarity measures. Specifically, we adopt a K-Nearest-Neighbors (KNN) model as emotion classifier. We compare MCD distances between the generated speech and the ground-truth speech from the same speaker, across eight different emotions. Subjective Metrics: We employ comparative mean option score (CMOS) and similarity mean option score (SMOS) in both two benchmarks to evaluate naturalness and similarity, respectively.
Evaluation Baselines. We compare NaturalSpeech 3 with baselines: 1) VALL-E . 2) NaturalSpeech 2 . 3) Voicebox . 4) Mega-TTS 2 . 5) UniAudio . 6) StyleTTS 2 . 7) HierSpeech++ . Please refer to Appendix A.3 for details.
2 Experimental Results on Zero-shot TTS
In this subsection, we compare NaturalSpeech 3 with baselines in terms of: 1) generation quality in Section 4.2.1; 2) generation similarity in Section 4.2.2; 3) robustness in Section 4.2.3. Specifically, for generation similarity, we evaluate in two aspects: 1) speaker similarity; 2) prosody similarity. Please refer to Appendix A.5 for latency analysis.
To evaluate speech quality, we conduct CMOS test, with native as the judges. We randomly select utterances from both LibriSpeech test-clean and RAVDESS benchmarks. As shown in Table 1, we find that 1) NaturalSpeech 3 is close to the ground-truth recording ( on Librispeech test-clean, and on RAVDESS), which demonstrates NaturalSpeech 3 can generate high-quality and natural speech; 2) NaturalSpeech 3 outperforms baselines by a substantial margin, verifying the effectiveness of NaturalSpeech 3 with factorization.
2.2 Generation Similarity
Speaker Similarity. We evaluate the speech similarity with both objective metrics (Sim-O and Sim-R) and subjective metrics (SMOS), with natives as the judges. We randomly select utterances for SMOS test. As shown in Table 1, we find that 1) NaturalSpeech 3 achieves parity in Sim-O and a increase in SMOS with ground truth, which indicates great speaker similarity achieved by our proposed method; 2) NaturalSpeech 3 outperforms all baselines on both objective and subjective metrics, highlighting the superiority of our method with factorization in terms of speaker similarity. Additionally, we notice certain discrepancy between Sim-O and SMOS. For instance, the SMOS is not as competitive as SIM-O for Voicebox model, likely due to some unnatural prosody.
Prosody Similarity. We evaluate prosody similarity with both objective metrics (MCD and MCD-Acc) and subjective metrics (SMOS) on the RAVDESS benchmark. We randomly select utterances for SMOS test. As shown in Table 2, NaturalSpeech 3 consistently surpasses baselines by a remarkable margin in MCD avg, MCD-Acc, and SMOS. It reveals that NaturalSpeech 3 achieves a significant improvement in terms of prosodic similarity. Please refer to Appendix A.7 for the MCD scores across emotions.
2.3 Robustness
We assess the robustness of our zero-shot TTS by measuring the word error rate of generated speech on the LibriSpeech test-clean benchmark. The results in Table 1 indicate that 1) NaturalSpeech 3 achieves a better WER than the ground truth, proving the high intelligibility; 2) NaturalSpeech 3 outperforms other baselines by a considerable margin, which demonstrates the superior robustness of NaturalSpeech 3.
2.4 Human-Level Naturalness on LibriSpeech Testset
We compare the speech synthesized by NaturalSpeech 3 with human recordings (Ground Truth) in Table 1 (more results can be found in Table 9 in Appendix A.4). We have the following observations: 1) NaturalSpeech 3 achieves -0.01 Sim-O and +0.16 SMOS compared to human recordings, which demonstrates that our method is on par or better on speaker similarity; 2) NaturalSpeech 3 achieves -0.08 CMOS and +0.16 UTMOS compared with recording, which demonstrates that our method can generate on-par or better voice quality; 3) Our method also achieves close WER with human recordings, which demonstrates the robustness of NaturalSpeech 3. Therefore, we can conclude that for the first time, NaturalSpeech 3 has achieved human-level quality and naturalness on the multi-speaker LibriSpeech test set in a zero-shot way. It is another great milestone after NaturalSpeech 1 has achieved human-level quality on the single-speaker LJSpeech dataset.
3 Ablation Study and Method Analyses
In this subsection, we conduct ablation studies to verify the effectiveness of 1) factorization; 2) classier-free guidance; 3) prosody representation. We also conduct ablation study to compare our duration diffusion model with traditional duration predictor in Appendix A.6.
Factorization. To verify the proposed factorization method, we ablate it by removing factorization in both codec and factorized diffusion model. Specifically, we 1) use the discrete tokens from SoundStream, a neural codec which does not consider factorization, and 2) do not consider factorization in generation. As shown in Table 3, we could find a significant performance degradation without the factorization, a drop of in Sim-O, in Sim-R, in WER, in CMOS and 0.42 in SMOS. This indicates the proposed factorized method can consistently improve the performance in terms of speaker similarity, robustness, and quality.
Classier-Free Guidance. We conduct an ablation study by dropping the classifier-free guidance in inference to validate its effectiveness. We double the iterations to ensure the same forward passes for fair comparison. Table 3 illustrates a significant degradation without classifier-free guidance, a decrease of in Sim-O, in Sim-R, in CMOS and in SMOS, proving that classifier-free guidance can greatly help the speaker similarity and quality.
Prosody Representation. We compare different prosody representations on zero-shot TTS task. In specific, we select handcrafted prosody features (e.g., the first 20 bins of mel-spectrogram ) as the baseline. We drop the prosody FVQ module and directly quantize the first 20 bins of the mel-spectrogram, without the normalized F0 loss. Table 4 shows that using “Mel 20 Bins” as prosody representation demonstrates inferiority in terms of prosody similarity compared to the prosody representations learned from codec (4.34 vs 4.28 in average MCD, 0.46 vs 0.52 in MCD-Acc).
3.2 Method Analyses
In this subsection, we first discuss the extensibility of our factorization. We then introduce the application of speech attributes manipulation in a zero-shot way.
Extensibility. NaturalSpeech 3 utilizes a non-autoregressive model for discrete token generation with factorization design. To validate the extensibility of our proposed factorization method, we further explore the autoregressive generative model for discrete token generation under our factorization framework. We utilize VALL-E for verification. We first employ an autoregressive language model to generate prosody codes, followed by a non-autoregressive model to generate the remaining content and acoustic details codes. This approach maintains a consistent order of attribute generation, allowing for a fair comparison. We name it VALL-E + FACodec. As shown in Table 6, VALL-E + FACodec consistently outperforms VALL-E by a considerable margin in all objective and subjective metrics, demonstrating the factorization design can enhance VALL-E in speech similarity, quality and generation robustness. It further shows our factorization paradigm is not limited in the proposed factorization diffusion model and has a large potential in other generative models. We leave it for future work.
Speech Attribute Manipulation. As discussed in Section 3.3, our factorized diffusion model enables attribute manipulation by selecting different attributes prompts from different speech. We mainly focus on manipulating duration, prosody, and timbre, since the content codes are dictated by the text in TTS, and the acoustic details do not carry semantic information. Leveraging the strong in-context capability of NaturalSpeech 3, the generated speech effectively mirrors the corresponding speech attributes. For instance, 1) we can utilize the timbre prompt from a different speech to control the timbre while keeping other attributes unchanged; 2) despite the correlation between duration and prosody, we can still solely adjust duration prompt to regulate the speed; 3) moreover, we can combine different speech attributes from disparate samples as desired. This allow us to mimic the timbre while using different prosody and speech speed. Samples are available on our demo pagehttps://speechresearch.github.io/naturalspeech3.
3.3 Experimental Results on FACodec
We compare the proposed FACodec in terms of the reconstruction quality with strong baselines, such as EnCodec , HiFi-Codec , Descript-Audio-Codec (DAC) , and our reproduced SoundStream . Table 5 shows that our codec significantly surpasses SoundStream in the same bandwidth setting ( in PESQ, in STOI, in MSTFT and in MCD, respectively). Check more details in Appendix B.2. Compared with other baselines, FACodec also get comparable performance. Additionally, since our codec decouples timbre information, it can enable zero-shot voice conversion easily, we provide the details and experiment results in Appendix B.3. Appendix B.4 shows some ablation studies about our FACodec.
4 Effectiveness of Data and Model Scaling
In this section, we study the effectiveness of data and model scaling on the proposed factorized diffusion model. We use the same FACodec trained on LibriLight dataset for fair comparison. We evaluate the zero-shot TTS performance in terms of speaker similarity (Sim-O) and robustness (WER) on an internal test set consisting of audio clips.
Data Scaling. With a fixed model size of 500M parameters, we trained the factorized diffusion model on three datasets: 1) a 1K-hour subset randomly drawn from the Librilight dataset, 2) a 60K-hour Librilight dataset, and 3) an internal dataset with 200K hours of speech. In Table 7, we observe that: 1) even with a mere 1K hours of speech data, our model attains a Sim-O score of and a WER of . It shows that with the speech factorization, NaturalSpeech 3 can generate the speech effectively. 2) As we scale up training data from 1K hours to 60K hours, and then to 200K hours, NaturalSpeech 3 displays continuously enhanced performance, with an improvement of and in terms of Sim-O, and and in terms of WER, respectively, thus confirming the benefits of data scaling. Note that our method trained on 200K hours is still underfitting and longer training will result in better performance.
Model Scaling. We scale up the model size from 500M to 1B parameters with the internal 200K hours dataset. Specifically, we double the number of transformer layers from to . The results in Table 8 show a boost in both speaker similarity ( in Sim-O) and robustness ( in WER), validating the effectiveness of model scaling. In the future, we will scale up the model size even larger to achieve better results.
Conclusion
In this paper, we develop a TTS system that consists of 1) a novel neural speech codec with factorized vector quantization (i.e., FACodec) to decompose speech waveform into distinct subspaces of content, prosody, acoustic details and timbre and 2) novel factorized diffusion model to synthesize speech by generating attributes in subspaces with discrete diffusion. NaturalSpeech 3 outperforms the state-of-the-art TTS system on speech quality, similarity, prosody, and intelligibility. We also show that NaturalSpeech 3 can enable speech attribute manipulation, by customizing speech attribute prompts. Furthermore, we demonstrate that NaturalSpeech 3 achieves human-level performance on the multi-speaker LibriSpeech dataset for the first time and better performance by scaling to 1B parameters and 200K hours of training data. We list the limitations and future works in Appendix C.
Boarder Impact
Since our model could synthesize speech with great speaker similarity, it may carry potential risks in misuse of the model, such as spoofing voice identification or impersonating a specific speaker. We conducted the experiments under the assumption that the user agree to be the target speaker in speech synthesis. To prevent misuse, it is crucial to develop a robust synthesized speech detection model and establish a system for individuals to report any suspected misuse.
References
Appendix A Details of Factorization Diffusion Model
The phoneme encoder uses a similar architecture as and comprises a -layer Transformer with attention heads, embedding dimensions, filter size and kernel size for 1D convolution, and a dropout of . In prosody, content and acoustic details diffusion, we adopt a shared -layer Transformer, with attention heads, embedding dimensions, filter size and kernel size for 1D convolution, and a dropout of . We additionally use conditional layer normalization in each Transformer block to support diffusion time input. In phoneme-level prosody and duration diffusion, we adopt a -layer Transformer with attention heads, embedding dimensions, filter size and kernel size for 1D convolution, and a dropout of . We also use conditional layer normalization in the model to support diffusion time input.
A.2 Training and Inference Details
We use Librilight , which contains K hours of KHz unlabeled speech data and around 7000 distinct speakers from LibriVox audiobooks, as the training set. We transcribe using an internal ASR system, convert transcriptions to phonemes via grapheme-to-phoneme conversion , and obtain duration with an internal alignment tool. We use A100 80GB GPUs with a batch size of K frames of latent vectors per GPU for M steps. We use the AdamW optimizer with a learning rate of , , and , K warmup steps following the inverse square root learning schedule.
During inference, we perform iterations in each diffusion process, including phoneme-level prosody, duration, prosody, content and acoustic details diffusion. We generate duration without classifier-free guidance, and generate others with a classifier-free guidance scale of . This strategy results a for phoneme-level prosody, for duration, for each token sequence of prosody, content and acoustic details, totaling forward passes due to the double computation with classifier-free guidance. We use a top-k of , with sampling temperature annealing from to . Following , Gumbel noises are added to token confidences when determining which positions to re-mask in , mentioned in Section 3.3.2.
A.3 Evaluation Baselines
We compare NaturalSpeech 3 with following strong zero-shot TTS baselines:
VALL-E . It use an autoregressive and an additional non-autoregressive model for discrete token generation. We report the scores directly obtained from the paper. We additionally reproduce it using discrete tokens from SoundStream on Librilight.
NaturalSpeech 2 . It use a non-autoregressive model for continuous vectors generation. We obtain samples through communication with the authors.
Voicebox . It use a non-autoregressive model for continuous vectors generation. We obtain samples through communication with the authors. We additionally reproduce it using mel-spectrogram on Librilight.
Mega-TTS 2 . It use a non-autoregressive model for continuous vectors generation. We obtain samples through communication with the authors.
UniAudio . It use an autoregressive model for discrete token generation. We obtain samples through communication with the authors.
StyleTTS 2 . It use a non-autoregressive model for continuous vectors generation. We use official code and checkpointhttps://github.com/yl4579/StyleTTS2.
HierSpeech++ . It use a non-autoregressive model for continuous vectors generation. We use official code and checkpointhttps://github.com/sh-lee-prml/HierSpeechpp. We do not use its super resolution model for fair comparison.
A.4 More Experimental Results on Zero-shot TTS
In this section, we report more evaluation results for NaturalSpeech 3 and other baselines on: 1) WER, inferred by an advanced ASR systemhttps://huggingface.co/nvidia/stt_en_conformer_transducer_xlarge; 2) UTMOS , which is a surrogate objective metric of MOS. The results are shown in Table 9.
A.5 Latency Analysis
In this subsection, we compare the inference latency of NaturalSpeech 3 with an autoregressive method (VALL-E) and a non-autoregressive method (NaturalSpeech 2). We also investigate the effect of reducing the number of iterations in each diffusion from 4 to 1, resulting in a total of 15 forward passes. We call this variant NaturalSpeech 3 one-step. We evaluate the performance on Librispeech test-clean in terms of speaker similarity (Sim-O/Sim-R) and quality (UTMOS https://github.com/tarepan/SpeechMOS, a surrogate objective metric of CMOS). The latency tests are conducted on a server with E5-2690 Intel Xeon CPU, 512GB memory, and one NVIDIA V100 GPU. The results are shown in Table 10. From the results, we have several observations. 1) NaturalSpeech 3 achieves a speedup over VALL-E and speedup over NaturalSpeech 2, while consistently surpasses these baselines on all metrics. This demonstrate NaturalSpeech 3 is both effective and efficient. 2) when using fewer diffusion steps, NaturalSpeech 3 can still maintain robust performance ( in Sim-O, in Sim-R, and in UTMOS) with a faster speed, proving the robustness of diffusion steps.
A.6 Ablation Study on Duration Diffusion Model
In this subsection, we conduct an ablation study to compare our duration discrete diffusion model with the traditional duration predictor, which regresses the duration in logarithmic domain. The ablation study focus on 1) Generation: multi-step generation vs. one-step generation. 2) Objective: classification-based cross-entropy loss vs. regression-based L2 loss. 3) Conditioning: with vs. without phoneme-level prosody conditioning. 4) Prompting: with vs. without duration prompting. We evaluate them on Librispeech test-clean in terms of speaker similarity (Sim-O/Sim-R), robustness (WER) and qualtiy (UTMOS). As shown in Table 11, we can find that 1) without multi-step generation, there’s a significant drop in performance (-0.05 in Sim-O, -0.03 in Sim-R, and -0.12 in UTMOS). 2) replacing cross-entropy loss with l2 loss affects the performance, causing a decrease of -0.05 in Sim-O, -0.04 in Sim-R, 0.44 in WER and -0.17 in UTMOS. 3) dropping phoneme-level prosody conditioning will affect both speaker similarity (-0.05 in Sim-O and -0.04 in Sim-R), robustness (0.55 in WER) and quality (-0.19 in UTMOS) 4) the duration prompting mechanism is crucial for speaker similarity, robustness and quality, with changes of -0.06 in Sim-O, -0.05 in Sim-R, 0.89 in WER and -0.22 in UTMOS. These results confirm that each design aspect of our duration predictor contributes to performance improvement.
A.7 Details of Prosody Similarity Evaluation
In Table 12, we present MCD on different emotions, comparing NaturalSpeech 3 with the baseline methods on the RAVDESS benchmark. NaturalSpeech 3 demonstrates robust performance across emotions, verifying the effectiveness and robustness in terms of prosody similarity.
Appendix B Details of FACodec
Model Architecture. The basic architecture of FACodec encoder and decoder follows and employs the SnakeBeta activation function . The timbre extractor consists of several conformer blocks. We use as the number of quantizers for each of the three FVQ , the codebook size for all the quantizers is 1024.
Loss Functions. We utilize the multi-scale mel-reconstruction loss as detailed in . For the adversarial loss , we employ both the multi-period discriminator (MPD) and the multi-band multi-scale STFT discriminator, as proposed by . Additionally, we incorporate the relative feature matching loss . For codebook learning, we use the codebook loss and the commitment loss from VQ-VAE . The training loss also includes the phone prediction loss , the normalized F0 prediction loss , and the gradient reverse losses of phone prediction , normalized F0 prediction , and speaker classification for disentanglement learning. The total training loss for the generator can be formulated as: where , , , , ,, , , and are coefficients for balancing each loss terms. In our paper, we set these coefficients as follows: , , , , , , , , and .
Training Details. We use Librilight as the training set. We train the codec using 8 NVIDIA TESLA V100 32GB GPUs with a batch size of 32 speech clips of 16000 frames each per GPU for K steps. We use the Adam optimizer with a learning rate of , , and .
B.2 Reconstruction Performance Comparison
We evaluate the reconstruction performance with the following objective metrics: Perceptual Evaluation of Speech Quality (PESQ), Short-Time Objective Intelligibility (STOI), Multi-Resolution STFT Distance (MSTFT), and Mel-Cepstral Distortion (MCD). These metrics collectively measure the difference between the original and the reconstructed samples. We select the following open-source codec models as baselines: EnCodec https://github.com/facebookresearch/encodec, HiFi-Codec https://github.com/yangdongchao/AcademiCodec, and Descript-Audio-Codec (DAC) https://github.com/descriptinc/descript-audio-codec. We additionally reproduce SoundStream following the original paper’s implementation and experimental setup. Table 13 shows that 1) FACodec significantly surpasses SoundStream in the same bandwidth setting ( in PESQ, in STOI, in MSTFT and in MCD, respectively). Moreover, FACodec achieves on-par performance with SoundStream even when its bandwidth is doubled ( in PESQ, in STOI, in MSTFT and in MCD, respectively). 2) For a fair comparison, we compare FACodec with other baselines in a similar bandwidth. FACodec achieve comparable or better result on most metrics than these strong baselines, which means that we can still achieve excellent reconstruction speech quality when disentangling speech attributes.
B.3 Zero-shot Voice Conversion
Voice conversion aims to transform speech from a source speaker into that of a target speaker, preserving content while altering timbre. Zero-shot voice conversion achieves this by utilizing a prompt speech sample from the target speaker to convert the source speaker’s speech. FACodec achieves zero-shot voice conversion by extracting the speaker embedding from the prompt speech to replace the speaker embedding from the source speech, and utilizing content codes , prosody codes , and detail codes from the source speaker to reconstruct the target speech . We compare FACodec with some previous SOTA models: YourTTS , Make-A-Voice (VC) , LM-VC , and UniAudio . We use VCTK dataset for comparison. We use Sim-Ohttps://huggingface.co/microsoft/wavlm-base-plus-sv to compare speaker similarity to baselines and WER to evaluate speech quality. Table 14 shows the evaluation results. The experimental results demonstrate that FACodec solely achieves comparable similarity and superior intelligence compared to the state-of-the-art zero-shot VC models, which need additional training on this task. This implies that FACodec achieves superior disentanglement, especially in timbre.
B.4 Ablation Study
In this subsection, we study 1) the impact of the information bottleneck on the disentanglement of our FACodec; 2) the effect of gradient reversal on the disentanglement of our FACodec; 3) the role of the acoustic details quantizers; 4) the effects of different prosody representations for TTS generation.
Information Bottleneck for Disentanglement.
We investigate the impact of the information bottleneck on speech disentanglement through qualitative analysis. We find that without using information bottleneck (quantize in original dimensional space rather than low dimensional space) can lead to incomplete disentanglement. For example, we conduct zero-shot voice conversion in the same experimental setting using the FACodec without information bottleneck, as mentioned in Appendix B.3. We observe that the timbre of the converted speech is the interpolation between the source and target, indicating its poor timbre disentanglement. Table 15 demonstrates that without the information bottleneck, the speaker similarity of zero-shot voice conversion decreases by 0.13.
We investigate the impact of gradient reversal on the disentanglement of the FACodec through qualitative analysis. We observe that not using gradient reversal diminishes the disentangling ability of FACodec. For instance, removing the content and prosody gradient reversal from the acoustic detail module results in some content and prosody information leaking into the detail acoustic. We can confirm this by solely reconstructing the speech using detail codes and timbre embedding, where partial content and pitch variations can be heard.
Although content, prosody, and timbre information already encompass the majority of speech information, Table 16 demonstrates that employing acoustic details quantizers enhances the speech reconstruction quality of FACodec. We find 1) without using acoustic details quantizers (only utilizing three codebooks), FACodec achieves comparable or better results compared to SoundStream with using three codebooks, which means that content codes, prosody codes, and timbre embedding already contain most of the necessary information for speech reconstruction; 2) adding acoustic details achieves better reconstruction quality, which suggests that acoustic details codes primarily serve to supplement high-frequency details.
Appendix C Limitation and Future Works
Despite our proposed TTS system has achieved great progress, we still have the following limitations:
Attribute Coverage. In this work, we propose the factorization design for speech representation and generation, and have achieved significant improvement by factorizing speech into content, prosody, duration, acoustic details and timbre. However, these attributes can not coverage all speech aspects. For example, we can not extract the background sounds, which is a common challenge for speech disentanglement. In the future, we will explore more attributes including: 1. energy, 2. background sounds, and etc.
Data Coverage. Although we have achieved remarkable improvement on zero-shot speech synthesis on speech quality, similarity and robustness, NaturalSpeech 3 is trained on English corpus from LibriVox audiobooks. Thus, it can not coverage real word people’s diverse voice and can not support multilingual TTS. In the future, we will address this limitation by collecting more speech data with larger diversity.
Neural Speech Codec. Although our FACodec can factorize speech into attributes and reconstruct with high quality, it still has the following limitations: 1) we need phoneme transcription for content supervision, which limits the scalability; 2) we only verified the disentanglement in zero-shot TTS task. In the future, firstly, we will explore more general methods to achieve better disentanglement, especially without supervision. Secondly, we would like to explore more tasks with the FACodec, such as zero-shot voice conversion and automatic speech recognition.