SpeechX: Neural Codec Language Model as a Versatile Speech Transformer

Xiaofei Wang, Manthan Thakker, Zhuo Chen, Naoyuki Kanda, Sefik Emre Eskimez, Sanyuan Chen, Min Tang, Shujie Liu, Jinyu Li, Takuya Yoshioka

I Introduction

The technology of generative models has undergone rapid and transformative advancements in various machine learning applications, encompassing text , vision , and audio . These advancements have had significant implications for both the industry and society at large. Notably, generative models using multi-modal input have emerged as a remarkable innovation .

In the speech domain, one prominent speech generation task that leverages audio-text input is zero-shot text-to-speech (TTS). Zero-shot TTS involves converting a given text into speech with the voice characteristics and speaking style of a desired talker by using only a brief audio sample of that person. Early studies in zero-shot TTS employed fixed-dimensional speaker embeddings . This approach limited their usage to TTS alone and did not adequately support speaker cloning capabilities.

In contrast, recent approaches have embraced more generic formulations, such as masked speech prediction or neural codec language modeling . These novel approaches directly utilize the target speaker’s audio without compressing it into a fixed-dimensional representation. Consequently, these models have not only achieved remarkable zero-shot TTS performance but also demonstrated additional capabilities, including voice conversion and speech editing . This enhanced flexibility holds tremendous promise for unlocking new possibilities in speech generation models.

However, despite their impressive achievements, these recent generative models still have certain limitations, particularly when it comes to addressing various audio-text-based speech generation tasks involving transforming input speech. For instance, existing speech editing models are restricted to handling clean signals only, lacking the ability to modify spoken content while preserving background sounds. Additionally, to perform denoising, the model discusssed in necessitates the noisy signal to be surrounded by clean speech segments, imposing significant constraints on its practical applications. In the context of transforming non-clean speech, another particularly useful task is target speaker extraction . Target speaker extraction involves extracting the voice of a desired speaker from a speech mixture containing multiple talkers. The desired speaker can be specified using a short voice recording of that individual. Despite its potential significance as discussed in , this task remains unaddressed by existing generative speech models.

It is noteworthy that traditional approaches to speech enhancement tasks, such as denoising and target speaker extraction, have relied on regression models for faithful signal recovery. However, these prior methods typically required distinct expert models for each task, which is not ideal, given the potential diversity of acoustic disturbances . Furthermore, there has been a lack of comprehensive audio-text-based speech enhancement models that leverage reference transcriptions to generate intelligible speech, except for limited studies focusing only on particular speech enhancement tasks .

Given the aforementioned considerations and the successful precedents in other domains, the creation of audio-text-based generative speech models unifying generation and transformation capabilities assumes crucial research importance. These models should possess an overarching capability to tackle a diverse array of speech generation tasks. We propose that such models should be equipped with the following key properties:

Versatility: Similar to unified or foundation models developed in other machine learning domains, the unified audio-text-based generative speech models must handle a wide range of tasks involving speech generation from audio and text inputs. These tasks should encompass not only zero-shot TTS but also various forms of speech transformation, including speech enhancement and speech editing, to name a few.

Robustness: It is essential for the unified models to exhibit robustness to various acoustic distortions since they are likely to be applied in acoustically challenging environments. By ensuring reliable performance, these models can be deemed highly usable in real-world scenarios where background sounds are prevalent.

Extensibility: The unified models must employ flexible architectures, allowing for seamless extensions of task support. One approach to achieving this involves accommodating additional elements, such as input tokens or extra modules. Such flexibility will empower the models to adapt to future speech generation tasks efficiently.

In pursuit of this objective, this paper introduces a versatile speech generation model capable of performing multiple tasks, including zero-shot TTS, noise suppression using an optional transcript input, speech removal, target speaker extraction using an optional transcript input, and speech editing for both quiet and noisy acoustic environments (Fig. 1). We refer to our proposed model as SpeechXX stands for transformation to highlight that our model performs various speech transformation tasks in addition to zero-shot TTS.. As with VALL-E, SpeechX adopts a language modeling approach that generates codes of a neural codec model, or acoustic tokens, based on textual and acoustic inputs. To enable the handling of diverse tasks, we incorporate additional tokens in a multi-task learning setup, where the tokens collectively specify the task to be executed. Experimental results, using 60K hours of speech data from LibriLight as a training set, demonstrate the efficacy of SpeechX, showcasing comparable or superior performance compared to expert models in all the aforementioned tasks. Notably, SpeechX also exhibits novel or expanded capabilities, such as preserving background sounds during speech editing and leveraging reference transcriptions for noise suppression and target speaker extraction. Audio samples showcasing the capabilities of our proposed SpeechX model are available at https://aka.ms/speechx.

II Related Work

Generative models based on a language modeling approach using autoregressive Transformers, also known as decoder-only Transformers, have garnered significant success in various application domains. Notable examples of such models include the GPT series and DALL-E . The autoregressive approach has also been extended to the audio and speech domains. AudioLM and MusicLM are pioneering efforts that exploit multiple types of tokens, each with a distinct time scale and degree of semantic granularity, allowing for hierarchical token generation. This hierarchical structure, comprising both coarse and fine-grained tokens, enables the synthesis of sounds with both nuanced details and long-term regularities.

For zero-shot TTS, VALL-E and SPEAR-TTS employ the autoregressive Transformers by representing textual (semantic) and acoustic tokens as a single data stream. This approach enables the models to perform zero-shot speaker adaptation, facilitating the generation of TTS voices that mimic a specific person’s voice. It was demonstrated that these models could perform zero-shot TTS from speech clips as short as three seconds. A notable advantage of these autoregressive speech generation models is their ability to perform TTS without requiring a separate duration model. This streamlined architecture simplifies the training process and potentially offers increased flexibility needed to subsume various speech generation tasks. For this reason, we opt to build our SpeechX models by using autoregressive Transformers.

II-B Multi-task generative speech models

Several papers have recently reported efforts in developing audio-text-based speech generation models that support zero-shot TTS and several related tasks. These tasks include voice or style conversion (Make-A-Voice , NaturalSpeech2 , and Voicebox ), speech editing (Mega-TTS and Voicebox), and denoising (NaturalSpeech2 and Voicebox). Voicebox has showcased noteworthy advancements by facilitating a multitude of tasks through its masked speech prediction principle. Nevertheless, its capabilities are still limited to clean speech generation alone, falling short of effectively dealing with noisy speech or encompassing conventional audio enhancement tasks such as noise suppression and target speaker extraction.

In this study, we deal with both clean and noisy speech and unify the generation and transformation tasks. To accomplish this, we extend VALL-E by performing multi-task learning with task-dependent prompts. The resulting model, which we call SpeechX, exhibits versatility in various speech processing tasks. The model excels not only in speech generation tasks like zero-shot TTS and speech editing but also performs effectively in enhancement tasks such as noise suppression and target speaker extraction. It also realizes novel capabilities, such as editing spoken content while retaining the background noise or effectively leveraging transcriptions for enhancement tasks.

III Method

Fig. 1 illustrates an overview of the SpeechX architecture. Building upon the principles introduced in VALL-E, SpeechX employs a neural codec language model based on Transformers. The model learns to perform conditional generation of a neural code sequence, denoted as O\mathcal{O}, based on two input prompts: textual prompt T\mathcal{T} and acoustic prompt A\mathcal{A}. The neural codes may also be referred to as acoustic tokens.

The textual prompt T\mathcal{T} is a sequence of phonemes obtained by applying grapheme-to-phoneme conversionhttps://github.com/Kyubyong/g2p to an input text. The textual prompt conveys the semantic information, and thus it is called semantic tokens. Conversely, the acoustic prompt A\mathcal{A} encapsulates the acoustic information of an input speech signal. It is obtained by converting the input audio into a sequence of acoustic tokens with an encoder of the neural codec model. Furthermore, to specify the task to be executed, or equivalently the desired output, we incorporate additional tokens in the acoustic prompt. The details will be explained in Section III-C. The output O\mathcal{O} is a sequence of neural codes of the desired signal, which is then translated into a waveform signal with the codec decoder.

We use EnCodec as the neural codec model, following the prior work. Encodec is based on an encoder-decoder architecture with LL quantization layers. In our experiments, we use L=8L=8 to be consistent with the configuration of . Each layer of the EnCodec model produces discrete codes consisting of 1024 entries at a sampling rate of 75 Hz.

We emphasize that the proposed simple architecture capitalizes on the end-to-end modeling capability of the neural language modeling approach. In contrast to other zero-shot TTS or speech generation methods, this approach eliminates the need for a separate model, such as a speaker embedding model or a duration model, apart from the neural codec model. This key property allows SpeechX to acquire knowledge of diverse tasks with varying requirements and input-output relationships, thereby facilitating a versatile and highly extensible speech generation process.

III-B Neural codec language model

As with VALL-E , SpeechX makes use of auto-regressive (AR) and non-auto-regressive (NAR) Transformer models. Specifically, the AR model is used to output the neural codes corresponding to the first quantization layer of EnCodec. On the other hand, the NAR model generates the neural codes of all the layers above the first layer, namely the second through eighth layers. Combining the AR and NAR models provides a reasonable trade-off between generation flexibility and inference speed, as discussed in .

where o<t,1=[o1,1,⋯ ,ot−1,1]\mathbf{o}_{<t,1}=[o_{1,1},\cdots,o_{t-1,1}], while θAR\theta_{\textit{AR}} represents the AR Transformer model parameters. Different embedding projections are applied to the textual and acoustic tokens, and they are superimposed by sinusoidal positional embeddings. Note that the AR model in SpeechX is conditioned on the concatenated embeddings of both the acoustic and textual prompts. This formulation differs from that of VALL-E, where the AR model is only conditioned on the textual prompt and the past acoustic history.

After obtaining the first layer codes with the AR model, the NAR model is used to generate the llth layer codes based on the text and acoustic prompts as well as the output codes for the first l−1l-1 layers, which have already been produced. The model is used repeatedly for l=2,⋯ ,8l=2,\cdots,8. Since we use the same NAR model for the remaining seven layers, the NAR model is trained to minimize the following negative log-likelihood function:

where θNAR\theta_{\textit{NAR}} represents the NAR model parameters, while o:,l\bm{o}_{:,l} denotes the entire sequence of ot,lo_{t,l} for the llth layer, and o:,<l=[o:,1,⋯ ,o:,l−1]\bm{o}_{:,<l}=[\bm{o}_{:,1},\cdots,\bm{o}_{:,l-1}]. In this formulation, in order for the single NAR model to process each of the seven layers, the acoustic tokens from the first to (l−1)(l-1)th layers, o:,<l\textbf{o}_{:,<l}, are embedded and summed up.

III-C Task-based prompting

SpeechX aims to handle multiple tasks with one model. To this end, we adopt task-based prompting, as illustrated in Table I and explained in detail below.

In practical speech editing scenarios, the input text is often obtained by first applying automatic speech recognition (ASR) to the input speech and then having a user edit the transcription. In such situations, it is simple to identify the positions at which and must be inserted. Also, it is noteworthy that, in clean speech editing, the use of allows the model to adaptively change the output speech length in such a way that the output speech sounds natural in terms of speaking speed.

The outlined task-based prompting strategy equips the SpeechX model with the ability to uniquely decide the desired output during inference. This approach enables flexibility for incorporating additional tasks. Adding new tasks entails integrating corresponding prompting schemes and continuing model training from an existing checkpoint, where only embeddings for newly introduced task-specific tokens are randomly initialized. This can be performed without changing the underlying model architecture.

III-D Model training

During training, we randomly sample the task for each model update at an equal probability. This is intended to ensure the model does not unduly favor any particular tasks. For noise suppression, speech removal, and target speaker extraction tasks, we include the textual prompt at a 50% probability so that the model equally experiences both text and text-less scenarios.

To help model to acquire basic generation capabilities, we first train the model only for zero-shot TTS and then continue the training process using all the tasks to perform multi-task learning. In other words, we initialize the model with an existing VALL-E model checkpoint. Precisely speaking, the SpeechX model trained solely for zero-shot TTS exhibits slight divergence from VALL-E. This difference arises from the fact that the former explicitly incorporates a distinct enrollment audio, originating from the same speaker, for each training sample, while the latter does not. Nevertheless, for the sake of simplicity, we refer to this initialization approach as VALL-E initialization. When starting the multi-task training stage, randomly initialized embeddings are appended for the special tokens related to the task-dependent prompts. This two-stage training strategy substantially enhances performance across all tasks, as evidenced by our experimental results.

IV Evaluation Setups

Evaluating versatile speech generation models like SpeechX requires performing an array of tests, each focusing on individual tasks. To keep the experiments manageable as well as ensure consistency across the tasks, we used evaluation datasets that were derived from the test-clean split of LibriSpeech for all evaluations. In this section, we provide the details of our evaluation setups. Following previously established practices , we selected the test samples with durations between 4 and 10 seconds.

Zero-shot TTS: For each test sample, we used the reference transcription to create the textual prompt. The acoustic prompt was generated by randomly choosing another utterance of the same speaker and extracting a 3-second-long clip.

Noise suppression: We mixed each test sample with a noise sample randomly picked from the MUSAN dataset at a signal-to-noise ratio (SNR) which was randomly determined from the range between 0 dB and 20 dB. The task was to recover the uncorrupted speech from the noisy speech. The acoustic prompt was obtained by applying EnCodec to the noisy signal. As regards the textual prompt, we considered both text-less (i.e., using no semantic prompt) and text-guided noise suppression, where we used the reference transcription for the text-guided setting.

Target speaker extraction: We mixed each test sample with an utterance of a different speaker at a signal-to-interference ratio (SIR) which was randomly determined from the range between 0 dB and 20 dB. Also, we randomly chose one or more other utterances of the same speaker to create a 3-second-long enrollment clip to help models identify who the desired speaker is. Both the mixed and enrollment signals were used to derive the acoustic prompt as described in Section III-C. The task was to recover the original uncorrupted speech of the target speaker. As with the noise suppression task, we considered both text-less and text-guided settings.

Clean speech editing: For each test sample, we randomly selected a period of length between 10% and 50% of the whole utterance. We replaced the speech of the selected period with another randomly chosen speech sample of the same speaker. Given the partially replaced, speaker homogeneous speech and the reference transcription, the task was to generate a speech signal that follows the transcription without changing the speaker characteristics and the unreplaced portion of the input signal. In our experiments, we used the correct and locations based on the knowledge of the replaced segment.

Noisy speech editing: We added a randomly picked MUSAN noise sample to each test sample of the clean speech editing task. The SNR was chosen from the range of 0 dB to 20 dB. Given the noise-corrupted partially replaced speech and the reference transcription, the task was to generate a noisy speech signal that follows the transcription without changing the background noise, the speaker characteristics, and the unreplaced portion of the input speech.

Speech removal: The same dataset was used as the one used for noise suppression. Given a noisy speech signal, the task was to extract the noise signal by removing the speech. We considered only the textless case. Consequently, the input exclusively comprised the acoustic prompt corresponding to the noisy speech.

IV-B Metrics

For consistency and reproducibility, we opted to use objective metrics for individual tasks as described below.

Word error rate (WER): We employed the WER as a metric to evaluate the fidelity of the generated audio in adhering to the provided transcription. The ASR system utilized for our experiments was NeMo’s stt_en_conformer_transducer_large modelhttps://huggingface.co/nvidia/stt_en_conformer_transducer_xlarge, which is based on the Conformer Transducer architecture . We selected this particular ASR model based on its superior stability and robustness against noise and processing artifacts in comparison to other publicly available ASR models, as was observed during our preliminary experiments. Robustness in ASR is particularly crucial for tasks such as noise suppression and noisy speech editing. The WER metric was employed across all tasks, with the exception of speech removal.

Speaker similarity score (SIM): The speaker similarity score served as a metric to assess the coherence of the generated speech in relation to the speaker’s characteristics. This score was calculated as the cosine similarity between the speaker embeddings of the generated speech and the desired speech signals. The computation of speaker embeddings was performed using NeMo’s TitaNet-Largehttps://huggingface.co/nvidia/speakerverification_en_titanet_large. We employed the original audio data instead of utilizing an EnCodec-processed signal for the speaker similarity measurement to capture and reflect any potential speech deformation effects that may arise due to the use of the EnCodec model. SIM was used in zero-shot TTS, clean speech editing, and noisy speech editing.

DNSMOS: For evaluation in the noise suppression and target speaker extraction tasks, we utilized DNSMOS , a well-established model-based metric for predicting the perceived quality of acoustically corrupted speechhttps://github.com/microsoft/DNS-Challenge/tree/master/DNSMOS. Specifically, we employed the OVRL score from the DNSMOS P.835 model. To evaluate the performance of target speaker extraction, we employed a personalized DNSMOS model, which was tailored for this particular task and is available on the same webpage.

Perceptual Evaluation of Speech Quality (PESQ): For the noise suppression and target speaker extraction tasks, we also utilized PESQ . Unlike DNSMOS, PESQ is an intrusive metric that necessitates the clean reference signals. Consequently, PESQ is expected to assess the fidelity of the generated audio with respect to the original clean data.

Mel-cepstral distortion (MCD): MCDhttps://pypi.org/project/pymcd is a metric used to quantify the dissimilarity between two sequences of mel cepstra. We employed this metric to objectively measure the speech removal accuracy by comparing the estimated noise with the ground truth noise audio.

V Experiments

We sourced clean speech data from LibriLight, comprising 60 thousand hours of untranscribed English reading speech from over 7,000 speakers , as was performed in the zero-shot TTS experiment using VALL-E . To meet the specific training requirements for each task, data simulation was performed by following the methods employed for creating the evaluation data, as elaborated below. Note that, as discussed in Section III-D, we formed individual training mini-batches based on randomly selected tasks for each iteration.

For the noise suppression and speech removal tasks, we mixed the clean speech with noise samples from the DNS challenge corpus at SNRs between -5 dB and 20 dB. Our models were trained to recover the acoustic tokens of the clean speech and noise for noise suppression and speech removal, respectively. For the target speaker extraction task, we mixed the individual clean speech samples with those of other randomly chosen speakers with SIRs ranging from -5 dB to 20 dB. As regards clean speech editing, for each clean utterance, we randomly selected a subsegment of length ranging from 10% to 70%, and then substituted it with another audio segment from the same speaker with different content. We saved the start and end times of the replaced segment, which were used to insert the and tokens to the correct positions in the acoustic prompt during training. Furthermore, to create training samples for noisy speech editing, we added noise samples used in the noise suppression task to the partially replaced clean audio. As a result, we obtained pairs of noisy partially replaced speech and the corresponding original noisy speech, which served as the training data for the noisy speech editing task. The SNR range used for noisy speech editing training was also $$ dB.

Since LibriLight does not provide reference transcriptions, we adopted a pseudo-labeling approach to derive the semantic prompts, i.e., the phoneme sequences of the individual training samples, by following . Specifically, we transcribed the LibriLight training data with an off-the-shelf Kaldi model that was trained on the 960-hour Librispeech data with 3x speed perturbationhttps://kaldi-asr.org/models/m13.

V-B Model and training configurations

Both the SpeechX AR and NAR models share the same Transformer architecture, featuring 12 layers, 16 attention heads, an embedding dimension of 1024, a feed-forward layer dimension of 4096, and a dropout rate of 0.1.

We conducted experiments employing two initialization methods: random initialization and VALL-E initialization (refer to Section III-D for details). In the random initialization scenario, we trained the SpeechX model for 800K iterations. The model optimization utilized the AdamW optimizer, with the learning rate undergoing a warm-up phase for the initial 32K updates, peaking at 5×10−45\times 10^{-4}, before transitioning into a linear decay phase. Conversely, with VALL-E initialization, we opted for 400K iterations, as the initial model already underwent zero-shot TTS training over 400K iterations. In this instance, the learning rate scheduler was retained, but the warm-up period was shortened to the first 20K updates.

V-C Baseline expert models

We employed expert models for different tasks to establish comparison baselines. For zero-shot TTS, we utilized VALL-E by following the model configuration outlined in the original paper . For the noise suppression task, we employed a non-causal Deep Complex Convolutional Recurrent Network (DCCRN) , which is a widely recognized model for noise suppression. Our training data for DCCRN came from Microsoft’s internal dataset, and we further fine-tuned the model using the ASR objective based on the training recipe of . For target speaker extraction, we leveraged VoiceFilter , employing a bidirectional LSTM configuration. We relied on a publicly available implementation of VoiceFilterhttps://github.com/Edresson/VoiceSplit. Finally, for speech editing, we employed A3T as the baseline. The implementation of A3T that we used is also publicly accessiblehttps://github.com/richardbaihe/a3t.

V-D Results

Table. II shows the performance analysis of SpeechX in various tasks compared to the individual expert models. We can see that initializing the model parameters using an exiting VALL-E model checkpoint was beneficial across all tasks, especially in terms of WER.

In noise suppression and target speaker extraction, SpeechX exhibited superior performance in terms of WER compared to the respective expert models. Conventional regression-based noise suppression and target speaker extraction models are known to suffer from processing artifacts, which our WER results confirmed. SpeechX was able to avoid this detrimental effect thanks to the audio-text-based generation capability. On the other hand, in terms of DNSMOS and PESQ scores, it lagged behind the expert models. This can largely be attributed to the impact of the codec model used, as discussed in detail in Section V-D5. The investigation into the speech removal task revealed that SpeechX demonstrated substantial improvement in MCD, showcasing its efficacy in removing speech. These results underscore the versatility of the SpeechX model in handling enhancement-related tasks, while also highlighting the usefulness of the audio-text-based speech generation capability that SpeechX provides.

In the zero-shot TTS task, SpeechX demonstrated a slight advantage over the baseline VALL-E model in terms of WER while concurrently achieving a comparable speaker similarity scoreTo avoid potential confusion, it should be noted that our experimental setup corresponds to the non-continual evaluation configuration utilized in the original VALL-E work.. Furthermore, for the clean speech editing task, SpeechX exhibited significant improvement over the baseline A3T model. The WER observed in the speech editing task was slightly higher than the WER obtained in the zero-shot TTS task, even though one might anticipate that they should fall within the same range. This discrepancy could be attributed to certain test samples where the length of non-edited speech was shorter than three seconds. These results highlight that SpeechX is equally effective in tasks primarily focusing on speech generation capability, rather than transformation ability.

V-D2 Speech editing for clean and noisy speech

Table II also compares the speech editing results between clean and noisy speech in terms of WER and SIM. Editing noisy speech poses greater challenges than clean speech, as it requires modifying the spoken content while preserving background noise. This difficulty is evident from a WER gap of 38.29% vs. 42.48% observed between the clean and noisy audio signals to be edited as well as A3T’s limited WER improvement from 42.48% to 32.17%.

Nonetheless, the SpeechX model successfully edited the noisy speech, reducing the WER to 13.95% after processing. This demonstrates the model’s robustness to acoustic noise in the input signal. The high SIM score of 0.65 shows the model largely preserved speaker characteristics, even with noise present. Our observation revealed the model retained background noise, as confirmed by our provided demo samples. Fig. 2 compares mel spectrograms for two exemplary pairs of input and generated speech signals. In the first example, the input speech contained periodic noise in the middle frequency range. SpeechX preserved this background noise over the full input duration while selectively modifying only the foreground speech during the period beginning at two seconds. A similar observation can be made for the second example, wherein the alteration was applied to the first half of the speech content. In summary, the results demonstrate SpeechX model’s effectiveness at noisy speech editing while maintaining speaker identity and background noise. Future work should develop a metric to quantitatively evaluate noise cloning capability.

V-D3 Effectiveness of text input in noise suppression and target speaker extraction

With SpeechX, it is feasible to perform noise suppression and target speaker extraction using solely the acoustic prompt as input. To assess the efficacy of incorporating additional text input in the SpeechX model, we conducted noise suppression and target speaker extraction experiments where we employed only the acoustic prompt as the model input. Specifically, the input for noise suppression comprised the noisy speech, while for target speaker extraction, it consisted of the mixed speech and the target speaker’s enrollment audio.

The experimental results are presented in Table III. For both tasks, omitting the text input resulted in a noticeable increase in WER, whereas the degradation in DNSMOS and PESQ scores was modest. These findings suggest that leveraging the text input was particularly beneficial for enhancing the intelligibility of the output speech. In target speaker extraction, a significant impact on the DNSMOS score was observed, indicating that the text input aids in disentangling the target speaker’s voice from the interfering talker. Notably, while relying solely on the acoustic prompt led to WER degradation, the achieved WERs were still comparable to those of the baseline expert models.

V-D4 Effect of multi-task training

We also conducted experiments where we used subsets of the tasks during training to explore potential interactions between different tasks. Specifically, in addition to the VALL-E model and the fully-trained SpeechX models that used the complete set of the tasks, we trained two additional SpeechX models: one trained exclusively for zero-shot TTS and speech editing tasks, and the other trained on the zero-shot TTS, speech editing, noise suppression, and speech removal data.

Table IV shows the experimental results. The inclusion of speech editing during training led to an enhancement in WER for zero-shot TTS while allowing the model to learn about the speech editing task. Considering the strong parallels between zero-shot TTS and speech editing, this improvement can be attributed to the speech editing training task introducing additional variations to the distribution of the training data. Further inclusion of the noise suppression and speech removal tasks during training resulted in degradation in clean speech editing performance, while concurrently enhancing the performance for noisy speech editing. This suggests that exposing the model to noisy speech samples from these additional tasks improved the model’s robustness to acoustic noise at the expense of clean speech generation. Also, it is noteworthy that introduction of the target speaker extraction tasks to the training data did not compromise the model’s proficiency in noise suppression and speech removal.

V-D5 Limitation of current neural codec model

The performance of SpeechX is inherently constrained by the accuracy of the neural codec model employed for acoustic tokenization. It should be noted that, in all previous experiments, we compared SpeechX’s results with the reference (i.e., no-processing and expert model) results obtained without any neural codec processing. To gain a more precise interpretation of SpeechX’s results, we conducted an additional experiment where we applied compression and decompression to the LibriSpeech test-clean data without any intermediate processing, measuring EnCodec’s impact on performance metrics.

Table V shows the experimental results. It is evident that processing the signals with the codec model resulted in varying degrees of performance regression across all metrics. Notably, the PESQ score dropped from 4.64 to 2.69 for the clean speech input. Our assessment indicates that while EnCodec produced slightly noticeable speech quality degradation, the significant PESQ degradation may be partly attributed to the mismatch between the PESQ algorithm and EnCodec’s training objective. While we utilized EnCodec due to its accessibility and prior usage, future work should address this issue by developing an acoustic tokenization model more suitable for handling speech under various acoustic conditions.

VI Conclusion

In this paper, we described SpeechX, a novel versatile speech generation model capable of handling diverse audio-text-based speech generation tasks, including zero-shot TTS, noise suppression, speech removal, target speaker extraction, and speech editing. For noise suppression and target speaker extraction, the proposed model provides a unified way for incorporating the knowledge of transcriptions. Also, regarding speech editing, SpeechX enables modifying the spoken content of a speech signal that contains a fair amount of background noise. SpeechX adopts a language modeling approach to generate acoustic tokens conditioned on textual and acoustic prompts, where additional task-dependent tokens are incorporated in a multi-task learning framework to support various speech transformation capabilities beyond zero-shot TTS. We demonstrated SpeechX’s efficacy through comprehensive experiments. The proposed model represents an important step toward unified generative speech models. Further research can build on this work by expanding the tasks supported, enhancing robustness, and developing more advanced conditioning mechanisms.

References