VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
Tianrui Wang, Long Zhou, Ziqiang Zhang, Yu Wu, Shujie Liu, Yashesh Gaur, Zhuo Chen, Jinyu Li, Furu Wei
Introduction
Recent years have revealed a significant convergence in model architectures, training techniques, and inference methods across research domains Vaswani et al. (2017); Devlin et al. (2018); Brown et al. (2020); Wang et al. (2022); OpenAI (2023). Although different model frameworks, e.g., encoder-only BERT Devlin et al. (2018), decoder-only GPT Radford et al. (2018), and encoder-decoder based BART Lewis et al. (2019) and T5 Raffel et al. (2020), have made remarkable improvements in natural language processing (NLP) tasks, emerging advances like ChatGPT OpenAI (2022) demonstrate that generative models have strong modeling abilities and potential to address various NLP tasks, especially for few-shot or zero-shot problems.
The success of generative models in NLP also extends to speech processing, namely speech synthesis tasks. For instance, the VALL-E Wang et al. (2023) model transforms continuous speech signal to discrete codec codes with EnCodec model Défossez et al. (2022), and trains a neural codec language model by scaling up TTS training data to 60K hours. VALL-E X Zhang et al. (2023b) extends this approach by introducing language ID to support cross-lingual TTS and speech-to-speech translation. The results show that with large training data, decoder-only models like VALL-E (X) can show in-context learning capacity and perform significantly better than baseline systems in zero-shot TTS tasks. However, VALL-E (X) is only designed for the single speech synthesis task, and can not apply to other speech tasks. This inspires us to ask one question: Is one decoder-only generative model all you need for speech recognition, synthesis, and translation?
In this paper, we extend the application scenario of the decoder-only network from speech synthesis to all speech-processing tasks, including speech recognition, synthesis, and translation. Specifically, we present a multi-lingual multi-modal auto-regressive Transformer model VioLA to unify the speech-to-text, text-to-text, text-to-speech, and speech-to-speech tasks into a conditional codec language model task, as illustrated in Figure 1. By converting continuous speech signals to discrete codec codes via an offline neural codec model EnCodec, we can treat speech representation as textual tokens and employ a decoder-only model to optimize multi-modality tasks effectively. Moreover, we leverage task IDs (TID) and language IDs (LID) to enhance the capability of distinguishing different languages and tasks. We hope the proposed method can promote the era of codec-based speech processing with generative models.
We trained VioLA model under a multi-task learning framework, including automatic speech recognition (ASR), machine translation (MT), and text-to-speech synthesis (TTS) tasks. Massive evaluation on speech-to-text recognition and translation tasks, tex-to-text synthesis tasks, tex-to-speech translation tasks, and speech-to-speech translation tasks demonstrates that our proposed VioLA can be effectively applied to these single-modal and cross-modal tasks. Besides, the proposed model is able to maintain the in-context learning ability for speech synthesis tasks. To the best of our knowledge, it is the first work to quantitatively explore codec-based auto-regressive Transformer decoder models for various speech processing tasks.
Related Work
Our work is built on an auto-regressive Transformer network for speech recognition, translation, and synthesis tasks. In previous speech processing, different models with varying architectures were used for different tasks. The primary end-to-end model structures used for ASR Li (2022); Prabhavalkar et al. (2023) are connectionist temporal classification (CTC) Graves et al. (2006), attention-based encoder-decoder network (AED) Chiu et al. (2018), and transducer networks with RNN or Transformer Chen et al. (2021a). For speech synthesis tasks, AED methods such as Tacotron Wang et al. (2017), Tacotron2 Shen et al. (2018), and TransformerTTS Li et al. (2019), and duration-based methods, including Fastspeech Ren et al. (2019), Fastspeech2 Ren et al. (2020), and RobuTrans Li et al. (2020) are widely used. At the same time, encoder-only models Devlin et al. (2018), decoder-only models Radford et al. (2018), and encoder-decoder models Sutskever et al. (2014); Lewis et al. (2019) have also been extensively explored, and applied to various natural language processing tasks, such as machine translation, text summarization, question answer, and so on.
In terms of building one model for all speech tasks, HuBERT Hsu et al. (2021) and WavLM Chen et al. (2022) learn a universal speech representation model using unsupervised masked unit prediction methods, like BERT Devlin et al. (2018). After pre-training, the pre-trained model, regarded as a speech representation extractor, and added specifical-task modal parameters, are fine-tuned using supervised data of downstream tasks. Inspired by T5 Raffel et al. (2020), SpeechT5 Ao et al. (2021) adopts a universal encoder-decoder based sequence-to-sequence model for all spoken language tasks, e.g., speech recognition, speech translation, speech synthesis, etc. To process different modalities, SpeechNet Chen et al. (2021b) uses multiple modal-specifical encoders and decoders for speech and text modalities. Whisper Radford et al. (2022) also employs an encoder-decoder framework trained with large-scale supervised data, but focuses on speech recognition and translation tasks.
Different from the above work, our proposed VioLA model tries to employ a Transformer decoder-only model for all spoken language tasks. Hence our work is also related to text-based and speech-based GPT models. Generative pre-training (GPT) Radford et al. (2018) methods have demonstrated the great potential capability for few-shot and zero-shot problems in ChatGPT OpenAI (2022) and GPT-4 OpenAI (2023). Moreover, researchers also apply the decoder-only model to text-to-speech synthesis tasks, namely VALL-E Wang et al. (2023), and VALL-E X Zhang et al. (2023b). They transform speech signals into discrete tokens, input cascaded discrete text and speech sequences to the Transformer decoder network, and show strong zero-shot transfer ability. In parallel with our work, SpeechGPT Zhang et al. (2023a) was recently proposed to perceive and generate multi-model content. However, it performs speech discretization with hidden units from HuBERT, which contains less acoustic information than codec codes, and only lists some audio examples without quantitative results for cross-modal instruction following and speech dialogue tasks. Our VioLA is based on VALL-E (X) and extends it to more speech tasks using the decoder-only Transformer model. It is also the first to quantitatively explore codec-based auto-regressive models for various spoken language tasks.
Multi-Task Codec Language Model
Our proposed VioLA is built upon the neural codec language models VALL-E and VALL-E X, which treat text-to-speech synthesis as a language model task, like GPT, and employ acoustic tokens (audio codec codes) as an intermediate representation of original speech. VALL-E (X) generates the codec sequences (8 codes for each frame), based on which a codec decoder generates the final waveforms. VALL-E (X) contains two key components, the auto-regressive codec language model and the non-autoregressive codec language model. The former is responsible for predicting the acoustic tokens of the first codec code for each frame based on the semantic tokens (phoneme sequences) in an auto-regressive manner, and the latter is used to generate the other 7-layer codes according to the sequence of the first-layer codes in parallel with the layer-level iterative generation method. In this work, we aim at extending VALL-E from the single TTS task to more speech tasks, e.g., speech recognition and speech-to-text translation, and focus on boosting the capability of the auto-regressive codec language model.
2 Problem Formulation
VioLA regards all speech tasks as conditional codec language modeling. The core of our approach is to convert all the speech data into sequences of tokens (codec codes) as discrete text, and the speech-to-text, text-to-text, and text-to-speech tasks turn to token-based sequence conversion tasks with the decoder-only model.
Give the dataset , where denotes different tasks, is the index of training samples, and denote audio sample and its transcription for ASR task, source text and target text for MT tasks, text transcription and audio sample for TTS task respectively. The speech data and text data in dataset is first converted to codec codes and phonemes with an offline audio codec model (EnCodec) and grapheme to phoneme (G2P) conversion tool.
In this work, our goal is to leverage a conditional codec language model to model various spoken language tasks, and the model is optimized by maximizing like generative pre-training methods in GPT, with a multi-task learning framework. The trained VioLA is excepted to have the capability to conduct all the ASR, MT, TTS, cascaded speech-to-text translation, and cascaded speech-to-speech translation tasks.
3 Model Framework
Following VALL-E (X), an auto-regressive codec Transformer language model is leveraged as the core architecture of VioLA, shown in the left of Figure 2. It comprises an embedding encoding module, a Transformer decoder block, and a prediction layer. Furthermore, we leverage language IDs to distinguish different languages and task IDs to accomplish different tasks, including speech recognition, machine translation, and speech synthesis.
The proposed auto-regressive codec LM is optimized by multiple tasks, using ASR corpus (, ), MT corpus (, ), and TTS corpus (, ), as multiple inputs and outputs of Figure 2 (a). The multiple inputs, including semantic and acoustic tokens, are first processed using an embedding encoding module, which will be introduced in the next subsection (3.3.2) in detail. After that, the standard Transformer decoder, consisting of the crucial multi-head attention network and feed-forward network, is employed to model the dependency relationship of semantic tokens and acoustic tokens. Note that all tokens in multiple inputs can attend over all positions in the input tokens, and each token in multiple outputs can attend to previous multiple outputs and all input tokens. Finally, we use a prediction layer to convert hidden states into the discrete index of vocabulary.
3.2 Pre-Net and Pos-Net
For speech synthesis tasks, the above auto-regressive code LM only estimates the first-layer acoustic tokens. Following VALL-E (X), we adopt another non-auto-regressive codec Transformer language model to generate the all-layer acoustic tokens. For simplicity, we refer readers to read the non-auto-regressive codec LM part in VALL-E (X). Different from the original non-autoregressive codec LM, the LSTM module is also introduced for encoding acoustic tokens in our non-autoregressive codec language model, as Equation 3.
4 Training Objectives
The multi-task auto-regressive codec language model is optimized by performing speech-to-text recognition, text-to-text translation, and text-to-speech synthesis tasks, as shown in Figure 2. Particularly, we integrate additional task IDs (e.g., , , ) into the proposed model to accomplish different tasks.
aims to predict the phoneme sequences according to the acoustic tokens from the original speech. Here, eight layers of the acoustic tokens are employed as the input to introduce more acoustic information. The training is implemented as teacher forcing and auto-regressive strategies with an ASR task ID , as
where the acoustic tokens, task ID, and semantic tokens are concatenated and then fed into the language model.
is to translate the source phoneme sequences into the target language phonemes . Similar to speech recognition, MT task is optimized by maximizing the following probability,
focuses on generating the first-layer quantized acoustic tokens depending on their corresponding transcription with auto-regressive manner, as follows,
We optimize the auto-regressive codec LM with a multi-task learning framework, namely,
where these three tasks are simultaneously trained.
5 Inference Methods
After training, our model can be applied to various tasks, including ASR, MT, TTS, speech-to-text translation (S2TT), and speech-to-speech translation (S2ST) tasks. For the last two tasks, we adopt a cascaded pipeline, consisting of the first three base tasks. We leave end-to-end speech-to-text translation and speech-to-speech translation for future work. It is noted that we use sampling methods for speech synthesis tasks, and beam search for other tasks. For speech synthesis tasks, we can use speech utterances with the same language or different languages as speech prompts, like VALL-E and VALL-E X. We perform five synthesis inferences for one text, and the utterance with the highest score is selected according to speaker similarity (SS) score, aka Strategy I, or the combination of speaker similarity and word error rate (WER) scores, aka Strategy IIStrategy I chooses the utterance with the highest speaker similarity. Strategy II chooses the utterance with the highest score of SS minus WER (SS and WER are the normalized scores of the five results), and the calculation of SS and WER will be introduced in Section 4.5..
Experiments
We evaluate our proposed VioLA on tasks of ASR, MT, S2TT, zero-shot TTS, and zero-shot S2ST.
Our proposed VioLA is trained on two speech recognition datasets, which can also be used to train TTS tasks as VALL-E (X), and two machine translation datasets. The Chinese ASR data are from WenetSpeech Zhang et al. (2022), a total of 10,000 hours of multi-domain labeled speech. The English ASR data are from LibriLight Kahn et al. (2020), a total of 60,000 hours of unlabeled audiobook speech, where the transcripts are generated by an ASR modelhttps://github.com/kaldi-asr/kaldi/tree/master/egs/librispeech trained on the LibriSpeech Panayotov et al. (2015). AI Challengerhttps://challenger.ai/competition/translation and WMT2020https://www.statmt.org/wmt20/translation-task.html datasets are employed for MT training, containing 63M English-Chinese sentence pairs in conversion, drama, and news domains.
We evaluate our proposed method on WenetSpeech, EMIME Wester (2010), LibriSpeech, and the WMT2020 datasets. EMIME contains 25 pairs of bilingual sentences recorded by seven female and seven male native Chinese speakers with two microphones. We evaluate speech-to-phoneme ASR on the development set of WenetSpeech. The MT performance is evaluated on the test set of WMT2020. The Chinese-to-English S2TT (ZHEN S2TT), zero-shot English TTS prompted by Chinese speech (ZHEN TTS), and zero-short Chinese-to-English S2ST (ZHEN S2ST) are evaluated on the EMIME dataset. The zero-shot English TTS prompted by English speech (EN TTS) is evaluated on the Librispeech dev-clean and test-clean sets.
2 Data Pretreatment
The proposed VioLA is trained on two types of discrete tokens, semantic tokens (phoneme sequences) and acoustic tokens (codec codes). The transcription of ASR data and the bilingual text of MT are converted into phonemes via the lexicon provided by ASR datasets and the International Phonetic Alphabet (IPA)-based unified phoneme set called BigCiDianhttps://github.com/speechio/BigCiDian. We also use the Kaldi force-alignment toolhttps://github.com/kaldi-asr/kaldi/tree/master to generate the alignment information for clipping utterances during training. The speech data are quantized into acoustic tokens using the EnCodechttps://github.com/facebookresearch/encodec with 6 kbps bandwidth in our method, resulting in 8 tokens for each frame. We also pre-process speech into the 80-dimensional feature with a frameshift of 10ms for baselines.
3 Model Architecture
The auto-regressive codec language model employs a Transformer decoder with an attention dimension of 1024 and the FFN dimension of 4096. Sinuous position embedding is separately computed for multiple input and output sequences in Figure 2 (a). In embedding encoding modules, we use layer-specific 1024-dimensional embeddings for each layer of acoustic tokens, and one 1024-dimensional embedding for semantic tokens. Besides, 3-layer unidirectional LSTM is employed for this pre-net. Since VioLA is required to support more tasks, which may need larger model capability with more parameters, we trained 12-layer and 18-layer decoder-only auto-regressive models for comparison, called VioLA (12L) and VioLA (18L), with 178M and 250M parameters, respectively.
We built two strong baselines, including encoder-decoder (AED) models and decoder-only language model (LM) for comparison. AED models adopt 6 Transformer encoder and decoder layers, with additional cross-attention modules, for MT and Fbank-based ASR tasks. Two-layer convolution with 4-times down-sampling of time dimension is employed to process the Fbank before Transformer. We trained the Fbank-based ASR models without SpecAugment Park et al. (2019) to compare the performance of different input features for ASR fairly. The configuration of LMs on single tasks is the same with 12-layer VioLA.
4 Training Details
Our multi-task auto-regressive codec language model is trained on ASR, MT, EN TTS, and ZH TTS total of 4 tasks simultaneously for 800K steps on 32 V100 GPUs with a batch size of 25 seconds per GPU (the losses of 4 tasks are accumulated for one step). We re-segment the training data to an average utterance duration of 12 seconds for effective training, and the maximum sentence length is 20 seconds. Each batch loads MT data with the same number of semantic tokens as the number of 25s acoustic tokens. The maximum learning rate is with warm-up step of 80K. We follow the configuration of VALL-E X to train our non-autoregressive language model, which integrates additional embedding encoding modules as introduced in Section 3.3.2.
5 Evaluation Metrics
We employ phoneme error rate (PER) and BLEU to evaluate the performance of models on speech recognition tasks and three translation tasks (MT, S2TT, and S2ST). Instead of human evaluation, we use the automatic evaluation metrics, including the word error rate (WER), speaker similarity (SS), and speech naturalness (SN) for EN TTS to evaluate the generated speech for simplicity and convenience. For synthesized speech in TTS and S2ST tasks, we first utilize a HuBERT-Large ASR modelhttps://github.com/facebookresearch/fairseq/tree/main/examples/hubert finetuned on LibriSpeech to transcribe it into text, then calculate the above WER and BLEU scores. Given generated and prompt speech utterances, the SS is measured by an automatic speaker verification (ASV) modelhttps://github.com/microsoft/UniSpeech/tree/main/downstreams/speaker_verification, ranging from -1 to 1. The larger the SS, the more similar the speakers of the two utterances are. SN score of generated speech is measured by the open-source NISQAhttps://github.com/gabrielmittag/NISQA Mittag and Möller (2021).
6 Main Results
We first evaluate the speech recognition performance of different models on the development set of WenetSpeech. The results are summarized in Table 1. The Fbank-based AED achieves the lowest PER. Fbank-based decoder-only LM causes a slight increase in the PER, which proves that the LM can still achieve comparable performance on the speech recognition task. As can be seen from the results of VioLA (12L), integrating multiple tasks into a single model slightly impairs the performance of speech recognition tasks. However, VioLA (18 L), with comparable parameters to the AED, achieves acceptable results for codec-based speech recognition with a PER of 11.36.
We verify the performance of different models on the machine translation task based on the test set of WMT2020 after processing in subsection 4.2, and the results are shown in Table 2. LM (decoder-only model) achieves comparable results to the AED (encoder-decoder model) on the machine translation task. VioLA (18L) can obtain the improvement of +2.53 BLEU scores than VioLA (12L). Furthermore, VioLA (18L) achieves the highest BLEU score (56.97) with a comparable number of parameters as AED models (250.6M vs. 242.5M).
We perform ZHEN S2TT tasks on the EMIME dataset by cascading ASR and MT tasks of VioLA. As listed in Table 3, the BLEU score of the codec-based cascaded LM is only 0.28 lower than the BLEU of 55.98 achieved by the Fbank-based cascaded AED models. We speculate that comparable MT performance can compensate for the shortcomings of ASR performance for our model. Moreover, VioLA (18L) achieves a comparable BLEU score on the S2TT task (55.85 vs. 55.98) with a 48.7% reduction in parameters compared to the Fbank-based cascaded AED models.
As mentioned in subsection 3.5, we synthesize the English speech of corresponding text prompted by an English speech utterance on selected samples of dev-clean and test-clean setsTo keep the consistency between training and inference, we select 1802 samples of 4 to 10 seconds with the same speaker’s speech prompt of 3 to 5 seconds from the LibriSpeech dev-clean and test-clean sets to evaluate different models on the zero-short EN TTS task.. As shown in Table 4, VioLA (18L) significantly outperforms VioLA (12L) in all ASR, MT, and TTS tasks, which suggests that LM with more tasks and larger data may require the larger model to train. Compared to the VALL-E X, the speaker similarity of VioLA (18L) is relatively improved by 2.0%, the WER is relatively decreased by 14.6%, and the speech naturalness is improved by 0.02. Our model also shows powerful in-context learning capabilities in keeping the speaker’s similarity of prompt speech.
Different from EN TTS, zero-shot cross-lingual ZHEN TTS is to synthesize English speech by using both the English text and the Chinese speech as prompts, like VALL-E X. Experimental results are summarized in Table 5. VioLA (18L) improves the average SS by 0.01 and BLEU by 2.20 compared to the VALL-E X, and generates more stable codecs for natural English speech with a 3.35 naturalness score.
We conduct cascaded ASR, MT, and TTS tasks of VioLA to accomplish ZHEN S2ST. Compared with the codec-based cascaded LMs, the Fbank-based cascaded AED with VALL-E X achieves an average 0.76 improvement in the BLEU score with its advantage in speech recognition. VioLA (18L) with 1.5 times parameters than VioLA (12L) balances three single tasks and achieves the best performance on the zero-short ZHEN S2ST task with an average SS of 0.50, an average BLEU of 47.27, and an average speech naturalness of 3.55.
7 Analysis
In this section, we will analyze the effect of the layer number of acoustic tokens and LSTM networks in embedding encoder modules for speech recognition and synthesis, respectively.
As introduced in Section 3.3.2, we average the acoustic embeddings of the eight codecs for each frame, resulting in the 1024-dimensional frame-level input for Transformer. We first explore the effect of the number of acoustic codecs per frame for speech recognition, and the results are shown in Table 6. As we can see from the results, the increase in the layer number of acoustic tokens significantly improves speech recognition performance. Compared with one layer of acoustic tokens, eight layers of acoustic tokens provide more speech information, and the phoneme error rate is relatively reduced by 54.5%, with only seven additional 1024-dimensional embeddings.
Inspired by the decoder of EnCodec, VioLA employs an additional LSTM network to process the acoustic embedding. The results in Table 6 show that, given the 8-layer acoustic tokens, the introduction of LSTM for embedding encoding modules in the auto-regressive model (LSTM) relatively reduces the PER by 8.7% on ASR task. We also investigate the influence of LSTM networks in TTS tasks for both auto-regressive and non-auto-regressive models. The results of the zero-shot TTS ablation experiments are shown in Table 7, which demonstrates that the integration of LSTM networks in both auto-regressive and non-autoregressive models improves the performance of the decoder-only model on speech synthesis tasks. The average WER of VioLA with the is relatively reduced by 12.5% compared to the VALL-E X.
Conclusion
In this paper, we propose VioLA, an auto-regressive codec language model for multi-modal tasks involving speech and text. VioLA is the first work to quantitatively explore and apply a Transformer decoder-only model and codec-based methods to various spoken language tasks. By discretizing speech signals to discrete acoustic tokens as textual semantic tokens, all spoken processing tasks can be formulated into one conditional language modeling task with the same optimization objective. Massive experiments on speech-to-text recognition and translation, machine translation, text-to-speech synthesis, and speech-to-speech translation demonstrate the effectiveness and superiority of the proposed model.
Limitations
Although the proposed conditional codec language model VioLA is successfully applied to various speech tasks, e.g., speech recognition, synthesis, and translation tasks, it still suffers from some limitations: (1) It is optimized with supervised data and does not make full use of unsupervised data, such as large-scale text corpora and unlabeled speech. (2) The model only shows in-context learning ability on speech synthesis tasks, rather than all speech processing tasks. (3) Current VioLA only supports cascaded inference methods for speech-to-text translation tasks, and speech-to-speech translation tasks, without end-to-end capability.