Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
Soumi Maiti, Yifan Peng, Shukjae Choi, Jee-weon Jung, Xuankai Chang, Shinji Watanabe
Introduction
In recent years text language models (textLMs) have emerged as a powerful generative model in natural language processing (NLP) . These textLMs can accommodate multiple tasks within a single model, leading to improvement in performance across a variety of tasks. On the other hand, with advances in discrete speech representations speech language models (speechLMs) have also been proposed. However, prior speechLMs focus on individual tasks, such as speech continuation or text-to-speech (TTS) . Our hypothesis is that by unifying diverse speech tasks into a generative language model (LM), we can potentially address multiple speech tasks using a single model with improved generalization thanks to multitask learning.
Traditionally, speech applications such as automatic speech recognition (ASR) and text-to-speech (TTS) use encoder-decoder architectures . These architectures consist of an encoder, for input processing and a decoder, for generating the output. For example, speech-to-text involves a speech encoder and a text decoder, whereas text-to-speech employs a text encoder and a speech decoder. Integrating task-specific and modality-specific encoder-decoder components complicates the incorporation of multiple tasks . In contrast, we can simplify multitask integration with a joint speech-text decoder-only model (depicted in Fig. 1).
In this work, we investigate two main questions. Firstly, can we cast diverse speech tasks as language modeling? ASR and TTS are used as example speech tasks. Secondly, can we combine speech tasks in a joint speech-text language modeling framework? To this purpose, we introduce a novel LM framework VoxtLM (Voice-text Language Model). VoxtLM combines multiple speech tasks within a single autoregressive decoder model. Specifically, we combine four tasks: speech recognition (speech-to-text), speech synthesis (text-to-speech), text generation (text-to-text), and speech generation (speech-to-speech). We create a Voxt (voice + text) vocabulary by merging self-supervised discrete speech tokens with the text vocabulary and incorporate sub-word modeling to efficiently process long sequences of speech. We show that VoxtLM can model both ASR and TTS as conditioned language model. In addition, combining four tasks leads to improvement in speech generation, ASR, and TTS. Significant improvement is observed in the TTS task with improvement in both intelligibility (28.9 to 5.6) and neural-predicted quality (2.68 to 3.90). Additionally, we demonstrate that improved initialization with pretrained textLM and scaling model parameters help in ASR. To ensure reproducibility, we use publicly available datasets, open-source our training and inference in the form of open-source toolkit ESPnet recipehttps://github.com/ESPnet/ESPnet and make model checkpoints available. TTS samples are also available.https://soumimaiti.github.io/icassp24_voxtlm/
Related Work
Discrete speech representations. Speech signals can be represented as two types of discrete tokens: semantic tokens and acoustic tokens. Semantic tokens are quantized from self-supervised learning features (e.g., HuBERT , w2v-BERT ) through clustering, which mostly captures the linguistic content. Acoustic tokens are generated by audio codec models . They capture rich acoustic information which is suitable for high-quality speech synthesis, but they consist of multiple code streams and are thus difficult to model. In this work, we follow GSLM to use semantic tokens derived from HuBERT.
Joint modeling of speech and text. Several studies propose to learn shared speech-text representations in a self-supervised manner. However, they employ separate encoders and decoders for different modalities. They also require additional losses like an alignment loss to encourage cross-modal transfer between speech and text. Recent concurrent studies employ a single model for multiple speech and text conversion tasks , which are similar to our approach. SpeechGPT uses a three-stage adaptation to combine audio generation with textLMs. PolyVoice applies speechLM to speech-to-speech translation (S2ST) with three decoder-only LMs. VioLA extends VALL-E for ASR and S2ST. Among them, VioLA is the most related method to this work. However, VioLA does not incorporate speech or text continuation tasks and requires additional sequence modeling for speech representations, which makes it more complicated than our approach. Moreover, we utilize textually pre-trained OPT for better initialization inspired by and leverage different speech tokens. Also in comparison to other works, our work is fully reproducible.
Method
Consider is a text utterance from a vocabulary with length . The probability of can be expressed as Now, when dealing with a continuous speech signal, we can convert it into discrete speech tokens (dst), represented as using a tokenizer. In this context is the vocabulary of discrete speech tokens. These discrete speech tokens can be treated as spoken language within and modeled in a manner similar to text. We combine text and speech in a new vocabulary Voxt vocabulary by . Therefore, we can model the probability of both speech and text tokens as , where . This probability is expressed as:
Here, can represent discrete speech tokens or text tokens or various combinations of and .
Fig. 2 illustrates the model’s overall architecture. Input of VoxtLM can be both speech and text within the vocabulary. To process speech, we use two additional modules to convert between continuous and discrete domains in speech. The speech tokenizer maps to , while the speech token decoder maps generated back to . Similar to , our speech tokenizer uses -means clustering to derive discrete features from the pretrained HuBERT . It is worth noting that selecting a small value may capture linguistic information effectively, but might fall short in representing other acoustic aspects particularly crucial for speech synthesis. We experiment with different to assess the impact. Furthermore, within vocabulary, we apply subword modeling to replace frequent patterns with metatokens. Such subword modeling technique is used to include more contextual information in text or to reduce the long sequence length of speech .
We use special tokens to guide the model in performing various tasks. Four such tokens are used: start-text and start-speech indicate the beginning of text or speech conditioning in the language model. generate-speech and generate-text instruct the model whether to generate speech or text. Table 1 shows examples of the Voxt data format for various tasks during training. Ideally, we can extend to more tasks with additional task-specific tokens.
1.2 Training
Initialization with pretrained textLM. Previous work shows that in speechLM initializing with a pre-trained textLM achieves better performance and faster convergence. Motivated by this, we use the pretrained textLM OPT to initialize VoxtLM weights and learn the embedding table from scratch. The same model configuration is used as the pretrained model except for . OPT is used due to training on publicly available data and the availability of smaller pretrained models.
1.3 Inference
The prediction from VoxtLM is expressed as:
For TTS, condition is the test text utterance and prediction is speech tokens . In ASR, condition is test speech tokens and prediction is the recognized text . For speech continuation, condition involves prefix speech tokens and prediction is continued speech tokens . For text continuation, the condition is text and prediction is continued text (summarized in Table 1. We use beam search in the inference phase.
2 Evaluation Metrics
For speech and text generation, we use perplexity (PPL) for evaluating models with same vocabulary size. For different vocabulary size models, we use spot-the-word error using sWUGGY and syntactic score using sBLIMP dev set . sWUGGY and sBLIMP are chosen as other speech LM works also report them.
For ASR, we use the word error rate (WER).
For TTS, we measure intelligibility with character error rate (CER) and quality using the neural-predicted mean opinion score (MOS) with MOSNet . We choose neural MOS prediction model because it scales to large number of evaluations and shows high-correlation with TTS evaluations in English.
Experiments
Dataset. We use a combination of speech-only, text-only, and paired speech-text datasets from public corpora.
Speech-only data: we use LibriLight (LL) with 60K hours of audiobook speech from 7K speakers (12M utterances).
Text-only data: we use the Librispeech (LS) external textLM dataset (40M text utterances).
For ASR, we use Librispeech with 960 hours of data ( 281K utterances). For an additional supervised data experiment, we use English Multilingual Librispeech (MLS) with 44K hours of data from 5490 speakers (11M utterances).
For TTS, we use LibriTTS (LT) with 580 hours of audiobook data from 2456 speakers and VCTK (VC) with 44 hours of studio recorded data from 109 speakers (404K utterances).
We standardized the data by downsampling speech to a 16kHz rate, converting text to lowercase, and removing punctuation. We use separate test/dev sets for each task. For textLM and speechLM, we use the test set from LS and dev sets from sWUGGY and sBLIMP, text for textLM, and speech counterpart for speechLM. For ASR we use speech-text test set from LS test-clean and test-other and report both test-clean/test-other separately. In TTS for computational efficiency, we create a test set of 100 utterances from two speakers from the LT test-clean. The test speakers are chosen via random sampling (specifically, speaker ids 1089 and 1284).
Experimental setup. To train the sub-word model, we use paired text-speech from ASR and TTS datasets. We experiment with three values (introduced in Sec. 3.1), 50, 200, and 1000, denoted as VoxtLM-. We also vary BPE sizes, setting them at , , and for values 50, 100 and 200, respectively. We use three configurations, small (=12, =768, =12), medium (=24, =1024, =16), and large (=24, =2048, =32), with , and detailed in Sec. 3.1.2. We use 4 A100 GPUs for training small/medium and 8 A100 GPUs for large with Adam optimizer and warmup learning rate schedule. Training data size varies considerably between different tasks. For example, the paired data for ASR and TTS are smaller than text-only data and smaller than speech-only data. We can assume that achieving optimal performance across all tasks requires balanced data for each of them. It is also worth noting that text-only data is more readily available compared to speech-only and paired data. Nonetheless, to assess the effect of different dataset sizes for tasks, we consider balanced and unbalanced data sets for training, as summarized in Table 2.
Single vs multitask. We compare multitask and four single-task models using VoxtLM-50. Single-task LMs are trained separately for each task (ASR, TTS, speechLM, and textLM) and are reported in the first row of Table 3, with each column representing a separate single-task model. Compared to single-task, VoxtLM shows competitive results for all four tasks, although the best model differs. For textLM exhibits higher sWUGGY but lower sBLIMP score. In speechLM, has the best scores in both sWUGGY and sBLIMP, followed by . In TTS, all multitask models show improvement compared to single task. ASR reports improvement in . We note that ASR is most affected in the unbalanced case: probably due to the lower ASR data ratio to textLM/speechLM (/ less). A smaller degradation in ASR is also observed in where ASR data ratio to textLM/speechLM is relatively better ( less).
Initialization with pretrained textLM. We compare with and without initialization with OPT for VoxtLM-50 with and report in Table 4. Initialization improves the performance of three tasks: textLM, speechLM, and ASR. For TTS, a slight degradation in CER is observed, whereas objective quality improves. In particular, better initialization aids ASR performance in the unbalanced scenario (reducing test-clean WER from 21.0 to 13.1).
Effect of token vocabulary size. We compare =50, 200 and 1000 as outlined in Table 5. Comparisons are made in and . For ASR and TTS, performance of =50 is poor. For speechLM with best sores on sWUGGY and sBLIMP are observed with the =200 model. TextLM, as expected, does not show a significant pattern with varying .
Scalability. Next, we explore whether model size can help with data balancing by comparing medium and large models with =200, presented in Table 6. All metrics in TextLM, speechLM, and ASR show improvement with larger model. TTS shows a very small degradation in intelligibility () and quality (). To mitigate the smaller ratio of paired data, we incorporate more supervised data for ASR in . We compare with = 200 and = 1000 and observe an improvement in the ASR task.
Comparison with single-task state-of-the-arts. Furthermore, we compare with state-of-the-art models in TTS, ASR, and speechLM. Note that these models aren’t fully comparable due to differences in training data, strategies, and architecture. Following models are used: for speechLM, we use GSLM and AudioLM , for TTS, we use VITS and for ASR we use E-Branchformer . For ASR, we compare two models: one using spectrogram as input (ASR-Fbank) and another using discrete speech tokens as input (dst-ASR-Hubert), trained following the procedure and the same speech tokenizer as VoxtLM-1000. We use a pretrained VITS model with LibriTTS. For speechLM (Table 7), GSLM-200 which uses the same tokenizer and a similar one-stage model, sBLIMP score is lower compared to VoxtLM. However, in AudioLM which uses two token representations (acoustic and semantic) and a three-stage model, both sWUGGY and sBLIMP scores are higher, suggesting potential for further improvement with hierarchical tokens and multistage training. For ASR, compared to dst-ASR-Hubert, which used the same tokenizer as VoxtLM, we observe a lower WER. Compared to ASR-Fbank (no tokenizer), WER is higher, such a trend is also observed in other discrete ASR models . In TTS (Table 8), compared to VITS, VoxtLM reports better intelligibility and quality. Although VoxtLM is trained with a larger data set compared to VITS, it is interesting to note that for traditional TTS diverse training data with more noise and more speakers degrade performance but here improvement is observed.
Finally, our experimental results show that both ASR and TTS can be modeled as language modeling tasks. Moreover, using special tokens we can combine ASR and TTS with joint speech-text language modeling framework. Although the four tasks are quite different, combining four tasks leads to improvement.
Conclusion
The integration of speech and text tasks within a joint language modeling framework presents a promising avenue for speech processing. We present a special token-based approach to combine four speech and text tasks: speech recognition, speech synthesis, text and speech generation. Our results demonstrate that by integrating different speech tasks into one generative model, we can improve the performance of the tasks. In particular, TTS shows impressive performance compared to the state-of-the-art VITS. We will expand this work to include more speech tasks in the future.
Acknowledgements Experiments of this work used the Bridges2 system at PSC and Delta system at NCSA through allocations CIS210014 and IRI120008P from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, supported by National Science Foundation grants #2138259,#2138286, #2138307, #2137603, #2138296.