SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data
Ziqiang Zhang, Sanyuan Chen, Long Zhou, Yu Wu, Shuo Ren, Shujie Liu, Zhuoyuan Yao, Xun Gong, Lirong Dai, Jinyu Li, Furu Wei
Introduction
Speech and text are two important carriers of human communication, and they can be converted into each other through speech recognition and synthesis systems. In past years, the unimodal self-supervised representation learning has been well explored in natural language (Devlin et al., 2019; Dong et al., 2019) and speech (Schneider et al., 2019; Hsu et al., 2021). According to neuroscience, humans first pre-process speech and text with different cortices, and then extract the meaning with the same area, called the Wernicke-Geschwind area (Tremblay and Dick, 2016). Motivated by this, it is a very promising direction to design two pre-nets and a unified representation space (similar to the Wernicke area) so that the speech model would benefit greatly from text modality.
In terms of joint speech-text modeling, most approaches employ a speech encoder and a text encoder to map the speech and text inputs to hidden states, based on which, a shared encoder is used to learn cross-modality content information (Bapna et al., 2021, 2022; Chen et al., 2022b). To align the speech and text modalities, two alignment losses (TLM and STM) in SLAM (Bapna et al., 2021) are introduced with supervised ASR data. Extending SLAM to the multilingual scenario, mSLAM (Bapna et al., 2022) introduces CTC losses and uses SpanBERT (Joshi et al., 2020) to replace the BERT objective for pre-training on character-level text. Based on the RNN-T framework, Maestro (Chen et al., 2022b) learns shared representations with modality matching, duration prediction, and sequence alignment. Almost all previous work follows the same structure with a speech/text encoder and a shared encoder, however, the interface between the speech encoder and the text encoder is not well studied, which probably leads to the outputs of the two encoders in different spaces, and suffers from transfer interference and capacity dilution for the shared encoder (Bapna et al., 2021).
In this paper, we aim at unifying speech and text modalities via a well-defined interface, with which the model can benefit from additional textual data. We argue that such an interface should provide a shared semantic space for both speech and text, and preferably have strong interpretability and learnability. To this end, we explore two alternative representation spaces satisfying the above characteristics of the interface, which are based on phoneme units and hidden units. Specifically, we introduce two discrete tokenizers named phoneme-unit tokenizer and hidden-unit tokenizer. All tokenizer models are obtained with unsupervised data or a small amount of ASR data and are used offline before pre-training. With them, we convert the speech and text to a shared intermediate modality (phoneme/hidden units), and decouple the joint speech-text modeling into two sub-modules, i.e., the learning of mapping between speech/text and the discrete units. Specifically, we propose two pre-training tasks. One is Unit-based Masked Language Modeling (UMLM) trying to predict the unit tokens from the masked speech. The other one is Unit-based Connectionist Temporal Classification (UCTC) task, aiming at reconstructing the whole text sequences from the masked unit sequences. To better align the representations of speech and text, we also adopt a Random Swapping Mechanism for the UMLM task, swapping the intermediate representations of the speech and the corresponding discrete units before feeding them into the shared subsequent network.
The contributions of this paper are summarized as follows.
We propose two alternative tokenizers which can convert unlabeled speech and text into the shared discrete space and relieve the influence of modality difference.
The proposed SpeechLM (Speech and Language Model), equipped with two simple and clear learning objectives and the random swapping mechanism, can unify and simplify the cross-modal speech-text pre-training.
Experiments demonstrate that SpeechLM enhanced by textual data significantly outperforms its speech-only counterparts on various spoken language tasks, e.g., ASR, speech translation (ST), and universal representation evaluation framework SUPERB (Yang et al., 2021).
Related Work
Unlike natural language processing (NLP), speech signals are continuous, making it not straightforward to find the predictive labels for pre-training. To tackle this issue, a tokenizer, also referred to as a quantizer, is required to map continuous speech features into discrete tokens (Baevski et al., 2020a; Hsu et al., 2021; Chung et al., 2021). HuBERT (Hsu et al., 2021) is the pioneer in the exploration of predictive speech representation learning (SSL), which utilizes a k-means model on the middle layer of the Transformer as the tokenizer to convert speech into discrete tokens. Chung et al. (2021) tries to combine a contrastive loss and a masked prediction loss in a self-supervised speech representation learning framework. In addition to the unsupervised tokenizers, Wang et al. (2022a) proposes a supervision-guided tokenizer, which is an acoustic model trained on limited labeled data, and can generate frame-level aligned phonemes as the predictive targets for SSL. In contrast, our goal is to take advantage of textual data to improve speech representation learning.
Joint Speech-Text Modeling
With the rapid development of unimodal pre-training in speech and natural language processing (Devlin et al., 2019; Hsu et al., 2021), joint speech-text pre-training obtains more and more attention from research and industrial communities (Kim et al., 2021; Qian et al., 2021; Bapna et al., 2021; Ao et al., 2022; Tang et al., 2022; Zhang et al., 2022). Most previous studies (Kim et al., 2021; Qian et al., 2021; Ao et al., 2022) let speech and text share some parameters of a neural network in pre-training, however, the speech and the text are not guaranteed to lie in the same space, suffering from transfer interference and capacity dilution (Bapna et al., 2021). To alleviate this issue, SLAM (Bapna et al., 2021) and mSLAM (Bapna et al., 2022), which are most related work to our SpeechLM, leverage extra supervised speech-to-text tasks to enhance the speech-text alignment. However, these approaches still leave unpaired speech and text data modeled separately by using different pre-training targets, which might lead the model to use individual capacities to handle each modality, and can not guarantee speech and text lie in the same space.
Our work is also related to MAESTRO (Chen et al., 2022b), which learns shared representations from speech and text modalities with a modality matching algorithm in RNN-T framework, but the modality matching could be only performed on a few paired speech-text data, limiting the effectiveness of alignment learning. Unlike SLAM and MAESTRO, we utilize trained tokenizers to convert all unpaired speech and text into the same discrete space and eliminate the influence of modal difference, so that the two modalities can interact naturally via the shared interface during the pre-training. SpeechUT (Zhang et al., 2022) also leverages hidden units as the bridge of speech and text modalities, while it only works in an encoder-decoder architecture where the encoder mainly models the speech and the decoder mainly models the text. Instead, this work focuses on the encoder architecture which is more suitable for general speech representations, and explores different tokenizers and pre-training tasks.
Methods
Given unpaired speech and text data, SpeechLM is pre-trained to learn a unified representation of speech and text modalities with the help of offline discrete tokenizers. In this section, we will present the overall framework of SpeechLM, as well as the pre-training procedures and the tokenizers.
Speech and language are two different modalities with different characteristics. We explore bridging speech and text pre-training with an explicitly defined discrete representation, where speech and text could be tokenized into a shared discrete space easily. Leveraging phoneme/hidden units as the bridge between speech and text has the following advantages: First, it is easier to separately align speech and text into a shared intermediate representation than to align them directly. Second, we can make full use of additional unpaired data to improve the alignment; Thirdly, we can leverage more fine-grained alignment information, i.e., at the frame level, to facilitate joint modeling.
To achieve this goal, we implement two tokenizers for both speech and text, a phoneme-unit tokenizer and a hidden-unit tokenizer, which will be described in detail in Section 3.4. The former aims to convert speech and text into the phoneme space, while the latter converts them into an acoustic clustering space. Given a speech sample or a text sample , a tokenizer ( for speech, for text) yields a sequence of discrete units ,
where and are the lengths of the unit sequences from speech and text, respectively.
2 Model Architecture
SpeechLM consists of a Speech Transformer and a Shared Transformer, which are enhanced with the random swapping mechanism, as illustrated in Figure 1. Next, we will introduce the main modules with the input of unpaired speech and text .
Following HuBERT (Hsu et al., 2021), we use a standard Transformer (Vaswani et al., 2017) as the backbone of the Speech Transformer, equipped with relative position embedding (Shaw et al., 2018). A speech waveform is first processed into a sequence of speech features by a stack of 1-D convolutional layers. We follow HuBERT to mask the speech feature with the mask probability of 8% and the mask length of 10. Then the masked features, , are fed into the Speech Transformer for higher-level representations,
where means the layer and indicating the input. Let be the total number of layers of all Transformer modules and the Speech Transformer accounts for half, consequently, the output is .
Shared Transformer
The Shared Transformer has the same architecture with the Speech Transformer and handles two types of input with respect to speech and text. The first input is the previous output of the Speech Transformer, , and it is processed by the Shared Transformer into . The second input is the unit embedding sequence that is derived from the text tokenized units by the unit embedding layer,
It is then processed by the Shared Transformer into , where indicates the input. Consequently, and are used as the encoded representations for speech and text. For textual representations, we further employ a CTC layer (Graves et al., 2006) that converts to character-level representations.
Random Swapping Mechanism
3 Pre-Training Tasks
SpeechLM is jointly optimized by a unit-based masked language modeling task with unlabeled speech data and a unit-based connectionist temporal classification task with unlabeled text data.
The unit-based masked language modeling task is designed for speech pre-training, like HuBERT (Hsu et al., 2021) and ILS-SSL (Wang et al., 2022b). Given -layer speech representations , UMLM tries to predict the corresponding tokenized units at the masked positions. The probability of the predicted unit at position is calculated with
where is a projection matrix, is an embedding matrix, = 0.1 is the temperature coefficient, and is the set of phoneme/hidden-unit categories. Similar to ILS-SSL, the UMLM loss is computed on both the outputs of Speech Transformer () and Shared Transformer (), with the loss formulated as,
where is the corresponding speech unit at position and is the set of masked positions.
Unit-based Connectionist Temporal Classification (UCTC)
Connectionist temporal classification (CTC) (Graves et al., 2006) is first proposed to address the sequence label problem where the output is shorter than the unsegmented input sequences. Here, we take the phoneme-unit or hidden-unit sequences tokenized and upsampled from the unlabeled text as the input, and aim at recognizing the original text through the Shared Transformer and CTC layer. The input sequence is masked in the same way as the input of the speech signal. Given a text label sequence , the unit-based CTC loss is calculated as,
where is modeled by the CTC layer, whose goal is to transform the encoded unit representation into the target characters .
By taking advantage of unlabeled speech and text data, SpeechLM performs multi-task pre-training with UMUM and UCTC tasks,
where is used to control the weight of two losses. Through joint optimization and the random swapping mechanism, SpeechLM is expected to align speech and text into a unified representation.
4 Unified Tokenizers
Figure 2 shows the overview of the proposed phoneme-unit tokenizer and hidden-unit tokenizer. Besides, the tokenizers are offline models, which are used to pre-process the unlabeled speech and text data before the pre-training.
Inspired by PBERT (Wang et al., 2022a), which leverages phoneme labels as the pre-training targets, we introduce the phoneme-unit tokenizer () to discretize speech signals () as well as text sequences (). For speech data, the tokenizer is composed of an acoustic model, whose goal is to convert acoustic features into phoneme units through a weight finite-sate transducer (WFST) based decoder (Mohri et al., 2002). We implement it using the open-source Kaldi toolkithttps://github.com/kaldi-asr/kaldi with a small amount of paired ASR data and language model (LM) data, with details described in Appendix A.1 due to the space limitation. For text data, we can directly convert words into phonemes by looking up the provided lexicons. We further upsample the phoneme sequences of text by randomly repeating each phoneme many times to make sure they have similar lengths to the phoneme sequences of speech.
Hidden-Unit Tokenizer
We follow HuBERT to tokenize speech into hidden units with a k-means cluster model, called , where the clustering feature is the intermediate hidden states of the 2nd round HuBERT model. Inspired by Zhang et al. (2022), to tokenize text data into the same hidden-unit space, we propose a non-autoregressive text to hidden-unit model (), which is based on FastSpeech (Ren et al., 2019). The model consists of a text encoder, a duration model, and a unit decoder, as shown in Figure 2 (b). is trained with a small amount of text-to-unit pairs from ASR data, where the text side is the phoneme transcriptions with phoneme’s durations, and the units are tokenized from the corresponding speech by . At inference time, only consumes non-aligned phoneme sequences converted from raw text since the duration is automatically estimated.
Experiment
SpeechLM is evaluated on various spoken language tasks, including automatic speech recognition (ASR), speech translation (ST), and the universal representation evaluation benchmark SUPERB (Yang et al., 2021). According to the tokenizers, the model can be divided into SpeechLM-H and SpeechLM-P using hidden-unit and phone-unit as discrete tokens, respectively.
We use unlabeled speech data from LibriSpeech (Panayotov et al., 2015) and LibriLight (Kahn et al., 2020) to pre-train Base and Large models respectively. LibriSpeech contains 960 hours of labeled speech where the labels are not used in pre-training. LibriLight has about 60,000 hours of unlabeled speech in the same domain as LibriSpeech. The unpaired text data are from LibriSpeech LM corpushttp://www.openslr.org/11/, containing about 40M English sentences. The paired data for optimizing the tokenizers are the full LibriSpeech data in the Large setting and the 100-hour subset (train-clean-100) in the Base setting. For downstream tasks, we use LibriSpeech for ASR evaluation, and four translation directions of CoVoST-2 (Wang et al., 2020) for ST evaluation. For all tasks of SUPERB evaluation, the data details can be found in Yang et al. (2021).
2 Pre-Training Setup
The network architecture of SpeechLM follows that of HuBERT (Hsu et al., 2021) for a fair comparison. Specifically, the Base model consists of =12 Transformer layers where both the Speech Transformer and the Shared Transformer have 6 layers. The Large model doubles the number of Transformer layers. The convolutional layers downsample the input waveform to a frame rate of 20ms. The CTC layer consists of a single 1-D convolutional layer followed by a linear layer, which outputs the probabilities of text characters. All models are pre-trained on 32 GPUs for 400K steps. To align with HuBERT, the update frequency is set to 4 for Large models to simulate 128 GPUs. The batch size for the Base model is 4375 tokens after down(up)-sampling for both speech and text input, and for the Large model it is set to 2800. The text loss () is weighted by 0.1The effect of different weights () is reported in Appendix A.3.. More details about the model configuration and training details can be found in Appendix A.2.
3 Evaluation on Speech Recognition
We first verify the pre-trained SpeechLM on ASR tasks, where the Speech Transformer, the Shared Transformer, and the CTC head are fine-tuned with a speech-to-text CTC loss. Base models are fine-tuned on the train-clean-100 subset and Large models are fine-tuned on the full 960h LibriSpeech. We measure the quality of ASR by the word error rate (WER) evaluated on the standard test-clean/other sets. Table 1 shows that in the Base setting, by taking advantage of textual data, SpeechLM significantly outperforms previous models, such as wav2vec 2.0 (Baevski et al., 2020b), HuBERT (Hsu et al., 2021), and data2vec (Baevski et al., 2022). Particularly, the proposed SpeechLM obtains 26% and 12% relative WER reductions over HuBERT and data2vec on test-other set, respectively. We notice that using 400K instead of the full 40M text data is better for SpeechLM-H models, as discussed later in 4.6. Furthermore, our SpeechLM Large model achieves competitive or even better performance than previous workSLAM and MAESTRO use 2 model size, larger amount of paired data, or different inference framework (e.g., RNN-T in MAESTRO), whose results (see Appendix A.4) are not comparable with the setting in Table 1..
4 Evaluation on Speech Translation
We then evaluate SpeechLM on speech-to-text translation tasks. Following Wang et al. (2021), we use four language directions from English to German (de), Catalan (ca), Arabic (ar), and Turkish (tr) in CoVoST-2 (Wang et al., 2020). When fine-tuning, the pre-trained model serves as the encoder, followed by a randomly initialized decoder consisting of 6 Transformer layers with a model dimension of 768. We use character vocabulary for target languages in all translation tasks, and report the case-sensitive detokenized BLEU (Papineni et al., 2002) on the test set. The results are shown in Table 2, including the baselines that are fine-tuned from other pre-trained models. The numbers in brackets represent the standard deviation of three fine-tuning results. Table 2 shows that by boosting the quality of speech representation learning with textual data, SpeechLM-H and SpeechLM-P achieve comparable results in the Base setting, with 2.4 BLEU improvement over HuBERT Base. Surprisingly, the SpeechLM-P Large model substantially outperforms previous work with a smaller encoder, such as SLAM X-Large (Bapna et al., 2021).
5 Universal Representation Evaluation
We further evaluate our SpeechLM models on SUPERB (Yang et al., 2021), which is designed to provide a standard and comprehensive testbed for pre-trained models on various speech tasks, including Speaker Identification (SID), Automatic Speaker Verification (ASV), Speaker Diarization (SD), Phoneme Recognition (PR), Automatic Speech Recognition (ASR), Out-Of-Domain Automatic Speech Recognition (OOD-ASR), Keyword Spotting (KS), Query by Example Spoken Term Detection (QbE), Speech Translation (ST), Intent Classification (IC), Slot Filling (SF), Emotion Recognition (ER). These tasks can be grouped into five aspects of speech: content, speaker, semantics, and paralinguistics (ParaL). Table 3 shows the universal speech representation evaluation results. Compared to the previous self-supervised learning methods, SpeechLM achieves good performance on several content-related and semantic-related tasks, such as PR, ASR, ST, and SF. Particularly, the proposed SpeechLM-P model obtains 36% and 20% relative PER/WER reductions on PR and ASR tasks. Meanwhile, we can observe performance degradation for the speaker and paralinguistics-related tasks, especially for SpeechLM-P. It indicates that with our joint speech and text pre-training method, the model learns more about extracting the content-related information while discarding the other aspects of speech signals.
6 Analysis
To better understand the effectiveness of the proposed method, we conduct several experiments to investigate its main components, such as the random swapping mechanism, the comparison of two tokenizers, the amounts of unpaired text data, and further visualization analysis More ablations such as the effect of the speech/text pre-training ratio could be found in Appendix A.3. .
The proposed random swapping mechanism is the key component of SpeechLM to align the speech and text modalities in the same space. Here, we explore its effectiveness by removing it. As shown in lines 1-2 of Table 4, without the random swapping mechanism, the performance declines dramatically from 8.1 WER to 9.1 WER in test-other set without LM, and it confirms our suspicions.
Comparison of Two Tokenizers
To further compare the influence of the two tokenizers, we pre-train two models with only speech data, with results shown in lines 3-4 of Table 4. Lines 3-4 show that two tokenizers perform comparably for downstream ASR tasks. Moreover, we explore whether we can obtain improvement by not relying on paired speech-text data for training tokenizers. We conduct an experiment (line 6) in which the speech side predicts the HuBERT hidden units and the text side is trained with masked phoneme-to-character CTC loss. Compared to the results using pair data (line 5), the performance is degraded drastically, indicating the paired data are necessary for aligning the modalities.
Effect of Text Data Size
Since the text corpus contains up to 40M sentences which is much larger than the number of speech samples (960-hour Librispeech contains about 30K sentences), we conduct experiments to explore the effect of text data size for pre-training, by randomly sampling subsets from the original text corpus. Surprisingly, Figure 3 shows that the performance does not degrade much until the text data are reduced to 40K sentences. We speculate that the text data here are modeled at the lexical level, i.e., the transformation from phoneme/hidden units to characters, and 40K data is sufficient to build a lexicon. It is also noted that the WER of dev-other set is getting worse as the amount of text data increases for the SpeechLM-H models, while such degradation is not observed for SpeechLM-P. It is possible due to the hidden-unit tokenizer trained on 100h clean unit-to-text data, since the tokenization errors can accumulate as the amount of text data increases.
Visualization Analysis
Figure 4 illustrates the data distributions from different layers of the Shared Transformer in the SpeechLM-P Base model. The dimension is reduced to 2-D by T-SNE (Van der Maaten and Hinton, 2008). Data points are randomly sampled from unpaired speech and text samples from LibriSpeech dev-clean set. Layer=6 denotes the input. It is shown that as the layer increases, SpeechLM is able to align speech and text representations into a shared space.
Conclusion
In this work, we present SpeechLM, a text-augmented speech pre-trained model, which achieves competitive performance on various spoken language tasks, such as automatic speech recognition and speech translation. To make full use of unpaired data, we propose two alternative discrete tokenizers based on phoneme units and hidden units to tokenize speech and text into the same semantic space. With the shared interface, SpeechLM can learn better speech representations with the help of text modality. Quantitative and qualitative analyses demonstrate the superiority and effectiveness of the proposed method. For future work, we would like to advance the work by deeply integrating the language model ability and extending to natural language tasks.
Limitations
While the proposed SpeechLM achieves competitive performance on various spoken language tasks, it still has some limitations: (1) the current method needs paired data, or phoneme lexicon to build the tokenizers. The lexicon might be language-specific, which restricts the cross/multi-lingual application; (2) the effectiveness of applying SpeechLM to other speech domains (e.g., noisy, conversation-style speech) and the minimum amount of paired data required to build well-performing tokenizers need to be further investigated; (3) due to our computation limits, the performance of SpeechLM X-Large models are not explored.
Ethics Statement
This work presents a text-augmented speech pre-trained model SpeechLM. We evaluate our methods on standard benchmarks of the research community. The datasets used in this study contain LibriSpeech Panayotov et al. (2015), LibriLight Kahn et al. (2020), LibriSpeech LM Corpus Panayotov et al. (2015), and CoVoST Wang et al. (2020). And the SUPERB benchmark is from Yang et al. (2021). They are all public datasets or benchmarks that are widely used in the research community.
References
Appendix A Appendix
In the Base setting, we train a hybrid GMM-HMM ASR model (tri4b) on 100 hours of labeled LibriSpeech data following Kaldi recipe (Povey et al., 2011). To boost performance, we then use the tri4b model to decode the remaining 860 hours of speech and train the tri6b model on all the pseudo-labeled data, which is finally used for the phone-unit tokenizer. In the Large setting, we train a neural network instead of GMM with 960- hour labeled LibriSpeech data, which can boost the performance and alignment accuracy. Once the hybrid model is trained, unlabeled speech data is decoded and transduced to the best phoneme-level alignment paths. The frame shift is 10ms for the Base model setup, and 30ms for the Large model setup, respectively. We then re-sample the phonemes to a frame rate of 20ms by linear interpolation.
Phone-unit tokenizer for text
We use the 200K word-to-phone lexicon provided by LibriSpeech to convert words to phonemes, the OOV words are replaced by
Hidden-unit tokenizer for speech
We use the released HuBERT (Hsu et al., 2021) model following a K-Means model as the tokenizer for speech. The K-Means model has 500 classes with a frame rate of 50.
Hidden-unit tokenizer for text
To build a text-to-hidden-unit tokenizer, we modify FastSpeech (Ren et al., 2019) by replacing the prediction head from predicting the spectrum to predicting the probability of hidden units. Specifically, the tokenizer has 4 layers of encoders and 4 layers of decoders, with a model dimension of 256. The input to the model is a phoneme sequence converted from raw text. Upsampling is performed by a duration model between the encoder and the decoder, which predicts the length of each phoneme and repeats the phonemes before feeding them into the decoder. We train the model on LibriSpeech train-clean-100 subset for 10K steps, with a learning rate of 5e-4 and a batch size of 10K phonemes. The final model achieves 41.3 and 34.6 BLEU scores on dev-clean and dev-other.
A.2 Experimental details
The Base model has 12 Transformer layers with the attention dimension of 768 and attention heads of 12, the Large model has 24 Transformer layers with the attention dimension of 1024 and attention heads of 16. The convolutional layers have 512 channels and kernel sizes of , resulting in a downsampling rate of 320. The CTC layer is a single 1-D convolutional layer with a kernel size of 2, whose channel matches the Transformer dimension. It is then followed by a linear projection to the text characters. All models are pre-trained on 32 GPUs for 400K steps including 32K warming-up steps. We use Adam (Kingma and Ba, 2014) with =0.9, =0.98 for optimization. The maximum learning rate is set to and decays linearly to zero after the warming-up steps.
Fine-tuning configuration
For Base models fine-tuned on 100-hour LibriSpeech, the total steps are 30K with a batch size of 800 seconds. For Large models fine-tuned on the full LibriSpeech, the total steps are 200K with a batch size of 1800 seconds. All LibriSpeech models are tuned with a maximum learning rate of 1e-5 and a tri-stage learning rate schedule with the warming-up, holding, and decay periods of . And for CoVoST-2, both the Base and the Large models are fine-tuned for 50K steps with a batch size of 1600 seconds. The learning rate warms up to 1e-4 in 5K steps and then decays linearly to zero. After fine-tuning, we select the model with the best accuracy on the valid set in the Base setting and average the top 5 models with the best accuracy on the valid set in the Large setting. The decoding beam size is 5 without external language model fusion.
A.3 Analysis
Table 5 shows the fine-tuning performance of the different pre-trained models with respect to the pre-training loss ratio . It is noticed that a lower weight (0.1) of the text pre-training task achieves the best performance in the dev set. Hence, we use for all other experiments.
A.4 ASR results on 960h LibriSpeech benchmark
Table 6 lists the ASR performance of SpeechLM on the full 960h LibriSpeech benchmark comparing with SLAM (Bapna et al., 2021) and Maestro (Chen et al., 2022b). Note that they use 2 model size, a larger amount of paired data, or a different inference framework (e.g., RNN-T in MAESTRO), making the results not fairly comparable with SpeechLM.