Connecting Speech Encoder and Large Language Model for ASR
Wenyi Yu, Changli Tang, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, Chao Zhang
INTRODUCTION
Large language models (LLMs) with rich knowledge and the ability to solve novel and complex tasks have revolutionised the field of natural language processing. More recently, significant attention has been drawn to enable LLMs to handle speech inputs . In addition to pipeline-based methods in which the LLM serve as a controller to manage a set of functional models , two categories of approaches have been developed. The first category of approaches discretises speech inputs and embeds the derived speech tokens into a vector space shared with the text tokens, then the LLMs are finetuned to fit into this new token space . The other category of approaches directly connects a speech encoder with an LLM using a connector that aligns the speech encoders with the LLMs . This paper focuses on the second category of approaches.
When aligning the speech encoder output and LLM input spaces, the choice of the connector is of vital importance, and should meet the following requirements. First, the connector should be able to retain as much information from the speech inputs as possible, since it determines the amount of information that LLMs can receive from the speech. Second, as the computation and storage costs of LLMs increase considerably when processing long input sequences, the connector should be able to achieve efficient and effective information compression to reduce the lengths of the LLM input sequences.
Focusing on automatic speech recognition (ASR) to show the ability of LLMs to recognise speech, this paper studies different structures of connectors that integrate LLMs with speech encoders. Specifically, an end-to-end ASR system was constructed by connecting a Whisper model encoder with a Vicuna LLM . Three types of connectors were compared in this paper, including the fully connected layers, multi-head cross-attention and Q-Former . To bypass the input length limitation of the pre-trained speech encoders and enable the LLM to process long speech inputs, a novel segment-level Q-Former connector is proposed. Experiments were conducted based on a training set with 4,000 hours data, and the LLMs with Q-Former connectors consistently outperform strong Whisper baseline ASR systems on the in-domain datasets, and can achieve over 12% relative word error rate (WER) reductions on the out-of-domain Eval2000 test set from the Switchboard corpus. The influence of the number of connector output tokens and model size are studied. Moreover, the proposed segment-level Q-Former structure achieved obvious WER reductions on long speech inputs, when compared with other connectors.
The rest paper is organised as follows. Sec. 2 summarises the work related to multimodal LLMs. Secs. 3 and 4 introduce the three connectors to compare and the proposed segment-level Q-Former. The experimental setup and results are presented in Secs. 5 and 6, followed by conclusions.
RELATED WORK
To enable LLMs to perform both speech perception and generation, SpeechGPT and AudioPaLM augment the vocabularies of LLaMA and PaLM LLMs with discrete speech tokens extracted by HuBERT and W2v-BERT or USM speech encoders respectively. Regarding the approaches to connect multimodal encoders to LLMs, X-LLM interfaced ChatGLM with audio and visual encoders. and connect speech encoders with reduced frame rates to LLMs, and achieve integrated multilingual ASR and speech translation respectively. Moreover, LLMs can also be prompt for domain adaptation and uncertainty estimation of the ASR results.
Several works studied visual LLMs . Following BLIP-2 , InstructBLIP and Video-LLaMA introduced Q-Former as the module connector. Alternative connectors were also investigated in . Regarding audio and music LLMs, the reasoning ability based on audio is studied . MU-LLaMA showed outstanding music understanding abilities.
MODULE CONNECTOR
As shown in Fig. 1, the proposed ASR model consists of three modules: a frozen speech encoder, a trainable module connector and a frozen LLM. This section introduces three connectors, including fully connected layers, multi-head cross-attention and Q-Former.
where consists of of a batch of samples. Actually, the vector stacking operation together with the first linear layer works the same as a 1-dimensional (-d) convolutional layer, .
2 Multi-head cross-attention
To bridge the gap between the multi-modal encoder output features and LLM input textual features , a multi-head attention layer denoted as is used in the multi-head cross-attention approach to align the two feature spaces . First, a layer reduces the length of the speech input by a rate of . Then the hidden states are converted to based on the textual embeddings using . That is,
3 Q-Former
SEGMENT-LEVEL Q-FORMER
Transformer-based speech encoders can have limitations on the input sequence duration . To enable LLMs to process with longer speech inputs, the whole sequence can be split into several shorter segments to transform by the speech encoder separately. Such segments can be concatenated to reform a single sequence at either the input or output end of Q-Former. In this paper, the structure of segment-level Q-Former (seg-QF) shown in Fig. 2 is proposed, which uses a Q-Former to transform each encoder output segment simultaneously and concatenates their fixed-length output token sequences before feeding into the LLM. Compared to performing the concatenation at the Q-Former input end and producing a fixed number of output tokens, seg-QF allows varying the number of output tokens according to the number of segments , which is more suitable for speech inputs with variable lengths in a wide range. Note the trainable query embeddings and Q-Former layers are shared among all the segments, and seg-QF can be initialised with a pre-trained standard Q-Former.
where means adding to each row of .
EXPERIMENTAL SETUP
Experiments were conducted on three setups. Specifically, models in Sections 6.1-6.4 were trained on LibriSpeech 960h dataset. Models in Section 6.5 were trained on around 4,000 hours of data including LibriSpeech 960h, Common Voice 7.0 (English) and GigaSpeech subset M. Models in Section 6.6 were finetuned on LibriSpeech train-clean-100 subset. While the test sets of the aforementioned data were used for in-domain evaluation, the Eval2000 set was used for out-of-domain evaluation.
2 Model specifications
The Vicuna LLMs and the Speech encoders from the Whisper models were used as the decoders and encoders , and were both frozen in training. In Section 6.4, speech encoders with different sizes were compared, including those from Whisper base, medium and large-v2. LLMs including Vicuna 7B and Vicuna 13B were compared. Based on Table 2, the best-performing setup with Whisper large-v2 encoder and Vicuna 13B were used in other sections.
All the models in Sections 6.1–6.5 were trained for 90k steps with a batch size of 24 with NVidia A100 80GB GPUs. Models in Section 6.6 were initialised with pre-trained Q-Former models or fully connected layers and trained for 10k steps with a batch size of 8. Sinusoid encoding matrix was used as the segment-level position embeddings. Checkpoints with the highest validation set accuracy obtained with teacher forcing were selected as the final models.
EXPERIMENTAL RESULTS
The Whisper encoder is built to have a fixed input window of 30 seconds, and zero vectors are padded to the input to increase its sequence length to match the window size. When connecting the Whisper encoder to LLMs using Q-Former and this padding strategy, high deletion errors are often produced for long inputs since the attention mechanisms of Q-Former are trained to ignore the ending elements in the sequence, which are often padded zeros in the training samples. To resolve this issue, a random concatenation strategy is applied, which is similar to that used in . Each input utterance in a training mini-batch is concatenated with a number of utterances randomly selected from the whole training set. The random concatenation continues as long as the total length of the concatenated utterance does not exceed a pre-set upper limit of seconds. The values of for the training samples are set to follow a uniform distribution of seconds. From the results in Table 1, it is demonstrated the strategy can considerably reduce deletion errors and improve the WERs, and therefore it is always used in the rest of the experiments.
2 Comparisons with different connectors
A desirable connector should be able to extract all useful information from the input without causing obvious increases in computation and storage. Regarding ASR, WERs can be used as an indicator to reflect the quality of the information extracted by the connector. Besides having a reasonable amount of model parameters, the connector is also expected to produce a reduced number of output tokens when required since the number of LLM input tokens is a key factor influencing both LLM computation cost and memory usage.
In Table 2, three connectors, including fully connected layers, multi-head cross-attention and Q-Former, are compared. Q-Former results in the lowest WERs on both test sets by producing only 80 output tokens. Although fully connected layers with 300 output tokens (with in Sec. 3.1) can achieve similar WERs to Q-Former, it requires much more calculations and memory usage. The WERs produced by the fully connected layers connector with 75 output tokens (with ) are obviously worse. The multi-head cross-attention connector has 133.4 million (M) model parameters, which are 6 times more than the others, but still produced the worst WERs. As a result, Q-Former is used in the rest of the study.
3 Trainable queries of Q-Former
Recapping Section 3.3, the number of trainable queries used in the Q-Former determines the number of its output tokens. This section compares Q-Formers with different numbers of queries. Table 3 shows that increasing the number of queries to 80 can considerably reduce WERs, which implies that using 80 tokens can retain sufficient information for ASR for inputs with less than 30 seconds (similar WER changes are found when only considering the utterances whose durations are close to 30 seconds), which reveals the strong information compression ability of Q-Former.
4 The sizes of LLMs and speech encoders
In this section, models with LLMs and speech encoders of different sizes are compared in this section. Q-Former connectors with 80 trainable queries are used for all models. From the results Table 4, increasing the sizes of both speech encoder and LLM can result in lower WERs, which is in line with expectations. Doubling the size of the speech encoder (Whisper medium to Whisper large) reduces the WERs more obviously than doubling the size of LLMs, which indicates that a stronger speech encoder is more important for ASR, which does not require much content understanding ability.
5 Results with large-scale training set
This section verifies the performance of the best-performing model configuration by training on a large-scale dataset with 4,000 hours of speech. A Vicuna 13B LLM is connected with a Whisper large-v2 encoder using a Q-Former connector with 80 trainable queries. The results in Table 5 show that the proposed model outperformed Whisper large-v2 by an average of 11% relative WER reduction on the three in-domain test sets. It also generalises well to out-of-domain data by achieving a 12% relative WER reduction over Whisper large-v2 on the Eval2000 test sets. These results verify the superior performance of the proposed model in ASR.
6 Results of segment-level Q-Former
To evaluate the models when recognising speech exceeding the duration limitation of the speech encoder, concatenated long-form test sets were built based on LibriSpeech test-clean and test-other test sets respectively. Utterances of the same chapter were concatenated sequentially to form sequences with a duration limit of 60, 90, or 120 seconds. The models were trained on LibriSpeech 100h subset using random concatenation with sampled from .
Different models were compared in Table 6. Without finetuning on data longer than 30 seconds, QF and FC cannot generalise to recognise longer speech. After finetuning, seg-QF outperformed both QF and FC considerably and showed some generalisation to 120-second-long speech inputs. From Table 6, WERs degrade with longer speech inputs. More extreme cases, such as repeated outputs and large chunks of deletions, were observed in case studies when feeding long speech inputs. The WERs for all utterances and those successfully decoded in Librispeech test-other are plotted in Fig. 3 for Seg-QF and Seg-QF*. It shows that the contexts in long speech inputs help ASR to reduce WERs if being decoded successfully. However, the overall WERs on the full test-other set increased due to more frequent extreme cases.
CONCLUSION
This paper studies to enable LLMs to recognise speech inputs by interfacing with a speech encoder. Three commonly used connectors including fully-connected layers, multi-head cross-attention and Q-Former were compared. The LLMs with Q-Formers demonstrated superior performance over LLMs with other connectors and the Whisper baseline ASR system on all of the in-domain and out-of-domain test sets. Moreover, a novel segment-level Q-Former was proposed to improve the performance with long-form speech inputs whose duration exceeds the limitations of the pre-trained speech encoder. Analyses show that although the rich context information in long-form speech inputs can improve ASR accuracy, overly long inputs can also aggravate the hallucination problem of LLMs.