SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, Xipeng Qiu
Introduction
Large language models (OpenAI 2023; Touvron et al. 2023) have performed astonishingly on various natural language processing tasks. Meanwhile, multi-modal large language models, such as GPT-4, PALM-E (Driess et al. 2023), and LLaVA (Liu et al. 2023), have explored the ability of LLMs to understand multi-modal information. However, a significant gap exists between current LLMs and general artificial intelligence (AGI). First, most current LLMs can only perceive and understand multi-modal content but cannot spontaneously generate multi-modal content. Second, continuous signals like images and speech cannot be adapted directly to LLMs that receive discrete tokens.
The current speech-language model mainly adopts a cascading paradigm (Huang et al. 2023a) i.e., the LLM is connected with an automatic speech recognition (ASR) model or a text-to-speech (TTS) model in tandem, or the LLM is employed as a control hub, with several speech processing models are integrated to cover multiple audio or speech tasks (Huang et al. 2023a; Shen et al. 2023). Some prior work on generative spoken language models involves encoding the speech signal into a discrete representation (Baevski et al. 2020; Hsu et al. 2021) and modeling it with language models (Lakhotia et al. 2021; Borsos et al. 2022; Zhang et al. 2023b; Wang et al. 2023).
While capable of perceiving and generating speech, the existing cascading methods or spoken language models still have several limitations. First, the LLM in the cascaded model only functions as a content generator. Since the representations of speech and text are not aligned, the LLM’s knowledge cannot be transferred to the speech modality. Second, the cascade approach (Shen et al. 2023; Huang et al. 2023a) suffers from the loss of paralinguistic signals such as emotion and prosody. Third, existing spoken language models (Wang et al. 2023; Zhang et al. 2023b) only synthesize speech but fail to comprehend its semantic information, preventing them from achieving true cross-modal perception and generation.
In this paper, we propose SpeechGPT, a large language model with intrinsic cross-modal conversational abilities, capable of perceiving and generating multi-model content. We perform speech discretization with a self-supervised trained speech model to unify the modality between speech and text. The discrete speech tokens are then expanded into the vocabulary of the LLM, thus endowing the model with an inherent competence to perceive and generate the speech.
To provide the model with the capacity to handle multi-modal instructions, we build the first speech-text cross-modal instruction-following dataset SpeechInstruct. Specifically, we discretize the speech to discrete units (Hsu et al. 2021) and construct the cross-modal unit-text pair based on the existing ASR dataset. Meanwhile, we construct hundreds of instructions for diverse tasks with GPT-4 to simulate actual user instructions as illustrated in Appendix B. In addition, to further enhance the model’s cross-modal capability, we designed the Chain-of-Modality instruction data, i.e., the model receives the speech command, thinks about the process in text, and then outputs the response in speech.
For better cross-modal transfer and efficient training, SpeechGPT undergoes a three-stage training process: modality-adaptation pre-training, cross-modal instruction fine-tuning, and chain-of-modality instruction fine-tuning. The first stage enables speech comprehension for SpeechGPT with the discrete speech unit continuation task. The second stage employs the SpeechInstruct to improve the model’s cross-modal capabilities. The third stage utilizes parameter-efficient LoRA (Hu et al. 2021) fine-tuning for further modality alignment.
To evaluate the effectiveness of SpeechGPT, we conduct a wide range of human evaluations and case analyses to estimate the performance of SpeechGPT on textual tasks, speech-text cross-modal tasks, and spoken dialogue tasks. The results demonstrate that SpeechGPT exhibits a strong ability for unimodal and cross-modal instruction following tasks as well as spoken dialogue tasks.
We build the first multi-modal large language model that can perceive and generate multi-modal contents.
We construct and release SpeechInstruct, the first large-scale speech-text cross-modal instruction-following dataset.
We build the first spoken dialogue LLM with strong human instruction following ability and spoken dialogue ability.
We show great potential to incorporate other modalities into LLMs through discrete representations.
Related Work
Multi-modal Large Language Model Current multi-modal LLMs predominantly focus on the visual domain, feeding continuous representations obtained from pre-trained visual encoders into LLMs, facilitating full-parameter or parameter-efficient training on visual-language data (OpenAI 2023; Huang et al. 2023b; Zhang et al. 2023a). Palm-E (Driess et al. 2023) integrates the 540B PaLM (Chowdhery et al. 2022) and 22B Vision Transformer (Dosovitskiy et al. 2021) into the largest vision-language model. LLaVA (Liu et al. 2023) leverages pre-trained CLIP (Radford et al. 2021) visual encoder and LLaMA (Touvron et al. 2023) and conduct instruct tuning on GPT4-assisted visual instruction data. X-LLM (Chen et al. 2023) converts multi-modalities into representations with X2L interfaces as the inputs of the large language model. However, such structures only enable LLMs to process multi-modal input, without ability to generate multi-modal output. Diverging from prior studies, our approach emphasizes the development of a speech-centric multi-modal LLM, endowing it with the proficiency to accommodate both multi-modal input and output.
Generative Spoken Language Model Discrete self-supervised representation based spoken generative language modeling is making remarkable progress on large-scale speech dataset training (Nguyen et al. 2022). AudioLM (Borsos et al. 2022) proposes to model speech based on audio codecs together with semantic codes, which can synthesize speech in a textlesss setting. VALL-E (Wang et al. 2023) builds a generative spoken language model on audio codecs and treat Text-to-Speech as a conditional generation task. However, these models are designed for a specific task and failed to benefit from LLMs. SpeechGPT is built upon the foundation of LLM and transfers LLM’s knowledge to speech modality, consequently obtaining better task generalization and human-instruction following ability.
Speech-Enabled LLM Interaction Following the emergence of ChatGPT, several studies have concentrated on the integration of expert speech models with LLMs to enable direct speech interaction with LLMs. HuggingGPT (Shen et al. 2023) facilitates task decomposition of human instructions by LLMs and allows the invocation of models from Huggingface to accomplish specific tasks, encompassing a range of automatic speech recognition (ASR) and text-to-speech models. AudioGPT (Huang et al. 2023a) leverages a variety of audio foundation models to process complex audio information and connect LLMs with input/output interface (ASR, TTS) for speech conversations. However, these models exhibit increased complexity, demand extensive resources, and are prone to the unavoidable error accumulation problems. Our approach enables speech interaction with LLMs without relying on ASR or TTS systems, circumventing the aforementioned drawbacks.
SpeechInstruct Construction
Due to the limitations in publicly available speech data and the lack of variety of speech-text tasks, we construct SpeechInstruct, a speech-text cross-modal instruction-following dataset. This dataset consists of two parts, the first part is called Cross-Modal Instruction, and the second part is called Chain-of-Modality Instruction. The construction process of SpeechInstruct is illustrated in Figure 2.
Data Collection We collect several large-scale English ASR datasets to construct Cross-Modal Instruction, including Gigaspeech (Chen et al. 2021), Common Voice (Ardila et al. 2020), and LibriSpeech (Panayotov et al. 2015). We employ mHuBERT https://dl.fbaipublicfiles.com/hubert/mhubert_base_vp_en_es_fr_it3.pt as the speech tokenizer to discretize speech data into discrete units and remove the repetitive units of adjacent frames to get reduced units. Ultimately, we obtain 9 million unit-text data pairs.
Task Description Generation We generate ASR and TTS task descriptions that are compatible with speech-text data pairs. Unlike the Self-Instruct method (Wang et al. 2022), we generate descriptions through a zero-shot approach. Specifically, we directly input the prompts shown in Appendix A into OpenAI GPT-4 to generate task descriptions. Our generation method yields 100 instructions for each task and some examples are shown in Appendix B.
Instruction Formatting For a discrete unit sequence and its associated transcription , we determine whether it will be used for constructing an ASR task or a TTS task based on the probability . Subsequently, we randomly select a description from the corresponding task description. This results in a triplet consisting of the task description, discrete unit sequence, and transcription, denoted as . Following this, the triplet is assembled into an instruction using the template: [Human]:. This is input:
2 Chain-of-Modality Instruction
Speech Instruction Generation Due to the lack of instruction data with speech input and speech output, we trained a text-to-unit generator to convert text instruction data into speech instruction data. Specifically, the text-to-unit generator adopts a Transformer encoder-decoder architecture. We trained it on LibriSpeech unit-text pairs in Cross-modal Instruction. We select 37,969 samples from the moss-002-sft-data dataset https://huggingface.co/datasets/fnlp/moss-002-sft-data whose response length is shorter than 35 words. And we convert both their instructions and responses into unit sequences through the text-to-unit generator. As a result, we obtained 37,969 quadruplets composed of speech instructions, text instructions, text responses, and speech responses, denoted as .
Instruction Formatting Using the above quadruplets, we could construct chain-of-thought style instructions for four input-output formats, namely Speech Instruction-Speech Response, Speech Instruction-Text Response, Text Instruction-Speech Response, and Text Instruction-Text Response. Their corresponding templates can be found in Appendix C.
SpeechGPT
A unified framework is designed to provide architecture compatibility across different modalities. As shown in Figure 2, our model consists of three main components: discrete unit extractor, large language modal and unit vocoder. Under this architecture, LLM can perceive multi-modal inputs and generate multi-modal outputs.
Discrete Unit Extractor The discrete unit extractor utilizes the Hidden-unit BERT (HuBERT) model (Hsu et al. 2021) to transform continuous speech signals into a sequence of discrete units, . HuBERT is a self-supervised model that learns by predicting discrete labels for masked audio segments based on k-means clustering applied to the model’s intermediate representations. It features a combination of 1-D convolutional layers and a Transformer encoder to encode speech into continuous intermediate representations, with a k-means model further converting these representations into a sequence of cluster indices. Subsequently, adjacent duplicate indices are removed, resulting in a discrete units sequence represented as , , , with denoting the total number of clusters.
Large Language Model We employ the Meta AI LLaMA (Touvron et al. 2023) model as our Large Language Model. LLaMA comprises an embedding layer, multiple transformer blocks, and an LM head layer. The total number of parameters in LLaMA ranges from 7B to 65B. Drawing from an extensive training dataset of 1.0 trillion tokens, LLaMA demonstrates competitive performance compared to the substantially larger 175B GPT-3 across various NLP benchmarks.
Unit Vocoder Due to limition of single speaker unit vocoder in (Polyak et al. 2021), we train a multi-speaker unit HiFi-GAN to decode the speech signal from the discrete representation. The HiFi-GAN architecture consists of a generator and multiple discriminators . The generator uses look-up tables (LUT) to embed discrete representations and the embedding sequences are up-sampled by a series of blocks composed of transposed convolution and a residual block with dilated layers. The speaker embedding is concatenated to each frame in the up-sampled sequence. The discriminator features a Multi-Period Discriminator (MPD) and a Multi-Scale Discriminator (MSD), which have the same architecture as (Polyak et al. 2021).
2 Training
To incorporate speech discrete representation into LLM, we expand the vocabulary and corresponding embedding matrix first. We divide the training process into three stages. The first stage is Modality-Adaptation Pre-training on unpaired speech data. The second stage is Cross-modal Instruction Fine-Tuning. The third stage is Chain-of-Modality Instruction Fine-Tuning.
Expanding Vocabulary Given original LLM vocabulary of size , to integrate speech discrete representations into LLM, we expand the vocabulary with an additional set of unit tokens , of size . The expanded vocabulary is the union of the original vocabulary and the new words :
Finally, we replace the original vocabulary and word embedding matrix with the new vocabulary and the word embedding matrix .
Stage 1: Modality-Adaptation Pre-training To enable LLM to handle discrete units modality, we utilize an unlabeled speech corpus to train LLM in a next-token prediction task. This approach aligns with the text pre-training objective of LLM. Given unlabeled speech corpus consisting of speech and LLM denoted as , the negative log-likelihood loss can be formulated as:
where is the number of speech in dataset , is the number of discrete unit token in speech , and represents the i-th unit token in the j-th speech.
Stage 2: Cross-modal Instruction Fine-Tuning In this stage, we align speech and text modalities utilizing paired data. We mix Cross-modal Instruction in SpeechInstruct with moss-002-sft dataset to derive mix dataset , which consists of samples . We fine-tune the model obtained from the first stage on .
Each sample consisting of is formed by concatenating a prefix and a text. The training objective is to minimize the negative log-likelihood and the loss calculation only considers the text part, ignoring the prefix, which can be formated as:
where is the number of samples in corpus , is the total number of tokens in sample , is the number of tokens in the prefix part of , and represents the i-th word in .
Stage 3: Chain-of-Modality Instruction Fine-Tuning After obtaining the model in stage 2, we utilizes parameter-efficient Low-Rank Adaptation (LoRA) (Hu et al. 2021) to fine-tune it on Chain-of-Modality Instruction in SpeechInstruct. We add LoRA weights (adapters) to the attention mechanisms and train the newly added LoRA parameters. We adopt the same loss function as stage 2.
Experiments
Datasets For modality-adaption pre-training, we use LibriLight (Kahn et al. 2020) which contains 60K hours of unlabelled English audiobook speech. For cross-modal instruction fine-tuning stage, we use Gigaspeech (Chen et al. 2021), Common voice (Ardila et al. 2020) and LibriSpeech (Panayotov et al. 2015) dataset and moss-002-sft-data dataset, which is illustrated in detail in 3.1. For chain-of-modality instruction fine-tuning stage, we use moss-002-sft-data dataset, which is illustrated in detail in 3.2.
Configuration We employ LLaMA-13B (Touvron et al. 2023) as our backbone model. For stage 1, we use 96 A100 gpu and train for 900 steps with batch size 768. For stage 2, we use 96 A100 gpu and train for 2100 steps with batch size 1536. For stage 3, we use 8 A100 gpu and train for 4200 steps with batch size 128. Details about training hyperparameters are shown in Appendix 3. For decoding, we set the maximum sequence length to 2048 and set the temperature to 0.8. We use Top- sampling with =60. We also use Top- sampling with p=0.8.
Evaluation We evaluate the capabilities of SpeechGPT in two aspects: cross-modal instruction following ability and spoken dialogue ability. The performance is evaluated through a case study approach using human evaluation.
2 Main Results
Cross-modal Instruction Following As shown in Table 1, when provided with various instructions, the model is capable of performing corresponding tasks and generating accurate outputs in accordance with these inputs.
Spoken Dialogue Table 2 shows 10 cases of speeech dialogue of SpeechGPT. The dialogue shows that in interactions with humans, SpeechGPT is capable of comprehending speech instructions and responding accordingly in speech, while adhering to the HHH criteria (Harmless, Helpful, Honest) (Askell et al. 2021).
Limitation
Despite SpeechGPT exhibiting impressive cross-modal instruction following and speech dialogue abilities, it still presents certain limitations: 1) It does not consider paralinguistic information in speech, such as the inability to generate responses in different emotional tones, 2) It necessitates the generation of a text-based response prior to the production of a speech-based one, 3) Due to the context length limitation, it is incapable of supporting multi-turn dialogues.
Conclusion
This work presents SpeechGPT, an inherent cross-modal multimodal large language model capable of perceiving and generating multimodal contents. In addition, to alleviate the scarcity of instruction datasets in the current speech domain, we propose SpeechInstruct. This first speech-text cross-modal instruction-following dataset contains cross-modal instruction data and spoken dialogue data based on the chain-of-modality mechanism. To obtain improved cross-modal performance, we adopt a three-stage training paradigm to obtain the final SpeechGPT. Experimental results indicate that SpeechGPT achieves promising results in various unimodal or cross-modal tasks and demonstrate that combining discrete speech tokens into the language model is a promising direction.