AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang, Xipeng Qiu
Introduction
LLMs have exhibited remarkable proficiency in comprehending and generating human language. Nevertheless, their capabilities are confined to textual processing. The real-world environment is inherently multimodal, with organisms perceiving and exchanging information through diverse channels, including vision, language, sound, and touch.
A promising objective in developing multimodal systems is to augment LLMs with the capacity for multimodal perception. The dominant methodology involves the integration of multimodal encoders with the language model, thus empowering it to process information across various modalities and utilize its sophisticated text-processing abilities to produce coherent responses. However, this strategy is limited to text generation and does not encompass multimodal output.
Pioneering efforts such as Emu (Sun et al., 2023b), SEED-LLaMA (Ge et al., 2023b) and SpeechGPT (Zhang et al., 2023a) have made significant strides by enabling multimodal understanding and generation within language models. Yet, these models incorporate only a single non-textual modality, such as images or audio. While aligning text with one additional modality is relatively straightforward, integrating multiple modalities () within a single framework—and achieving bidirectional alignment among them—poses a more formidable challenge.
Existing explorations in any-to-any multimodal generation have encountered obstacles: some (Tang et al., 2023b) lacked a robust core language model, which impeded the system’s reasoning and decision-making capabilities; Others, such as NExT-GPT (Wu et al., 2023), CoDi-2 (Tang et al., 2023a), and Unified-IO2 (Lu et al., 2023), have employed separately pre-trained encoders and decoders. This approach results in representational inconsistencies between the inputs and outputs of the LLMs, which in turn complicates both training and inference processes. Moreover, stabilizing training with such diverse modalities necessitates substantial modifications to existing models and techniques.
To overcome these challenges, we introduce AnyGPT, an any-to-any multimodal language model that employs discrete representations for unified processing. AnyGPT is equipped with multimodal tokenizers that compress raw multimodal data, such as images and audio, into a sequence of discrete semantic tokens. These discrete representations enable the core LLM to unify tasks such as perception, understanding, reasoning, and generation in an autoregressive manner at the semantic level. Subsequently, de-tokenizers convert the discrete representations back into the original modal representations at the perceptual level. Thanks to discrete representation, which filters out high-frequency, modality-specific perceptual information while preserving essential low-frequency semantic information (Ge et al., 2023a; Borsos et al., 2023a; Rombach et al., 2022), we can train our model stably without any alterations to the existing LLM architecture or training paradigms. Instead, our approach relies solely on data-level preprocessing. This allows for the seamless integration of new modalities into LLMs, akin to the addition of new languages, and permits the direct application of existing LLM tools, which enhances the efficiency of both the training and inference stages.
Furthermore, to mitigate the scarcity of multimodal alignment data encompassing all modalities, we build a text-centric multimodal alignment dataset for pre-training. Our goal is to use text as a bridge, by aligning other modalities with text, to achieve mutual alignment among all modalities, since natural language is the most refined modality of semantic representation and is present in the majority of multimodal alignment datasets. To endow the model with the capability to comprehend and generate content interwoven with multiple modalities, we employ advanced generative models to synthesize a multimodal instruction dataset, AnyInstruct-108k. This dataset, comprising 108k samples of multi-turn conversations, enables AnyGPT to handle arbitrary combinations of multimodal inputs and outputs.
Experimental results demonstrate that AnyGPT achieves zero-shot performance comparable to that of specialized models across various modalities. Furthermore, extensive case studies corroborate AnyGPT’s remarkable ability to facilitate any-to-any multimodal dialogue, substantiating the feasibility of using discrete representations to unify multiple modalities.
We propose AnyGPT, a token-based any-to-any multimodal language model which can understand and generate various modalities, including speech, text, images, and music.
One key challenge is the lack of multimodal interleaved instruction-following data. We develop a pipeline using generative models to build AnyInstruct-108k, a dataset comprising 108k multi-turn dialogues with interleaved multimodal elements.
We demonstrate discrete representations can effectively unify multiple modalities within a language model.
Related Work
To enable cross-modal perception in LLM, a common approach is to connect pre-trained encoders of other modalities as adaptors. However, these models are often limited to text generation.
To empower LLMs with multimodal generation capabilities, Tang et al. (2023b) introduces a frozen text-to-image diffusion model and learns the mapping between the LLM’s embeddings and the diffusion model. Sun et al. (2023a) utilizes continuous embeddings to represent the image, calculating either a loss for the next token prediction or the next visual embedding regression. In contrast, SEED-LLaMA (Ge et al., 2023b) trains an image discretization tokenizer to encode the original image into discrete tokens. Through a unified next token prediction task, it achieves unified image understanding and generation. Similarly, in the field of speech, SpeechGPT (Zhang et al., 2023a) enables LLMs to have inherent cross-modal conversation capabilities through discrete speech representation.
To achieve multimodal generation across various modalities on LLMs, NExT-GPT (Wu et al., 2023) utilizes existing high-performance encoders and decoders, connected by a small number of projection layer parameters. However, NExT-GPT does not train the LLM, which may result in suboptimal performance. Moreover, its representation of multimodal input and output lacks a unified form, which poses challenges in unified training and inference.
2 Multimodal Discretization
To create a unified multimodal language model, a common approach is to use discretization. A typical method is VQ-VAE (van den Oord et al., 2017). This involves maximizing the restoration of the original representation from the compressed tokens. Some studies (D’efossez et al., 2022; Zeghidour et al., 2021) incorporate residual quantization mechanisms to further enhance fidelity.
In addition to VQ-VAE based tokenizers, some tokenizers focus on extracting high-level semantic representations. Ge et al. (2023b) discretizes the image into semantic-level. The SpeechTokenizer (Zhang et al., 2023b), based on the RVQ-VAE structure, enables the first layer of tokens to retain the semantic information of speech, and the remaining layers to supplement residual information information, achieving a disentanglement of semantic and acoustic information.
AnyGPT
Our interest lies in facilitating the generation of any modality to any modality with LLMs. To realize this, we propose a comprehensive framework that can be uniformly trained. As illustrated in Figure 1, this framework is composed of three main components: () multimodal tokenizers, () a multimodal language model serving as the backbone, and () multimodal de-tokenizers. The tokenizers transform continuous non-text modalities into discrete tokens, which are subsequently arranged into a multimodal interleaved sequence. Then the sequences are trained by the language model using the next token prediction training objective. During the inference process, multimodal tokens are decoded back into their original representations by the associated de-tokenizers. To enrich the quality of generation, multimodal enhancement modules can be deployed to post-process the generated results, including applications like voice cloning or image super-resolution. In the following section, we will introduce the details of each module.
We utilize the SEED tokenizer (Ge et al., 2023a) for image tokenization. The SEED tokenizer consists of several components, including a ViT encoder (Dosovitskiy et al., 2021), Causal Q-Former, VQ Codebook (van den Oord et al., 2017), multi-layer perceptron (MLP), and a UNet decoder (Ronneberger et al., 2015). SEED takes a RGB image as input, and the ViT encoder encodes the image into patches, then the Causal Q-Former converts the patch features into 32 causal embeddings. A codebook with 8192 entries discretizes the embeddings into a sequence of quantized codes. An MLP is employed to decode the visual codes into a generation embedding, which is aligned with the latent space of the pre-trained unCLIP Stable Diffusion(unCLIP-SD) (Rombach et al., 2022). Finally, the UNet decoder is used to restore the generation embedding to the original image.
The tokenizer for speech we utilize is SpeechTokenizer (Zhang et al., 2023b), adopting an encoder-decoder architecture with residual vector quantization (RVQ). The SpeechTokenizer compresses single-channel audio sequences into a discretized matrix using eight hierarchical quantizers, each with 1,024 entries, and achieves a frame rate of 50 Hz. The first quantizer layer captures semantic content, while layers 2 to 8 encode paralinguistic details. A 10-second audio is thus transformed into a matrix, splitting into semantic and acoustic tokens. We adopt a SpeechTokenizer variant pre-trained on the Commonvoice (Ardila et al., 2020) and Librispeech (Panayotov et al., 2015) datasets.
In AnyGPT, the Large Language Model (LLM) is employed to model the semantic tokens, while a voice cloning model supplements the remaining paralinguistic information. As a result, the size of the voice vocabulary in the LLM is equivalent to the size of one codebook, which is 1024. Further details will be discussed on in Section 3.3.
Although speech and music share similar data formats, their substantial content differences lead us to treat them as distinct modalities, each equipped with its own tokenizer. For music, we employ Encodec (D’efossez et al., 2022), a convolutional auto-encoder with a latent space quantized using Residual Vector Quantization (RVQ), as the music tokenizer. We use an available off-the-shelf variant of the Encodechttps://huggingface.co/facebook/encodec_32khz pre-trained on 20k pieces of music tracks. This variant processes 32 kHz monophonic audio, and achieves a frame rate of 50 Hz. The embeddings generated are quantized using an RVQ with four quantizers, each with a codebook size of 2048, resulting in a combined music vocabulary size of 8192.
We encode 5 seconds music into 250 latent frames, ultimately generating a codes matrix. To enable the language model predict entire music clip, we flatten the 4-layer music codes into a causal sequence in a frame-by-frame manner. The language model begins by predicting the initial four tokens of the first frame and continues in a similar fashion for the subsequent frames.
2 Language Model Backbone
To incorporate multimodal discrete representations into pre-trained LLMs, we expand the vocabulary with new modality-specific tokens, and consequently extend the corresponding embeddings and prediction layer, the newly incorporated parameters are initialized randomly. The tokens from all modalities combine to form a new vocabulary, where each modality is trained within the language model to align in a shared representational space. The size of this enhanced vocabulary, denoted by , is the summation of the vocabulary sizes across all modalities, that is, , where signifies the vocabulary size of the -th modality.
Equipped with the modality-specific tokenizers, we can compress multimodal data into discrete token sequences, which can be trained by the language model using the next token prediction loss. This naturally enables the core LLM to unify tasks such as perception, understanding, reasoning, and generation in an autoregressive manner.
We employ the LLaMA-2 (Touvron et al., 2023) 7B as the backbone, which is pre-trained on 2 TB of text tokens. Apart from reshaping the embedding matrix and prediction layer, the rest of the language model remains unaltered.
3 Multimodal Generation
The generation of high-quality multimodal data, including high-definition images, and high-fidelity audio, presents a substantial challenge. These data typically necessitate a large number of bits for accurate representation, resulting in long sequences which is particularly demanding for language models, as the computational complexity increases exponentially with the length of the sequence.
To tackle this, we adopt a two-stage framework for high-fidelity generation, comprising semantic information modeling and perceptual information modeling. First, the language model is tasked with generating content that has undergone fusion and alignment at the semantic level. Then, non-autoregressive models convert multimodal semantic tokens into high-fidelity multimodal content at the perceptual level, striking a balance between performance and efficiency.
Specifically, we employ SEED tokens, aligned with the diffusion latent space, for visual language modeling. Semantic-level SEED tokens are decoded into high-quality images by a Diffusion Model, which is renowned for its superior generation capabilities. For speech, we utilize SoundStorm (Borsos et al., 2023b), a non-autoregressive Masked Language Model, trained to generate SpeechTokenizer’s acoustic tokens from semantic tokens. We train a variant of Soundstorm, which is trained using the SpeechTokenizer on the Multilingual LibriSpeech(MLS) dataset (Pratap et al., 2020). Subsequently, the SpeechTokenizer’s decoder transforms all speech tokens into raw audio data. This approach enables AnyGPT replicate the voice of any speaker using a 3-second speech prompt, while significantly reducing the length of the voice sequence for LLM. For music, we employ Encodec tokens to filter out high-frequency details beyond human perception, and then use the Encodec decoder to reconstruct these tokens into high-fidelity audio data.
Multimodal Data
To enable the generation from any modality to any other, it is crucial to have data that is well-aligned across these modalities. Unfortunately, such data is notably scarce. To address this challenge, we build a text-centric bi-modal alignment dataset. Here, text is employed as a vital intermediary to bridge the gap between various modalities. By aligning different modalities with the textual modality within a language model, we aim to achieve mutual alignment amongst all modalities.
The representational forms and types of information vary greatly across different modalities, To facilitate a standardized comparison of data volumes across various modalities, we have adopted a quantification approach based on token counts. Figure 2 presents all the data used in pre-training and their respective proportions. A certain level of oversampling is applied to modalities with comparatively lower data quantities, to attain a balanced representation of diverse data types within a single batch. More details are in Appendix 7.
We utilized image-text pairs from LAION-2B (Schuhmann et al., 2022), LAION-COCO (lai, 2022b), LAION-Aesthetics (lai, 2022a) and JouneyDB (Pan et al., 2023). LAION-2B provides images paired with noisy alt-texts sourced from the web, while LAION-COCO represents a 600M subset of this, captioned by BLIP. We refined these datasets by filtering for text quality, image aspect ratio, and clip score, etc., yielding a high-quality corpus of 300M pairs. To enhance the overall image generation fidelity, we supplement our data with the high-quality LAION-Aesthetics subset and the synthetic dataset JourneyDB from Midjourney.
We also incorporate image-text interleaved data to adapt the model to an interleaved mode. We deploy the Multimodal-C4 (MMC4) dataset (Zhu et al., 2023), an enhanced version of the text-only C4 (Raffel et al., 2020). Specifically, we utilize the MMC4-core split, consisting of 7.3M documents.
We collect several large-scale English Automatic Speech Recognition (ASR) datasets, including Gigaspeech (Chen et al., 2021), Common Voice (Ardila et al., 2020), and Multilingual LibriSpeech(MLS) (Pratap et al., 2020). These datasets are sourced respectively from online platforms, volunteer crowdsourcing, and audiobooks, collectively constituting a corpus of 57,000 hours of speech-text pairs, encompassing a wide variety of speakers, domains, and recording environments.
We embark on an extensive data collection process by crawling over one million music videos from the Internet. The core step involves matching the titles of these videos with corresponding songs using the Spotify API. Subsequently, we harvest a comprehensive set of metadata for each music audio, including video titles, descriptions, keywords, playlist names, and Spotify lyrics. This metadata is formatted into JSON and fed into GPT-4 (Achiam et al., 2023) for processing. GPT-4’s role is pivotal as an intelligent caption generator; it utilizes the noisy metadata to extract meaningful information and succinctly summarize it into coherent sentences. This approach allows us to generate high-quality text captions for a large amount of music audio, effectively minimizing the occurrence of hallucinations in the dataset.
To train the Language Model (LM), we employ various templates to construct multimodal sentences, which the LM then processes autoregressively. Further training details can be found in Appendix A.2. Additionally, We observe significant variation in sentence lengths across different modalities and datasets. To enhance training efficiency, samples from the same dataset are concatenated into a long sequence, adhering to the model’s maximum sequence length. Consequently, each token in the sequence contributes to the loss.
2 Multimodal Interleaved Instruction Data Construction
Effective human-machine interaction should permit the exchange of information in a variety of interleaved modalities. However, the increasing number of modalities in conversation significantly complicates the data collection process. To our knowledge, there is currently no large-scale instruction dataset involving more than two modalities. This poses a significant limitation on the development of a comprehensive model capable of managing dialogues with multiple, intertwined modalities.
To overcome this limitation, we draw inspiration from the most recent research on data synthesis (Wang et al., 2022; Wu et al., 2023), and build a dataset comprised of 108k multi-turn conversation samples with generative models. With careful curation, each synthetic conversation integrates multiple modalities—text, speech, images, and music—in an interleaved manner. Specifically, our data synthesis process is carried out in two stages, as illustrated in Figure 3.
In this phase, we employ GPT-4 to generate a series of text-based conversations. Notably, we incorporate non-text modality in the form of their textual descriptions within these conversations. To ensure high-quality data at scale, we divide this stage into three steps. (1) Initially, we brainstorm 100 meta topics to cover a broad spectrum of scenarios related to audiovisual elements and we employ GPT-4 to expand these meta-topics into 20,000 specific topics. (2) Subsequently, we prompt LLM to generate specific dialogue scenarios based on these topics. Acknowledging the intrinsic constraints of a text-based LLM in generating multimodal elements, we prepare several demonstrations that encompass as many modality combinations as possible. While generating scenarios, a subset is sampled from this demonstration pool, serving as examples for the LLM. This approach guides the model to effectively synthesize varied and contextually appropriate conversational scenarios. (3) Finally, we utilize GPT-4 to generate multi-turn conversations derived from scenarios. In these synthesized dialogues, multimodal elements, including images and music, are depicted through detailed textual representations. We curate a diverse range of conversation examples, similar to scenario generation, to prompt the model into creating dialogues with the widest possible variety of modalities. As a result, we compiled a substantial corpus of multimodal conversational data in solely textual format.
In this phase, we employ advanced generative models to convert textual descriptions into multimodal elements. We use OpenAI’s DALL-E-3 (Betker et al., 2023) for image generation, MusicGen (Copet et al., 2023) for music composition, and Microsoft Azure’s text-to-speech API (Microsoft, ) for speech synthesis from user’s instructions and model’s text responses.
After filtering, we obtain a dataset of 108k high-quality multimodal dialogues, featuring a variety of multimodal combinations. This dataset includes around 205k images, 503k voice recordings, and 113k music tracks. Additionally, we enhanced our dataset by extracting dialogues from existing text-only instruction datasets well-suited for spoken narration. This results in 100k voice dialogues through the employment of text-to-speech models.
The two-stage approach efficiently collected a diverse array of high-quality multimodal conversations at scale. Appendix D provides the prompts used during the data synthesis process.
Experiment
We evaluate the fundamental capabilities of the pre-trained base AnyGPT (Section 3), covering multimodal understanding and generation tasks for all modalities. This evaluation aimed to test the alignment between different modalities during the pre-training process. Specifically, we test both text-to-X and X-to-text tasks for each modality, where X is image, music, and speech separately.
To simulate real-world scenarios, all evaluations are conducted in a zero-shot mode. This means that AnyGPT will be not fine-tuned nor pre-trained on downstream training samples during the evaluation process. This challenging evaluation setting requires the model to generalize to an unknown test distribution, showcasing the generalist abilities of AnyGPT across different modalities. The evaluation results demonstrate that AnyGPT, as a generalist multimodal language model, achieves commendable performance on various multimodal understanding and generation tasks.
We assess the image comprehension capabilities of AnyGPT on the image captioning task. The comparison results are presented in Table 2. We utilize the MS-COCO 2014 captioning benchmark (Lin et al., 2014) and adopt the Karpathy split testset following previous studies (Li et al., 2023; Tang et al., 2023b).
The results of the text-to-image generation task are presented in Table 3. To ensure consistency with previous research (Koh et al., 2023; Ge et al., 2023b; Sun et al., 2023a), we randomly select 30k images from the MS-COCO validation set and use CLIPscore as the evaluation criterion. This metric computes a similarity score between the generated image and its corresponding caption from a real image, based on CLIP-ViT-L (Radford et al., 2021).
1.2 Speech
We evaluate the performance of AnyGPT on the Automatic Speech Recognition (ASR) task by calculating the Word Error Rate (WER) on the test-clean subsets of the LibriSpeech dataset (Panayotov et al., 2015). We use Wav2vec 2.0 and Whisper Large V2 as baselines. Wav2vec 2.0 is pre-trained with 60,000 hours of speech and fine-tuned on LibriSpeech, while Whisper Large V2 is evaluated in a zero-shot setting but is trained with 680,000 hours of speech. The results are shown in Table 4.
We conduct a zero-shot Text-to-Speech (TTS) evaluation on the VCTK dataset. The results are presented in Table 5. We evaluate the TTS systems with speaker similarity and Word Error Rate (WER), where WER is focused on speech quality. More experimental details can be found in Appendix C.
1.3 Music
we evaluate AnyGPT’s performance on the MusicCaps benchmark (Agostinelli et al., 2023) for both music understanding and generation tasks. We utilize the CLAPscore (Wu et al., 2022; Huang et al., 2023) score as the objective metric, which measures the similarity between the generated music and a textual description.
For the evaluation of music captioning, we found that existing objective metrics may be limited in expressing the performance in the music captioning task. The diversity and subjectivity of music lead to varying opinions from individuals. Only specific music genres and instruments possess distinctive characteristics that can be easily recognized. While recent studies (Gardner et al., 2023) have explored this issue, it remains a challenging problem to address. To ensure an objective evaluation, we compute CLAPscore of
2 Example Demonstrations
After fine-tuning on the AnyInstruct-108k dataset, AnyGPT demonstrates the capability and potential in any-to-any multimodal dialogue. We provide compelling conversation examples of AnyGPT in Appendix E. These examples showcase AnyGPT is capable of comprehending and reasoning contents across various modalities in any combination. Specifically, AnyGPT can comprehend instructions interwoven with multiple modalities, including text, voice, images, and music, and can adeptly select the appropriate multimodal combination for its reply. The two-stage framework of semantic-acoustic hierarchical modeling empowers AnyGPT to generate voice responses that matches the timbre and emotion of a 3-second speech prompt. For additional examples and to experience the speech and music content, we highly recommend visiting the demo page.
Conclusion
In this work, we introduced AnyGPT, an any-to-any multimodal language model that utilizes discrete representations for the unified processing of various modalities, including speech, text, images, and music. Discrete multimodal representations facilitate a seamless integration of new modalities—comparable to incorporating a foreign language—without necessitating alterations to the existing LLM architecture or training paradigms. To equip the model to handle arbitrary combinations of multimodal inputs and outputs, we synthesize the first large-scale any-to-any multimodal instruction dataset, AnyInstruct-108k, consisting of multi-turn conversations that intricately interweave various modalities. Experimental results indicate that AnyGPT achieves promising results in various cross-modal tasks and demonstrates that discrete representations can effectively and conveniently unify multiple modalities within a unified large language model.
Limitations and Future Work
The domain of any-to-any multimodal large language models (LLMs) is an emerging field of research. However, the lack of a dedicated benchmark to evaluate the models’ capabilities across multiple dimensions, as well as to mitigate potential risks, presents a considerable challenge. Consequently, the development of a comprehensive benchmark is imperative.
Although the multimodal LLMs with discrete representations can be trained stably, a higher loss is observed compared to unimodal training, preventing optimal performance in each modality. Potential strategies to improve multimodal fusion could involve scaling LLMs and tokenizers or adopting a Mixture-Of-Experts (MOE) architecture to better manage diverse data and optimize performance.
In multimodal LLMs employing discrete representations, the tokenizer’s quality sets a ceiling for the model’s comprehension and generative potential. Enhancing the tokenizer can be approached from various angles, including the adoption of superior codebook training methods, the development of more cohesive multimodal representations, and the application of information disentanglement across various modalities.".
Multimodal content, such as images and audio, often spans extensive sequences. AnyGPT, for instance, limits music modeling to 5 seconds, significantly restricting the practical usefulness of its audio output. Moreover, for any-to-any multimodal dialogue, an extended context allow for a higher number of conversational exchanges, thereby enriching the interaction’s depth and complexity.
References
Appendix A pretraining
A.2 pre-training
We employ various templates to construct multimodal sentences, ensuring a diverse spectrum within our pre-training data. Each non-text modality content is identified by special tokens placed at both the beginning and end.
Typically, the paired data comprises a non-text modality (X) - such as images, speech, or music - and its corresponding text, which could be a caption or transcription. We prompt OpenAI GPT-4 to generate hundreds of bidirectional instructions, specifically X-to-text or text-to-X such as "Please generate an image based on the provided text." Given a token sequence (S) and related text (T), we randomly pick a generation direction alongside an instruction (I) from our pre-established pool, forming a triplet (I, S, T). This triplet is then incorporated into a sequence using the template [Human]: {I}.{S}
For interleaved multimodal data, like a web document with interspersed images and text, we directly replace non-text content with the corresponding tokens sequence as they naturally form sentences.
As most of the image and music data are sourced from the web, there is a certain level of noise that can affect the quality of multimodal generation. Consequently, after the initial pre-training, we selectively utilized high-quality datasets—JourneyDB and LAION-Aesthetics for text-to-image generation, and LAION-COCO for image captioning. For music data, we incorporated the AnyInstruct-108k dataset. The remaining data were kept unchanged, and we continued to pre-train the model for an additional 4000 steps.
We report the detailed training hyperparameters of AnyGPT in Tab 8.
Appendix B Instruction Tuning
Appendix C Evaluation
We conduct a zero-shot Text-to-Speech (TTS) evaluation on the VCTK dataset. There is no overlap in speakers between our training data and the VCTK dataset. We randomly select a 3-second clip from each speaker as the vocal prompt along with a separate text as input.
The results can be found in Table 5. We evaluate the TTS systems with speaker similarity and WER. To evaluate the speaker similarity between the generated speech and the prompt speech, we employ WavLM-TDNNhttps://github.com/yangdongchao/UniAudio/blob/main/UniAudio/tools/evaluation/compute_similarity_vc.py. It can generate speaker embeddings for both the generated speech and the prompt speech, then compute the cosine similarity between these embeddings. WER is calculated using the Whisper medium model to transcribe the generated speech, with lower WER indicating higher quality of the synthesized speech.
We compare our model with VALL-E and USLM, both of which employ two autoregressive models for speech modeling. They utilize Encodec and SpeechTokenizer, respectively, as speech tokenizers.
Appendix D Prompts for Constructing Multimodal Interleaved Instruction Data
In the first stage of our pipeline to construct multimodal interleaved instruction data (Sec. 3) with GPT4. To facilitate reproducibility, we detail our prompts to the language model for brainstorming a topic pool (Fig. 5), constructing chatting scenarios (Fig. 6), and detailing the chat contents (Fig. 7), with multimodal content written as their text descriptions.