VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Chaoyou Fu, Haojia Lin, Xiong Wang, Yi-Fan Zhang, Yunhang Shen, Xiaoyu Liu, Haoyu Cao, Zuwei Long, Heting Gao, Ke Li, Long Ma, Xiawu Zheng, Rongrong Ji, Xing Sun, Caifeng Shan, Ran He

Introduction

Recent advancements in MLLMs have led to significant progress, particularly in the integration of visual and textual modalities. The introduction of visual information into LLMs has notably enhanced model capabilities across a range of multimodal tasks. However, with the growing appeal of human-computer interaction, the role of the speech modality has become increasingly prominent, especially in the multimodal dialogue system. In such a system, speech not only serves as a key medium for information transmission but also greatly improves the naturalness and convenience of interactions. Consequently, integrating both visual and speech modalities to achieve high-performance multimodal interactions has emerged as a critical research focus.

The integration of vision and speech in MLLMs is not straightforward due to their inherently differences . For example, visual data, such as images, convey spatial information, while speech data convey dynamic changes in time series. These fundamental differences pose challenges for simultaneous optimization of both modalities, often leading to conflicts during training. For instance, the inclusion of speech data may degrade performance on vision tasks, and vice versa. In addition, traditional speech-to-speech systems rely on separate modules for Automatic Speech Recognition (ASR) and Text-to-Speech, which can increase latency and reduce coherence, limiting their practicality in real-time applications .

In this paper, we introduce VITA-1.5, a multimodal LLM that integrates vision, language, and speech through a carefully designed three-stage training methodology. The training strategy progressively incorporates vision and speech data, relieving modality conflicts while maintaining strong multimodal performance. In the first stage, we focus on vision-language by training visual adapters and fine-tuning the model with descriptive caption and visual QA data. This step establishes the model’s foundational visual capabilities, enabling robust image and video understanding. The second stage introduces audio input processing by training an audio encoder using speech-transcription paired data, followed by fine-tuning with speech QA data. This stage equips the model with the ability to understand and respond to audio inputs effectively. Finally, in the third stage, we train an audio decoder to enable end-to-end speech output, eliminating the need for external TTS modules. This allows VITA-1.5 to generate fluent speech replies, enhancing the naturalness and interactivity of multimodal dialogue systems.

We have conducted extensive evaluations on various benchmarks related to image, video, and speech understanding, comparing the results with both open-source and proprietary models. VITA-1.5 demonstrates comparable perception and reasoning capabilities comparable to leading image/video based MLLMs, and shows significant improvements in the speech capability.

Related Work

Recently, thanks to the rapid development of language models such as GPTs , LLaMA , Alpaca , Vicuna , and Mistral , researchers have successfully extended text comprehension to multimodal understanding/reasoning through techniques like multimodal alignment and instruction tuning. For example, models such as LLaVA , Qwen-VL , Cambrian-1 , Mini-Gemini , MiniCPM-V 2.5 , DeepSeek-VL , and SliME have made significant advances in image perception and reasoning, while models like LongVA and Video-LLaVA have showcased the latest progress in video understanding. These models are increasingly capable of handling diverse data types, driving the continuous improvement of multimodal perception and understanding capabilities.

However, compared to proprietary models that support multiple modalities, including audio, image, and text (e.g., GPT-4o and Gemini-Pro 1.5 ), most open-source models have primarily focused on image and text modalities . Moreover, few open-source models have involved multimodal interaction capabilities, which is a relatively unexplored area. While works like VITA-1.0 have made initial attempts to introduce speech for human-computer interaction, introducing additional speech data poses challenges to the model’s original multimodal abilities. Furthermore, speech generation typically relies on existing TTS systems, which often results in high latency, thus impacting user experience. In this paper, we present VITA-1.5 that leverages a refined training strategies, excelling in perceiving data across four modalities (video, image, text, and audio), while also realizing near real-time vision and speech interaction.

VITA-1.5

The overall architecture of VITA-1.5 is depicted in Fig. 2. The input side is the same as that of the VITA-1.0 version , that is, adopting the configuration of “Multimodal Encoder-Adaptor-LLM”. It combines the Vision/Audio Transformer and the Multi-Layer Connector with an LLM for joint training, aiming to enhance the unified understanding of vision, language, and audio. With respect to the output side, VITA-1.5 has its own end-to-end speech module, instead of using the external TTS model like the original VITA-1.0 version.

Visual Encoder. VITA-1.5 adopts InternViT-300Mhttps://huggingface.co/OpenGVLab/InternViT-300M-448px as the visual encoder, with an input image size of 448×448 pixels, generating 256 visual tokens per image. For high-resolution images, VITA-1.5 employs a dynamic patching strategy to capture local details, improving the accuracy of image understanding.

Video Processing. Videos are treated as a special type of multiple-image input. If the video length is shorter than 4 seconds, 4 frames are uniformly sampled; for videos between 4 and 16 seconds, one frame per second is sampled; for videos longer than 16 seconds, 16 frames are uniformly sampled. No dynamic patching is applied to video frames to avoid excessive visual tokens that could hinder processing efficiency.

Vision Adapter. A two-layer MLP is used to map the visual features to visual tokens suitable for the subsequent understanding of LLM.

1.2 Audio Modality

Speech Encoder. Similar to , our audio encoding module consists of multiple downsampling convolutional layers (4x downsampling) and 24 Transformer blocks (with a hidden size of 1024). The downsampling layers help reduce the frame rate of the audio features, improving the processing speed of LLM. The audio encoder has about 350M parameters and an output frame rate of 12.5Hz. Mel-filter bank features are used as the input of the audio encoder, with a window size of 25ms and a shift of 10ms .

Speech Adapter. It consists of multiple convolutional layers with 2x downsampling.

Speech Decoder. TiCodec is used as our codec model, customizing a single codebook with a size of 1024. This single-codebook design simplifies the decoding process during the inference phase. The codec model is responsible for encoding continuous speech signals into discrete speech tokens with the frequency of 40Hz, and at the same time has the ability to decode them back into speech signals with the sample rate of 24,000Hz.

The current LLM can only output text tokens, and the speech generation capability requires the LLM to be able to output speech tokens. To this end, we add two speech decoders after the text tokens following : 1) Non-Autoregressive (NAR) Speech Decoder, which processes text tokens globally and models semantic features, with the aim of generating an initial distribution of speech tokens; 2) Autoregressive (AR) Speech Decoder generates higher quality speech tokens step by step, based on the speech information produced by the NAR decoder. The final sequence of speech tokens is then decoded into a continuous speech signal flow (waveform) using the speech decoder of the Codec model. We adopt 4 LLaMA decoder layers for both NAR and AR speech decoders, where the hidden size is 896 and the parameter size is about 120M.

2 Training Data

As shown in Table 1, the training data of multimodal instruction tuning encompass a wide range of categories, such as caption data and QA data, both Chinese and English. During different training phases, subsets of the overall dataset are selectively sampled to serve different objectives. Specifically, the datasets are categorized as follows:

Image Captioning Data. Datasets such as ShareGPT4V , ALLaVA-Caption , SharedGPT4o-Imagehttps://sharegpt4o.github.io/, and synthetic data are used to train the model to generate descriptive languages for images.

Image QA Data. Datasets like LLaVA-150Khttps://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K, LLaVA-Mixture-sample , LVIS-Instruct , ScienceQA , ChatQA , and subsets sampled from LLaVA-OV , such as general image QA and mathematical reasoning datasets, are utilized to train the model in answering image-based questions and performing visual reasoning tasks.

OCR & Diagram Data. This category supports the model in understanding OCR and diagram content, using datasets such as Anyword-3M , ICDAR2019-LSVThttp://icdar2019.org/, UReader , SynDOGnaver-clova-ix/synthdog-en, ICDAR2019-LSVT-QAhttp://icdar2019.org/, and corresponding data sampled from LLaVA-OV.

Video Data. Datasets like ShareGemini and synthetic data are used to train the model to handle video inputs and perform tasks such as captioning and video-based QA.

Pure Text Data. This category enhances the model’s capability to understand and generate languages, facilitating text-based QA tasks.

In addition to the image and video data listed in Table 1, 110,000 hours of internal speech-transcription paired ASR data, covering both Chinese and English, are incorporated to train the audio encoder and align the audio encoder with the LLM. Furthermore, 3,000 hours of text-speech paired data generated by a TTS system are used to train the speech decoder.

3 Three Stage Training Strategies

In order to ensure that VITA-1.5 performs well in tasks involving vision, language, and audio, we have to face a key challenge, i.e., training conflicts between different modalities. For example, adding the speech data could negatively impact the understanding of the vision data, as the features of speech differ significantly from those of vision, causing interference during the learning process. To address this challenge, we devise a three-stage training strategy as shown in Fig. 3. The core idea is to gradually introduce different modalities into the model, allowing it to increase the power of a new modality while maintaining the power of the existing modalities.

Stage 1.1 Vision Alignment. In this stage, our goal is to bridge the gap between vision and language. The features of the former are extracted from the pre-trained vision encoder InternViT-300M, and the latter is introduced through the LLM. We use 20% of the descriptive caption data from Table 1 for training, where only the visual adapter is trainable, while the other modules are frozen. This approach allows the LLM to initially align the visual modality.

Stage 1.2 Vision Understanding. In this stage, our goal is to teach the LLM to transcribe image content. Toward this end, we use all the descriptive caption data from Table 1. During this process, the encoder and adapter of the visual module, as well as the LLM, are trainable. The focus is to enable the model to establish a strong connection between vision and language by learning from descriptive texts about images, allowing it to understand image content via generating natural language descriptions.

Stage 1.3 Vision SFT. Following Stage 1.2, the model has acquired a basic understanding of images and videos. However, the instruction following ability is still limited, and it is difficult to cope with the visual QA task. To achieve this, we use all the QA data from Table 1 while retaining 20% of the descriptive caption data to increase the diversity of the dataset and the complexity of the tasks.

During training, the encoder and adapter of the visual module, as well as the LLM, are trainable. The key objective of this stage is to enable the model not only to understand visual content but also to answer questions following instructions.

3.2 Stage 2: Audio Input Tuning

Stage 2.1 Audio Alignment. After completing the training of Stage 1, the model has developed a strong foundation in image and video understanding. In this stage, our goal is to reduce the discrepancy between audio and language based on Stage 1, enabling the LLM to understand audio inputs. The training data consists of 11,000 hours of speech-transcription pairs. We follow a two-step approach: (a) Speech Encoder Training: We adopt a training framework used in common speech recognition systems, using a Connectionist Temporal Classification (CTC) loss function to train the speech encoder. The aim is for the encoder to predict the transcription text from the speech input. This step ensures that the audio encoder can extract speech features and map them to the text representation space. (b) Speech Adapter Training: After training the speech encoder, we integrate it with the LLM, using an audio adapter to introduce audio features into the input layer of the LLM. The training objective at this stage is to enable the LLM to output the transcription text of the speech data.

Besides, in step (b), we introduce special trainable input tokens to guide the speech understanding process. These tokens provide additional contextual information that guides the LLM used for the QA task to perform the ASR task.

Stage 2.2 Audio SFT. The focus of this stage is to introduce the QA functionality with speech questions and text answers. To achieve this, we sample 4% of the caption data and 20% of the QA data from Table 1. In terms of data processing, approximately half of the text-based questions are randomly replaced with their corresponding speech versions, generated using a TTS system.

In this stage, both the visual encoder and adapter, the audio encoder and adapter, as well as the LLM are trainable, aiming to improve the model’s adaptability with multimodal inputs. In addition, we add a classification head to the LLM’s output. This head is used to distinguish whether the input comes from speech or text. As a result, the model can more accurately interpret speech inputs and process different modalities efficiently and flexibly.

3.3 Stage 3: Audio Output Tuning

In the first two stages of training, the VITA-1.5 model has effectively developed its multimodal understanding capabilities. However, a crucial capacity, i.e., speech output, remains absent, which is essential for its role as an interactive assistant. To introduce speech output functionality without compromising the model’s fundamental abilities, we draw on the strategy , using 3,000 hours of text-speech data and employing a two-step training approach (see Fig. 3).

Stage 3.1 Codec Training. The goal of this step is to train a codec model with a single codebook using speech data. The encoder of the codec model has the ability to map speech to discrete tokens, while the decoder can map the discrete tokens back to speech stream. During the inference phase of VITA-1.5, only the decoder is used.

Stage 3.2 NAR + AR Decoder Training. The training of this stage uses text-speech paired data, where the text is fed into the tokenizer and the embedding later of the LLM to obtain its embedding vectors, and the speech is fed into the encoder of the codec model to obtain its speech tokens. The text embedding vectors are sent to the NAR speech decoder to get global semantic features, and then the features are sent to the AR speech decoder, which predicts the corresponding speech tokens. Note that the LLM is frozen during this stage, thus the multimodal performance is not affected.

Evaluation

Baselines. We compare a series of open-source MLLMs, including VILA-1.5 , LLaVA-Next , CogVLM2 , InternLM-XComposer2.5 , Cambrian-1 , MiniCPM-V-2.6 , Ovis1.5 , InternVL-Chat-1.5, InternVL-2 , LLaVA-OV , and Video-LLaVA , SliME , and LongVA , as well as 5 closed-source MLLMs, including GPT-4Vhttps://openai.com/index/gpt-4v-system-card/, GPT-4ohttps://openai.com/index/hello-gpt-4o/, GPT-4o-mini, Gemini 1.5 Pro , and Claude 3.5 Sonnethttps://www.anthropic.com/news/claude-3-5-sonnet.

Evaluation Benchmarks. To assess the image perception and understanding capabilities of VITA-1.5, we utilize several evaluation benchmarks, including MME , MMBench , MMStar , MMMU , MathVista , HallusionBench , AI2D , OCRBench , and MMVet . These benchmarks cover a wide range of aspects, including general multimodal capabilities (e.g., MME, MMBench, and MMMU), mathematical reasoning (MathVista), hallucination detection (HallusionBench), chart (AI2D) and OCR (OCRBench) understanding, providing a comprehensive evaluation results. For video understanding, we use representative evaluation benchmarks including Video-MME , MVBench , and TempCompass .

Vision-Language Capabilities. Table 2 presents a comparison of VITA-1.5’s image understanding performance. After the training of the three stages, VITA-1.5 performs comparably to the most advanced open-source models and even surpasses some closed-source models like GPT-4V and GPT-4o-mini. This result highlights the robust capabilities of VITA-1.5 in image-language tasks. As shown in Table 3, VITA-1.5 shows comparable performance to the top open-source models in the evaluation of video understanding. The notable gap compared to proprietary models suggests that VITA-1.5 still has significant room for improvement and potential for further enhancement in video understanding. Please note that after the training of Stages 2 (Audio Input Tuning) and 3 (Audio Output Tuning), VITA-1.5 retains almost its original visual-language capabilities in Stage 1 (Vision-Language Training).

2 Speech Evaluation

Baselines. The following three baseline models are used for comparison: Wav2vec2-base , Mini-Omini2 , Freeze-Omini , and VITA-1.0 .

Evaluation Benchmarks. The Mandarin Evaluation Sets consists of three datasets: aishell-1 , test net , and test meeting . These datasets are used to evaluate the model’s performance on Mandarin speech. The evaluation metric is the Character Error Rate (CER). The English Evaluation Sets include four datasets: dev-clean, dev-other, test-clean, and test-other , which are used to evaluate the model’s performance on English speech. The evaluation metric is Word Error Rate (WER).

ASR Performance. The evaluation results in Table. 4 indicate that VITA-1.5 achieves leading accuracy in both Mandarin and English ASR tasks. This demonstrates that VITA-1.5 has successfully integrated advanced speech capability to support multimodal interaction.

Conclusion

In this paper, we has presented VITA-1.5, a multimodal LLM designed to integrate vision and speech through a carefully crafted three stage training strategy. By relieving the inherent conflicts between modalities, VITA-1.5 achieves robust capabilities in both vision and speech understanding, enabling efficient speech-to-speech interactions without relying on separate ASR or TTS modules. Extensive evaluations demonstrate that VITA-1.5 performs competitively across multimodal benchmarks. We hope that VITA-1.5 can take over the banner of VITA-1.0 and continue to promote the progress of open-source models in the field of real-time multimodal interaction.

References