InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Pan Zhang, Xiaoyi Dong, Yuhang Cao, Yuhang Zang, Rui Qian, Xilin Wei, Lin Chen, Yifei Li, Junbo Niu, Shuangrui Ding, Qipeng Guo, Haodong Duan, Xin Chen, Han Lv, Zheng Nie, Min Zhang, Bin Wang, Wenwei Zhang, Xinyue Zhang, Jiaye Ge, Wei Li, Jingwen Li, Zhongying Tu, Conghui He, Xingcheng Zhang, Kai Chen, Yu Qiao, Dahua Lin, Jiaqi Wang
Introduction
The goal of developing AI systems that can understand and interact with environments over long periods, akin to human cognition, has been a central focus of research for decades. The rise of large-scale data corpora and multimodal large language models has driven significant advances in free-form multimodal question answering. Recent developments, such as Mini-Omni , VideoLLM-Online , and VITA , have made notable strides toward enabling more natural and immersive online interactions. However, challenges persist in creating systems capable of continuous interaction due to the intrinsic limitations of a single decoder-only large language model architecture.
Existing architectures encounter significant limitations in real-time and long-term streaming perception, reasoning, and memory. The sequence-to-sequence decoder-only architecture used in current MLLMs forces a switch between perception (e.g., seeing and hearing) and thinking, limiting the simultaneous processing of inputs and outputs. Additionally, existing works rely on the integration of multimodal memories within context windows. The reliance on long contexts to store historical information proves impractical for long-term use, especially in scenarios requiring continuous AI assistance. Multimodal data, like video streams, can quickly accumulate millions of tokens within a few hours, making it impractical to maintain context over multiple days of service. The cost and inefficiency of storing all historical clues within the context further limit the system’s capacity to provide continuous and long-term service. In contrast, the human brain can effortlessly integrate perception and cognition, preserving long-term multimodal memories. This is believed to be closely related to the functional partitioning design of the human brain cortex, where different areas of the cortex are responsible for distinct tasks, such as perception, memory, and cognition.
Inspired by the paradigm of Specialized Generalist AI , we propose a system InternLM-XComposer2.5-OmniLive (IXC2.5-OL) composed of fused specialized generalist models for streaming perception, reasoning, and memory, respectively. The system is designed to enable AI models to engage continuously with environments while retaining observations over time. By integrating short-term and long-term multimodal memory, our approach attempts to emulate human-like cognition, enabling more dynamic and sustained interactions.
As shown in Figure 1, the IXC2.5-OL system consists of three key modules: (1) Streaming Perception Module: This module processes the multimodal information stream on-the-fly. To ensure perception accuracy and efficiency, the video and audio streams are handled separately. A live video perception model processes the video stream, encoding the information and storing key details in memory. Meanwhile, an audio model recognizes the contents of human speech and other sounds, e.g., barking, knocking, or whistling. It triggers the reasoning process when human queries occur. (2) Multi-modal Long Memory Module: This component integrates both long-term and short-term memory, enabling the retrieval of detailed short-term information as well as long-term historical cues. It continuously compresses short-term memories into more information-rich long-term memories to enhance retrieval efficiency and accuracy. (3) Reasoning Module: The reasoning module, activated by the perception module, handles queries and performs reasoning tasks. As the component with the most model parameters, it serves as the core of the system’s deep cognitive processes.
The proposed system empowers AI with the ability to perceive, think, and memorize simultaneously. By overcoming the limitations of alternating perception and reasoning, IXC2.5-OL seeks to provide continuous, adaptive service, and long-term AI service. The proposed system will not only enhance the performance of AI assistants but will also contribute to the broader AI applications capable of continuously interacting and adapting to dynamic environments.
The IXC2.5-OL demonstrates strong performance across both audio and video benchmarks. Among the open-source models, IXC2.5-OL achieves competitive results on audio recognition (ASR) benchmarks such as Wenetspeech for Chinese and LibriSpeech for English. For video understanding benchmarks, IXC2.5-OL achieves state-of-the-art results among models with less than 10B parameters, obtaining an M-Avg of 66.2% on MLVU and an overall accuracy of 68.7% on MVBench . Additionally, it demonstrates competitive performance on Video-MME (60.6%) and MMBench-Video (1.42). On recent streaming video bench StreamingBench , IXC2.5-OL achieves new SOTA results on open-source models (73.79%), highlighting its exceptional capabilities for real-time video interactions.
To foster the development of the multimodal streaming interaction community, alongside the model parameters, the inference and deployment source code, encompassing both the web frontend and backend code, has also been released. All code and models of IXC2.5-OL are publicly available at https://github.com/InternLM/InternLM-XComposer/tree/main/InternLM-XComposer-2.5-OmniLive.
Related Works
MLLMs for Text-Image Conversation. Large Language Models (LLMs) have garnered significant attention for their remarkable capabilities in language comprehension and generation. Building on this success, Large Vision-Language Models (LVLMs) have been developed by integrating LLMs with vision encoders , extending their ability to comprehend visual content and enabling applications like text-image conversations. Earlier LVLMs were primarily designed for single-image, multi-round conversations, whereas recent advancements have expanded their capabilities to process and understand multi-image inputs.
MLLMs for Video Understanding. In addition to advancements in image understanding, the field of MLLMs has seen growing efforts in video analysis . To address the complexity of video inputs, existing approaches leverage techniques such as sparse sampling or temporal pooling , compressed video tokens , and memory banks . Additionally, some methods utilize language as a bridge for video understanding . Beyond these video-specific strategies, video analysis can also be framed as interpreting a high-resolution composite image generated from sampled video frames . Recent advancements have increasingly focused on online video understanding, aiming to simulate real-world scenarios where AI processes video streams in real-time to comprehend the environment on-the-fly. However, existing solutions still lack the capability to simultaneously perform perception, memory, and reasoning, limiting their applicability for consistent and long-term human-AI interactions.
MLLMs for Audio Understanding. Audio understanding can be effectively modeled as a sequence-to-sequence (Seq2Seq) task , which enables powerful integration with large language models by incorporating audio tokenizers and encoders . In addition to receiving the audio input, recent research investigates streaming duplex speech models that allow speakers to interrupt freely. Beyond audio-text models, emerging research delves into audio-visual models and unified architectures that process audio, visual, and text modalities .
MLLMs for Omni-Modal Understanding. Integrating multiple modalities into a single omni-modal foundation model represents a promising research direction. Existing works explore models capable of processing omni-modal inputs, typically combining video and audio, to produce outputs in various formats. These outputs include text , audio , and omni-modal contents . In the current design of IXC2.5-OL, we handle the audio and video modalities separately to mitigate potential influence during joint training. In future versions, our model will incorporate joint training across all modalities, enabling seamless omni-modality integration.
Method
As we briefly introduced in Sec.1, the IXC2.5-OL has three disentangled modules: 1) the Streaming Perception Module for on-the-fly visual and audio information processing, 2) the Multi-modal Long Memory Module for memory integration and retrieval, and 3) the Reasoning Module collect information from the perception and memory module, and handles queries and performs reasoning tasks. All the modules work simultaneously and interact asynchronously.
Besides nature language, the IXC2.5-OL could handle video and audio natively. To realize this, the Streaming Perception Module contains an Audio Translation Module and a Video Perception Module.
Audio Translation Module contains an audio encoder, an audio projector, and a Small Language Model (SLM). The audio encoder encodes the input audio sample into high-dimension features, and the audio projector further maps the feature to the input space of the SLM. The SLM outputs both the class (e.g. laughing, clapping, or raining) of the audio and the natural language within the audio (i.e. the automatic speech recognition). In practice, we use the Whisper model as the audio encoder and a Qwen2-1.8B as the SLM. The training contains two stages and we list the training data in Table 1.
Video Perception Module provides coarse-grained visual information to the Multi-modal Long Memory Module. It processes the real-time video input stream and encodes each frame into semantic features. For efficiency, we use the OpenAI CLIP-L/14 In practice.
2 Multi-modal Long Memory Module
The Multi-modal Long Memory Module is the core design to handle extremely long video input and helps the Reasoning Module to get rid of millions of tokens from its context window. It shares a similar idea from the VideoStreaming that encodes video clips into short-term memories and integrates them into long-term memory. With the given questions, it retrieved the most related video clips for the Reasoning Module. Formally, the Multi-modal Long Memory Module is trained with three tasks:
Memory Integration. Short-term memory represents the detailed information of each short video clip while the model still lacks a macro view of the video. To this end, with the short-term and global memory of a list of video clips, we integrate them into long-term memory by the Compressor in the following format:
Video Clip Retrieval. When users raise questions, the Multi-modal Long Memory Module retrieves the question-related video clips and provides both the video clips and their short-term memory to the Reasoning Module. In practice, we first encode the question to the feature space of the memory. We concatenate the long-term memory with the tokenized question as the Compressor input, and we view the last token of the output features as the memory-space-aligned question feature. Then we calculate the similarity between the question feature and each video’s global memory, and select the most related clips for the Reasoning Module.
Implementation Detail. We use Qwen2-1.8B as the LLMs and construct several kinds of training data for the three aforementioned tasks. As shown in Table. 2, we train the Video Clip Compression task with short video captioning data from multiple sources, using the same prefix captioning task designed in VideoStreaming . For the Memory Integration task and Video Clip Retrieval task, besides the off-the-shelf video grounding data, we also construct data for two unique tasks: ‘Semantics Implicit Question’ and ‘Reference Implicit Question’.
The ‘Semantics Implicit Question’ means the question does not point to some object directly, but mentions the usage or meaning of the object, and the model should find out the object by understanding the implicit question. For example, when the user asks ‘How about the weather today?’, the model should find out some weather-related object in the past video stream, such as an umbrella, a sun-glass, or something. Another example could be ‘I’m hungry, where can I heat my sandwiches?’, the model should find the microwave oven it has seen before.
The ‘Reference Implicit Question’ means the question uses pronouns rather than nouns. For example, ‘What is this’ means the models should retrieve the current frames, although it does not mention any exact objects.
Both kinds of implicit questions are commonly used in real-world communication while current models failed to handle them, so we construct corresponding training data to empower the model with these capabilities.
3 Reasoning Module
The Reasoning Module is initialized by an improved version of InternLM-XComposer2.5 (IXC2.5 in the following for simplified statement) and we add a memory projector to align the memory feature with IXC-2.5. For a given questions and both visual and memory information provided by the Memory Module, we formulate the input as:
In real-world usage, there exists some noisy input that should not be answered (e.g., the user says ‘enn…’ or ‘ok…’), the model should keep salient and wait for the next question. To realize this, we add an additional ‘Instruction Prediction’ process for each question to decide it should be answered or not.
4 System Pipeline
As illustrated in Figure 3, the system comprises the Frontend, SRS Server, and Backend Server.
Frontend. The frontend application, developed with JavaScript, enables the camera and microphone to capture video and audio stream inputs, which are then pushed to the SRS server. Concurrently, it establishes a WebSocket connection with the backend to listen for audio outputs and interrupt signals. When audio data is received, the frontend plays it. Upon receiving an interrupt signal, the frontend suspends the audio playback and discards the pending audio.
SRS Server. SRS (Simple Realtime Server) is a straightforward and efficient real-time video server, adept at supporting a multitude of real-time streaming protocols such as RTMP, WebRTC, HLS, HTTP-FLV, SRT, and others. It is renowned for its ability to reliably receive and deliver audio and video streams.
Backend Server. After establishing a WebSocket connection with the frontend, the backend will pull streaming from the SRS Server and initiate separate threads to read audio and video.
The audio reading thread will segment the audio stream into 4096-bit chunks and enqueue them into the Audio Queue. The Voice Activity Detection (VAD) thread continuously reads data from Audio Queue and detects the start and end of voice activity. Upon detecting the start of voice activity, the backend sends an interrupt signal to the frontend to pause the currently playing audio, and at the same time, dispatches a backup signal to the video process, directing it to save the current memory state. When detecting the end of voice activity, the entire voice segment will be enqueued into ASR Todo Queue. The ASR thread continuously reads audio segments from ASR Todo Queue, performs background noise classification and voice recognition on them, and then enqueues the results into LLM Todo Queue for use by the LLM.
The video reading thread reads video frames at a rate of 1 frame per second and enqueues them into Frame Queue. The compressor process reads video frames from the queue, recognizes them, extracts relevant memory, and stores it. Upon receiving a backup signal from the VAD thread, the compressor process will save the current memory state for later retrieval.
The LLM process reads text from the LLM Todo Queue and determines whether it is an instruction that requires a response from the model. For texts identified as instructions, the compressor process will use the current instruction and the backed-up memory to perform memory grounding, in order to retrieve memories related to the instruction. The LLM process will then generate a response based on the retrieved memories and the instruction, and enqueue the resulting output into TTS Todo Queue. An additional TTS thread (e.g., F5-TTS , MeloTTS ) will convert the text from the TTS Todo Queue into audio and send it to the frontend.
Experiments
In this section, we validate the benchmark performance of our InternLM-XComposer2.5-OmniLive (IXC2.5-OL), including both audio and video benchmarks.
We evaluate our audio models on two prominent automatic speech recognition (ASR) benchmarks: Wenetspeech for Chinese (CN) and LibriSpeech for English (EN). WenetSpeech includes two test sets: Test_Net, which represents high-quality and relatively clean Chinese speech, and Test_Meeting, which captures more challenging conversational scenarios. LibriSpeech consists of four splits: Dev_clean and Test_clean, which contain clean, high-quality English speech, and Dev_other and Test_other, which include noisier, more complex utterances.
As shown in Table 3, our IXC2.5-OL demonstrates superior performance compared to recent streaming audio LLMs such as VITA and Mini-Omni, particularly achieving lower Word Error Rates (WER) across both CN and EN benchmarks with merely a lightweight 1.5B LLM.
2 Video Benchmarks
In Tables 4, 5, 7 and 8, we compare IXC2.5-OL with both closed-source APIs and open-source models on conventional video understanding benchmarks, including MLVU , Video-MME , MMBench-Video and MVBench . Furthermore, we also assess the performance of different models on the recently proposed StreamingBench , which is designed to better evaluate performance for real-time video interactions. The results of this comparison are presented in Table 6. For the video benchmarks, the base model utilizes 64 sampled frames for each video during evaluation.
MLVU is a comprehensive benchmark designed for evaluating Multimodal Large Language Models in Long Video Understanding tasks. The videos range from 3 minutes to 2 hours and include nine distinct evaluation tasks. Here, we evaluate seven multi-choice tasks, including Topic Reasoning, Anomaly Recognition, Needle QA, Ego Reasoning, Plot QA, Action Order, and Action Count. The detailed comparisons are given in Table 4. The IXC2.5-OL exhibits state-of-the-art (SOTA) performance among closed-source APIs, and open-source models with parameters less than 10 billion, surpassing the previous SOTA by for Video-XL, for GPT-4o.
Video-MME
Video-MME is a high-quality video benchmark. The videos are collected from 6 primary visual domains with 30 subfields to ensure broad scenario generalizability, encompassing both short-, medium-, and long-term videos, ranging from 11 seconds to 1 hour. As demonstrated in Table 5, the IXC2.5-OL exhibits competitive performance on this benchmark, comparable to previous SOTA MiniCPM-V 2.6.
StreamingBench
StreamingBench is a streaming video benchmark designed for real-time video evaluation. It comprises 18 tasks, showcasing 900 videos and 4,500 human-curated QA pairs. In this context, we focus on assessing visual understanding in real-time. Table 6 illustrates the comparative analysis, demonstrating that IXC2.5-OL excels among all open-source models, achieving a improvement over the previous state-of-the-art model, LLaVA-OneVision, and falling just short of the closed-source API, Gemini 1.5 Pro. This performance solidifies IXC2.5-OL’s remarkable prowess in real-time video interaction.
MMBench-Video
MMBench-Video is a free-form QA video benchmark consisting of 600 videos and 2000 QA pairs. The duration of each video varies from 30 seconds to 6 minutes. Given the open-ended nature of the answers, the benchmark utilizes GPT-4-based evaluation to enhance quality in terms of accuracy, consistency, and alignment with human judgment. The results are presented in Table 7. IXC2.5-OL demonstrates state-of-the-art performance on perception tasks and comparable performance on overall evaluations.
MVBench
MVBench is a video benchmark that emphasizes temporal understanding. It encompasses 20 challenging video tasks that cannot be effectively addressed using a single frame. As shown in Table 8, IXC2.5-OL, despite having a smaller 7B parameter size, has outperformed both the GPT-4 series and the 72B open-source model LLaVA-OneVision, demonstrating its strong capability in understanding video temporal dynamics.
Conclusion
We have presented IXC2.5-OL, a real-time streaming model that advances multi-modal text, audio, and visual capabilities with long-term memory. IXC2.5-OL empowers users to engage in dynamic and interactive experiences. Our model’s real-time processing enables fluid and responsive interactions, allowing users to engage with ever-changing environments of multimodal data seamlessly, providing a more intuitive and efficient user experience. Our future work will focus on reducing system latency to provide a seamless user experience.