InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian, Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang, Songyang Zhang, Wenwei Zhang, Yining Li, Yang Gao, Peng Sun, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Hang Yan, Conghui He, Xingcheng Zhang, Kai Chen, Jifeng Dai, Yu Qiao, Dahua Lin, Jiaqi Wang
Introduction
Recent advancements in Large Language Models (LLMs) have sparked interest in the development of Large Vision Language Models (LVLMs) . Leading paradigms like GPT-4 , Gemini Pro 1.5 , and Claude 3 have achieved considerable success and significantly expanded the range of applications for LLMs. Open-source LVLMs are also being rapidly developed and can compete with proprietary APIs in several benchmarks. However, these open-source models still lag behind closed-source leading paradigms in versatility. They lack the ability to perform diverse vision-language comprehension and composition tasks, largely due to limited diversity in training corpus and challenges in managing long-context input and output.
To further bridge the gap between proprietary APIs and open-sourced Large Vision Language Models, we are introducing InternLM-XComposer-2.5 (IXC-2.5), a versatile LVLM supporting long-contextual input and output with diverse comprehension and composition capacities. IXC-2.5 excels in existing open-sourced LVLMs with two advantages. (1) Versatility: IXC-2.5 supports a wide range of tasks related to comprehension and composition, such as free-form text-image conversation, OCR, video understanding, article composition with illustrations, and webpage crafting. (2) Long-context capabilities in both input and output: It is natively trained with 24K interleaved image-text data, whose context window can be extended to 96K through positional encoding extrapolation , empowering the long-term human-AI interaction and content creation.
Benefiting from the long contextual capability, compared to its previous 2.0 version , IXC-2.5 has upgraded three comprehension abilities: (1) Ultra-High Resolution Understanding: IXC-2.5 enhances the dynamic resolution solution proposed in IXC2-4KHD with a native 560 × 560 ViT vision encoder, supporting high-resolution images with any aspect ratio. (2) Fine-Grained Video Understanding: IXC-2.5 treats videos as a ultra-high-resolution composite picture consisting of tens to hundreds of frames, allowing it to capture fine details through dense sampling and higher resolution for each frame. (3) Multi-Turn Multi-Image Dialogue: IXC-2.5 supports free-form multi-turn multi-image dialogue, allowing it to naturally interact with humans in multi-round conversations.
Besides comprehension, IXC-2.5 also supports two notable applications by incorporating extra LoRA parameters for text-image composition: (1) Crafting Webpages: IXC-2.5 can be readily applied to create webpages by composing source code (HTML, CSS, and JavaScript) following text-image instructions. (2) Composing High-Quality Text-Image Articles: Compared to IXC-2, IXC-2.5 leverages specially designed Chain-of-Thought (CoT) and Direct Preference Optimization (DPO) techniques to significantly enhance the quality of its written content.
We evaluated the versatility of InternLM-XComposer-2.5 (IXC-2.5) across a range of twenty-eight benchmarks, including five video benchmarks , nine structural high-resolution benchmarks , twelve general VQA benchmarks , one multi-true multi-image benchmark , and one webpage crafting benchmark . Compared to previous open-source LVLMs, IXC-2.5 achieved state-of-the-art results in 16 out of 28 benchmarks based on InternLM2-7B backend. As shown in Figure 1, the performance of IXC-2.5 matches or even surpasses proprietary APIs, e.g., GPT-4V and Gemini Pro , in 16 benchmarks.
IXC-2.5 web demo now supports audio input and output using open-source tools . You may try it at https://huggingface.co/spaces/Willow123/InternLM-XComposer.
Related Works
LVLMs for Text-Image Conversation. Large Language Models (LLMs) have received considerable attention because of their impressive performance in language comprehension and generation. Large vision-language models (LVLMs) have been developed by integrating LLMs with vision encoders to extend the ability to understand vision content, enabling the application of text-image conversation. Most existing LVLMs are trained for single-image multi-round conversations, while some works have the ability to understand multi-image inputs. However, IXC-2.5 focuses on providing a free-form long-contextual multi-turn multi-image interaction experience , which has not been addressed yet.
LVLMs for High Resolution Images Analysis. Understanding high-resolution images has significant potential applications such as OCR and document/chart analysis, which is attracting increased attention in the LVLMs area. In recent works, there are two main strategies to enable high-resolution understanding: (1) High-resolution (HR) visual encoders directly support higher resolution images. (2) Patchification: A high-resolution image is cropped into patches . Each patch is processed with a low resolution vision encoder, e.g., CLIP and visual embeddings of patches are further concatenated as inputs for LLM backends. IXC-4KHD scales the supported resolution of open-source LVLMs into 4K and beyond for the first time. IXC-2.5 combines both solutions with a vision encoder trained with a resolution of 560x560 and a dynamic resolution solution proposed in IXC2-4KHD , resulting in further improvements.
LVLMs for Video Understanding. In addition to image understanding, the LVLMs area has also witnessed emerging efforts in video analysis . To handle complex video inputs, existing works use sparse sampling or temporal pooling , compressed video tokens , memory banks , and language as a bridge for video understanding. Apart from these video-specific designs, video analysis can also be formulated to understand a high-resolution composite picture consisting of sampled video frames . Benefiting from the ability to comprehend ultra-high-resolution images and long context, IXC-2.5 exhibits strong performance on various video benchmarks for LVLMs.
Webpage Generation. Pix2Code presents an end-to-end solution for UI-to-code transformation leveraging CNNs and RNNs. This approach contends with the challenges posed by intricate visual encoding and extensive text decoding when applied to real-world UIs. In the sphere of recent advancements, works such as Sightseer , DCGen , and Design2Code have employed large vision-language models trained on synthetic screenshot-HTML paired datasets like WebSight v0.1 or v0.2 to facilitate HTML code generation. Nevertheless, the synthesized web page datasets have been critiqued for their simplicity and lack of diversity. These studies generally concentrate on the screenshot/sketch-to-code task. In contrast, our IXC-2.5 model extends these capabilities to include screenshot-to-code, instruction-aware webpage generation, and resume-to-homepage tasks. IXC-2.5 is trained using a combination of high-quality synthesized and real-world web data. Furthermore, IXC-2.5 is proficient in generating JavaScript code, thereby enabling the development of interactive front-end webpages.
Preference Alignment. Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning from AI Feedback (RLAIF) have shown great promise in aligning LLMs across various domains, including improving logical reasoning and generating helpful and harmless outputs. The typical approach involves training a reward model using human or AI preference data and fine-tuning the LLM to maximize the expected reward function with optimization algorithms like Proximal Policy Optimization (PPO) . Alternatively, Direct Preference Optimization (DPO) and the following works have emerged as leading methods that implicitly represent the reward score and eliminate the need for a separate reward model. Building on the success of RLHF and RLAIF in LLMs, recent studies have successfully extended RLHF/RLAIF algorithms for multimodal LVLMs to reduce hallucination. In this work, we investigate the application of preference alignment techniques to the text-image article composition task, with a focus on generating high-quality and stable response results.
Method
The model architecture of InternLM-XComposer-2.5 (IXC-2.5 in the following for simplicity) mainly follows the design of InternLM-XComposer2 and InternLM-XComposer2-4KHD (IXC2 and IXC2-4KHD for simplicity), including a light-weight Vision Encoder OpenAI ViT-L/14 , Large Language Model InternLM2-7B , and Partial LoRA for efficient alignment. We recommend the readers to the IXC2 and IXC2-4KHD papers for more details.
2 Multi-modal Input
Our IXC-2.5 supports diverse input modalities, including text, single/multiple images, and videos. As showin in Figure 5, a Unified Dynamic Image Partition strategy is adopted for both videos and multiple images with any resolutions and aspect ratios.
Image Processing. We mainly follow the Dynamic Image Partition and Global-Local Format design used in IXC2-4KHD with a few modifications. For the vision encoder, we reuse the ViT of resolution used in IXC2 and further increase its resolution to , so that each sub-image has 400 tokens.
For the high-resolution strategy, we unify the different strategies used in the IXC-4KHD into a scaled identity strategy. Given a maximum partition number , the image with size is resized and padded to the new image with size . This process is subject to the following constraints:
where is the scale factor, and represent the number of patches in each row and column, respectively.
For multi-image input, we assign an index to each image like IMAGE i, and format the image and text in an interleaved format.
Video Processing. We sample frames from the given video and concatenate them along the short side of the frame, leading to a high-resolution image. The frame index is also written in the image to provide the temporal relation.
Audio Processing. IXC-2.5 web demo supports audio input and output using open-source tools. For audio input, we employ Whisper to transcribe audio into text. For audio output, we utilize MeloTTS to convert the text back into audio.
3 Pre-training
During the pre-training phase, the LLM (InternLM2-7B ) is frozen while both the vision encoder and Partial LoRA are fine-tuned to align the visual tokens with the LLM. The data used for pre-training is shown in Table 1.
In practice, we employ the CLIP ViT-L-14-490 from IXC2 as the vision encoder and further increase its resolution to . For the Unified Dynamic Image Partition strategy , we set the maximum number for the pertaining. For the Partial LoRA , we set a rank of for all the linear layers in the LLM decoder block. Our training process involves a batch size of 4096 and spans across 2 epochs. The learning rate linearly increases to within the first of the training steps. Following this, it decreases to according to a cosine decay strategy. To preserve the original knowledge of the vision encoder, we apply a layer-wise learning rate (LLDR) decay strategy , and the decay factor is set to .
4 Supervised Fine-tuning
We fine-tune the model with data listed in Table 2. The maximum number of the Unified Dynamic Image Partition strategy is to handle extremely large images and videos. For video datasets, the IXC-2.5 is trained with large images concatenated by at most 64 frames. The largest training context is set to a 24,000 context window size, where the MMDU dataset can achieve this limitation. In practice, we jointly train all the components with a batch size of 2048 over 4000 steps. Data from multiple sources are sampled in a weighted manner, with the weights based on the number of data from each source. The maximum learning rate is set to , and each component has its own unique learning strategy. For the vision encoder, we set the LLDR to , which aligns with the pretraining strategy. For the LLM, we employ a fixed learning rate scale factor of . This slows down the update of the LLM, achieving a balance between preserving its original capabilities and aligning it with vision knowledge.
5 Webpage Generation
We enhance the capabilities of the IXC-2.5 to include automated webpage generation. Specifically, the IXC-2.5 is now equipped to autonomously construct web pages, utilizing HTML, CSS, and JavaScript, based on input in the form of a visual screenshot, a set of free-form instructions, or a resume document. Current open-source general-purpose large language models frequently demonstrate suboptimal performance in generating HTML and CSS relative to their proficiency in natural language generation. To address this limitation, we propose training the screenshot-to-code task using extensive datasets from WebSight v0.1/v0.2 , and Stack v2 . Subsequently, we fine-tune the model with a smaller, meticulously crafted dataset consisting of instruction-aware webpage generation and personal page generation examples.
Screenshot-to-code. In addition to the WebSight datasets, we preprocess the HTML and CSS code from the Stack v2 dataset to facilitate screenshot-to-code training. Initially, we combine the CSS and HTML code into a single file. Subsequently, we remove all comments, JavaScript code, and external links. Furthermore, we eliminate any CSS styles that are not referenced by the HTML code. We convert all files into screenshots, subsequently discarding those that did not render successfully. The remaining screenshots are then processed using the IXC2-4KHD model to assess the quality of the web pages. Following the exclusion of low-quality web pages, we retained a final set of three remaining about 250,000 high-quality web pages.
We conduct training on the LoRA model utilizing the three aforementioned datasets. The LoRA rank is set to 512. The training protocol employs a batch size of 512 and is executed over a single epoch. Initially, the learning rate is incremented linearly to within the first of the training iterations. Subsequently, the learning rate decreases to following a cosine decay schedule.
Instruction-aware Webpage Generation. A pivotal attribute of large language models lies in their capability to adhere to human instructions. To facilitate web page generation based on freeform instructions, we propose constructing data through querying closed-source large language models. Specifically, we utilize GPT-4 to generate diverse instructions and concepts for web page creation, encompassing elements such as type, style, and layout. Subsequently, these instructions are harnessed to query Claude-3-sonnet for the actual web page generation process. This approach results in 18,000 high-quality, instruction-aware samples. Additionally, we employ Tailwind CSS instead of traditional CSS, given its succinct nature.
Resume-to-homepage. In addition to instruction-aware webpage generation, we introduce a more practical task. Specifically, given a resume, the model is designed to generate a personal homepage. This homepage not only encapsulates the information present in the resume but also presents it with a well-structured and visually appealing format, improving both content organization and aesthetic layout. To generate corresponding datasets, we propose an idea-resume-homepage data generation pipeline. Initially, we leverage GPT-4 to produce resume ideas tailored for diverse personas, such as researchers, students, and engineers. GPT-4 is tasked with generating these resumes in markdown format based on the provided ideas. Upon obtaining the generated resumes, we then prompt Claude-3-sonnet to create corresponding homepages from these resumes. To enhance the interactivity of these webpages, Claude-3-sonnet is also utilized to generate JavaScript events based on the HTML code. In total, we have constructed a dataset comprising 2,000 samples.
Upon constructing the dataset for instruction-aware webpage generation and resume-to-homepage, we subsequently fine-tuned the LoRA model for 10 epochs. All other experimental settings were maintained consistent with those employed during the screenshot-to-code training phase.
6 Article Composing
Generating high-quality text-image articles (e.g., poetry, novels, short stories, and essays) is a crucial capability for AI assistants, with various applications in daily life, including education and entertainment. Building upon the IXC-2.5 SFT model in Section 3.4, we enhance creative writing capabilities for generating high-quality text-image articles. However, collecting high-quality text-image articles is a rare and expensive endeavor. Direct fine-tuning on scarce instruction data can lead to unstable responses from LVLMs in most cases. To overcome these challenges, we propose a scalable pipeline that integrates supervised fine-tuning, reward modeling, preference data collection, and DPO alignment for high-quality and stable article generation.
Supervised Fine-tuning. We begin with the SFT model (Section 3.4) and a collection of 5,000 instruction tuning data samples from IXC2 , focused on article writing. Due to the limited scale of the instruction data, we use the SFT model to rewrite the original prompts using the Chain-of-Thought (CoT) technique , generating step-by-step prompts to supplement the instruction tuning data as augmented data . We observe that the SFT model is more effective in generating long-form responses when using these augmented prompts. We then train the initial model on the augmented instruction tuning data via LoRA with the rank of and get the model to establish a starting point of our alignment pipeline.
Preference Data Collection. We use the fine-tuned model to generate diverse responses for each prompt in the augmented instruction tuning data , using different random seeds. This yields a collection of 80,000 prompt-response pairs. Next, we employ the GPT-4o model to label 2,000 responses with chosen or rejected decisions and give the reasons, which serve as our reward modeling data. We then train a reward model , sharing the same architecture of , on the reward modeling data. The reward model is used to make the chosen or rejected prediction on the remaining prompt-response pairs. These selected responses are then used to construct the pair data , while , and refer to the prompt, chosen response and rejected response, respectively. Ultimately, we obtain a total of 30,000 preference data for DPO alignment.
DPO Alignment. We use the DPO algorithm to update the SFT model on target policy from the preference data :
In practice, we use LoRA with a rank of to get the DPO model . We observe that our model tends to prioritize minimizing the likelihood of dis-preferred responses over maximizing the likelihood of preferred responses to avoid generating inappropriate or low-quality content.
In summary, our scalable pipeline consists of three primary components. First, we address the challenge of limited instruction tuning data by re-writing original prompts into augmented prompts. Next, we generate diverse responses using different random seeds, enabling the exploration of various creative possibilities. Finally, we apply the DPO algorithm to the chosen and rejected responses to refine our model’s performance. Through our pipeline, our model is capable of generating high-quality articles.
Experiments
In this section, we validate the benchmark performance of our InternLM-XComposer-2.5 (IXC-2.5) after supervised fine-tuning.
In Table 3 and Table 4, we compare our IXC-2.5 on a list of benchmarks with both closed-source APIs and SOTA open-source LVLMs (with comparable model size). Here we report video understanding results on MVBench , MLVU , MME-Video , MMBench-Video , TempCompass . For Structural High-resolution understanding, we report results on DocVQA , ChartQA , InfographicVQA , TextVQA , OCRBench , DeepForm , WikiTableQuestion (WTQ) , Visual MRC , and TabFact . For general visual question answering, we report results on MMStar , RealWorldQA, MathVista , MMMU , AI2D , MME , MMBench (MMB) , MMBench-Chinese (MMBCN) , MMBench-v1.1 (MMBv1.1) , SEED-Bench Image Part (SEEDI), MM-Vet , HallusionBench (HallB) . For Multi-True Multi-Image dialogue, we evaluate IXC-2.5 on MMDU benchmark. For webpage crafting, we report a subtask screenshot-to-code since benchmarks for others are not available in the community.
The evaluation is mainly conducted on the OpenCompass VLMEvalKit for the unified reproduction of the results.
Comparison on Video Understanding Benchmarks. As demonstrated in Table 3, IXC-2.5 exhibits competitive performance on fine-grained video understanding tasks, outperforming open-source models on 4 of the 5 benchmarks and being on par with Closed-Source APIs. For example, IXC-2.5 reaches 69.1 on the MVBench, higher than the previous SOTA method VideoChat2-7B and outperforms GPT-4V with . For the recent challenging MMBench-Video, IXC-2.5 reaches the SOTA performance on open-source models and performs close to Gemini-Pro.
Comparison on Structural High-resolution Benchmarks. Benefiting from the unified image partition strategy, IXC-2.5 could handle diverse kinds of images. Table 3 reports its performance on several structural high-resolution benchmarks. IXC-2.5 with only 7B parameters performs on par with the current large open-source LVLMs and close-source APIs. For example, IXC-2.5 gets on the DocVQA test set, the same as InternVL-1.5 which has nearly parameters. For the highly structured form and table understanding tasks, IXC-2.5 outperforms DocOwl 1.5-8B with , , on WikiTableQuestion, DeepForm and TableFace respectively.
Comparison on Multi-Image Multi-Turn Benchmarks. IXC-2.5 is capable of taking multiple images as input and conducting multi-round free-form dialogue based on them. We evaluate it quantitatively on the newly proposed MMDU benchmark . As shown in Table 4, the IXC-2.5 model demonstrates superior performance, outperforming the previous SOTA open-source model by a significant margin of 13.8%. This notable improvement highlights the effectiveness of our approach in advancing the capabilities of multi-image and multi-turn understanding.
Comparison on General Visual QA Benchmarks. IXC-2.5 is designed as a general LVLM to handle diverse multi-modal tasks. Here we report its performance on general visual QA benchmarks. As shown in Table 4, the IXC-2.5 shows superb performance on these benchmarks and on par with current large open-source LVLMs and closed-source APIs. For example, IXC-2.5 gets on the challenging MMStar and outperforms GPT-4V and Gemini-Pro. On the RealWorldQA, IXC-2.5 also performs better than Gemini-Pro and close to GPT-4V.
Comparison on Screenshot-to-code Benchmark. Table 5 presents the comparison results on the Design2Code benchmark that assesses the ability to translate visual design into code implementation. Our IXC-2.5 even surpasses the GPT-4v on average performance, which highlights the potential of IXC-2.5 to excel in bridging the gap between visual design and code implementation.
Conclusion
We have introduced InternLM-XComposer-2.5 (IXC-2.5), a cutting-edge Large Vision-Language Model (LVLM) boasting long-contextual input and output capabilities that enable advanced features such as ultra-high resolution image understanding, fine-grained video understanding, multi-turn multi-image dialogue, webpage generation, and article composing. Our comprehensive experiments demonstrate that IXC-2.5 achieves competitive performance, remarkably, with a relatively modest 7B Large Language Model (LLM) backend.
Our model sets out a promising research direction that can extend to a more contextual multi-modal environment, including long-context video understanding (e.g., long movies) and long-context interaction history, to better assist humans in real-world applications.
Acknowledgements: We deeply express our gratitude to Prof. Chao Zhang from Tsinghua University for suggestions about audio models and tools.