Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, Lixin Gu, Xuehui Wang, Qingyun Li, Yiming Ren, Zixuan Chen, Jiapeng Luo, Jiahao Wang, Tan Jiang, Bo Wang, Conghui He, Botian Shi, Xingcheng Zhang, Han Lv, Yi Wang, Wenqi Shao, Pei Chu, Zhongying Tu, Tong He, Zhiyong Wu, Huipeng Deng, Jiaye Ge, Kai Chen, Kaipeng Zhang, Limin Wang, Min Dou, Lewei Lu, Xizhou Zhu, Tong Lu, Dahua Lin, Yu Qiao, Jifeng Dai, Wenhai Wang
Introduction
In recent years, multimodal large language models (MLLMs) have emerged as a pivotal technology in artificial intelligence, capable of processing and understanding information from multiple modalities such as text, images, and videos. These models promise breakthroughs across fields like natural language processing, computer vision, and human-computer interaction. However, developing large-scale MLLMs remains a challenging task, requiring significant computational resources, sophisticated architectures, and the ability to effectively integrate diverse data types in a scalable manner.
Various attempts have been made to address these challenges, including enhancing model architectures , scaling vision encoders and language models , incorporating more diverse and high-quality datasets , and refining the test-time scaling process to boost performance. Notable commercial models, like GPT-4o and Claude-3.5-Sonnet , have demonstrated exceptional performance, their closed nature limits transparency and accessibility, leaving gaps in the open-source community. While open-source multimodal models such as the InternVL series and Qwen-VL series have provided high-performance, transparent alternatives, they still fall short in terms of achieving the desired levels of performance and efficiency.
In this work, we introduce InternVL 2.5, an advanced large-scale MLLM series that builds upon the foundational architecture of InternVL 2.0. Continuing the objectives of the entire InternVL series, we aim to bridge the performance gap between commercial closed-source models and open-source multimodal models. In InternVL 2.5, we systematically explore various factors in MLLM, including how changes in vision encoders, language models, dataset sizes, and inference times affect the overall performance of the model, demonstrating the relationship between scaling and performance in multimodal models. Specifically, we have some interesting findings: (1) Large vision encoders significantly reduce the dependency on training data when scaling up MLLMs. As shown in Table 3, compared to Qwen2-VL-72B equipped with a 600M vision encoder, our InternVL2.5-78B with a 6B vision encoder can achieve better performance using only 1/10 of the training tokens. This greatly reduces the exploration cost when scaling up MLLMs; (2) Data quality matters. Upgrading InternVL from 2.0 to 2.5 doubled the dataset size, but strict filtering greatly improved quality. For example, we carefully excluded the anomalous samples (e.g., repetitive patterns), achieving substantial improvements in Chain-of-Thought (CoT) reasoning tasks such as MMMU and complex challenges like the OlympiadBench . Note that, most existing open-source MLLMs tend to underperform when using CoT . (3) Test-time scaling is beneficial for difficult multimodal QA. For challenging tasks such as MMMU, the InternVL2.5-78B with CoT reaches 70.1%, which is 3.7 points higher than the direct response. Subsequently, we have successfully verified that CoT can be further combined with majority voting and bring additional improvements.
Our contributions can be summarized as follows:
(1) We release InternVL 2.5 to the open-source community, providing a powerful tool for the development and application of multimodal AI systems and encouraging further research in this domain.
(2) We investigate how scaling different components of the MLLMs such as vision encoders, language models, dataset sizes, and inference time affect performance.
(3) Through extensive evaluations on diverse benchmarks—including multi-discipline reasoning, document understanding, multi-image / video understanding, real-world comprehension, multimodal hallucination detection, visual grounding, multilingual capabilities, and pure language processing—InternVL 2.5 exhibits competitive performance, rivaling leading commercial models like GPT-4o and Claude-3.5-Sonnet . It is the first open-source MLLM to surpass 70% on the MMMU validation set , setting a new benchmark and highlighting the potential of open-source solutions in advancing multimodal AI.
Model Architecture
As shown in Figure 2 and Table 2, InternVL 2.5 retains the same model architecture as its predecessors, InternVL 1.5 and InternVL 2.0, following the “ViT-MLP-LLM” paradigm widely adopted in various MLLM studies .
In this new version, our implementation of this architecture integrates a newly incrementally pre-trained InternViT-6B or InternViT-300M with various pre-trained LLMs of different sizes and types, including InternLM 2.5 and Qwen 2.5 , using a randomly initialized 2-layer MLP projector. As in the previous version, to enhance scalability for high-resolution processing, we simply applied a pixel unshuffle operation, reducing the number of visual tokens to one-quarter of the original. Consequently, in our model, a 448448 image tile is represented by 256 visual tokens.
In terms of input data preprocessing, we adopted a similar dynamic resolution strategy as InternVL 1.5, dividing images into tiles of 448448 pixels based on the aspect ratio and resolution of the input images. The key difference, starting from InternVL 2.0, is that we additionally introduced support for multi-image and video data, as shown in Figure 2(b). Different data types correspond to different preprocessing configurations, which we will detail in Section 3.1.
2 Vision Encoder
InternVL employs InternViT as the vision encoder. To better document the training progression of InternViT, we have provided detailed information in Table 1. InternViT currently has two different model sizes, including InternViT-6B and InternViT-300M.
InternViT-6B. InternViT-6B-224px was first introduced in our CVPR paper , and its structure follows the vanilla ViT , with minor adjustments incorporating QK-Norm and RMSNorm . It had 5.9B parameters, 48 layers, a hidden size of 3200, and 25 heads, and it was trained using a contrastive loss . Due to the limited gains at that time, we adopted an incremental pre-training strategy to continuously refine its weights. Specifically, we connected InternViT-6B to an LLM via an MLP projector and, following a brief MLP warmup, jointly trained the InternViT-6B using a next token prediction loss (as shown in Figure 4(a)) to enhance its visual feature extraction capabilities. In the V1.0 and V1.2 versions, we used a fixed resolution of 448448 for training, but in later versions, we switched to dynamic resolution training to improve high-resolution processing. As detailed in the InternVL 1.5 report , we removed the last three layers of InternViT-6B-448px-V1.2, reducing its depth from 48 to 45 layers, as these layers were more tuned to the CLIP loss objective, prioritizing global alignment over local information. As a result, all subsequent versions, including the latest InternViT-6B-448px-V2.5, have 45 layers and 5.5B parameters.
InternViT-300M. InternViT-300M-448px-Distill is a distilled variant of the teacher model, InternViT-6B-448px-V1.5, utilizing a cosine distillation loss. This model comprises 0.3B parameters, 24 layers, a hidden size of 1024, and 16 attention heads. Unlike the 6B version, the 0.3B variant employs standard LayerNorm without QK-Norm . To reduce distillation costs, we initialized this model using CLIP-ViT-Large-336px where applicable, despite some architectural differences. After distillation, we integrated this model with an LLM and, following a similar procedure as described above, trained the vision encoder with dynamic high-resolution and the NTP loss. Then, we extracted the vision encoder and released it as InternViT-300M-448px. In this report, we further refined the InternViT-300M by incrementally pre-training the previous weights on a more diverse data mixture using the NTP loss, leading to the enhanced InternViT-300M-448px-V2.5.
3 Large Language Model
In Table 2, we provide an overview of the language models used across different versions of InternVL, including InternVL 1.5, InternVL 2.0, and the latest InternVL 2.5. As shown, earlier versions primarily built on language models such as InternLM 2 , Qwen 2 , Phi 3 , Yi , and Llama 3 . To achieve better performance, in the InternVL 2.5 series, we have comprehensively upgraded the language backbones to the latest state-of-the-art models, including InternLM 2.5 and Qwen 2.5 .
Training Strategy
In InternVL 2.0 and 2.5, we extend the dynamic high-resolution training approach introduced in InternVL 1.5 , enhancing its capabilities to handle multi-image and video datasets. The process mainly consists of the following steps:
Closest Aspect Ratio Matching. Given an input image with dimensions , the aspect ratio is computed as . The objective is to resize the image into tiles of size (where ) while selecting the closest aspect ratio that minimizes distortion. The number of tiles, , is constrained within a predefined range .
To find the optimal aspect ratio for resizing, we define the set of target aspect ratios as:
The closest aspect ratio is selected by minimizing the difference between the original aspect ratio and each target aspect ratio :
In cases where multiple aspect ratios produce the same difference (e.g., 1:2 and 2:4), we prioritize the aspect ratio that results in an area less than or equal to twice the original image size. This helps to some extent in preventing the excessive enlargement of low-resolution images.
Image Resizing and Splitting. Once the best aspect ratio is determined, the image is resized to new dimensions , where and are the factors corresponding to :
The image is then split into tiles of size , with the number of tiles calculated as . Each tile is cropped from the resized image to ensure consistent size.
Thumbnail Generation. Optionally, if the number of tiles , the original image is resized to a square of dimensions to generate an additional thumbnail . This thumbnail is appended to the list of tiles, providing a global view alongside the localized tiles. In cases where , there is no thumbnail to append, and the mechanism naturally skips this step.
Data Formats for Different Data Types. As shown in Figure 3, the dynamic high-resolution method in InternVL 2.0 and 2.5 extends beyond single-image datasets to also support multi-image and video datasets.
For single-image datasets, the maximum number of tiles is allocated to a single image, ensuring that it is processed at the highest possible resolution. In this scenario, visual tokens are enclosed within and tags, with no additional auxiliary tags used.
In the case of multi-image datasets, the total number of tiles is distributed across all images within one sample. Each image is identified by an auxiliary tag like Image-1 to clearly label individual images. The images themselves are enclosed within and tags, denoting the start and end of the image data. The number of tiles assigned to each image is proportional to the total number of images , following the equation:
For video data, this approach is simplified by setting . Each video frame is resized to a fixed resolution of 448448, eliminating the need for tiling. This is because, during training, a large number of frames (e.g., 32 or 64) are typically extracted from a single video. For our model, even without high-resolution input, this still results in 8,192 or 16,384 visual tokens. Each video frame, labeled with tags like Frame-1, is enclosed within the and tags, similar to image data.
2 Single Model Training Pipeline
The training pipeline for a single model in InternVL 2.5 is structured across three stages, designed to enhance the model’s visual perception and multimodal capabilities. Each stage progressively integrates vision and language modalities, balancing performance optimization with training efficiency.
Stage 1: MLP Warmup. As shown in Figure 4(a), the training begins with warming up the MLP projector, which is the initial bridge between visual and language representations. In this stage, only the MLP projector is trained while both the vision encoder (i.e., InternViT ) and language model are frozen. To achieve optimal performance, we begin using the dynamic high-resolution training strategy from this stage, even though it increases the training cost.
In this phase, we utilize the pre-training data mixture as outlined in Table 4. The data is formatted in a structured ChatML style and optimized with the NTP loss. Additionally, a higher learning rate is applied to accelerate convergence, allowing the MLP to quickly adapt to the LLM’s input space and establish robust cross-modal alignment. The MLP warmup phase ensures the model is well-prepared to handle multimodal tasks before unlocking additional trainable components in later stages, thereby improving training stability.
Stage 1.5: ViT Incremental Learning (Optional). As shown in Figure 4(a), Stage 1.5 introduces incremental learning for the vision encoder. During this stage, both the vision encoder and MLP projector are trainable, and training is conducted using the same pre-training data mixture and NTP loss as in Stage 1. The aim of this stage is to enhance the vision encoder’s ability to extract visual features, allowing it to capture more comprehensive information, especially for domains that are relatively rare in web-scale datasets (e.g., LAION-5B ), such as multilingual OCR data and mathematical charts, among others.
As shown in Table 3, a lower learning rate is used in this stage to prevent catastrophic forgetting, ensuring the encoder doesn’t lose previously learned capabilities. Additionally, the vision encoder only needs to be trained once unless new domain requirements or data are introduced. Once trained, it can be reused with different LLMs without retraining (see Figure 4(b) and Section 3.3), making Stage 1.5 optional. This is particularly beneficial when the encoder has already been optimized for some specific tasks, allowing it to integrate with LLMs of various sizes without significant additional costs.
Stage 2: Full Model Instruction Tuning. In the final stage, as illustrated in Figure 4(a), the entire model—comprising the ViT, MLP, and LLM—is trained on high-quality multimodal instruction datasets. Data quality is especially important here, as the LLM, responsible for generating the final user-facing output, is now trainable. Even a small amount of noisy data (e.g., a few thousand samples) can lead to abnormal model behavior, like repetitive output or specific erroneous results. To mitigate the degradation of the LLM, we implement strict data quality controls during this stage.
Additionally, the training hyperparameters in this stage are kept simple, with a unified learning rate applied to the entire model rather than different learning rates for various components. After completing this stage, InternVL 2.5’s full training process is finished. Although further improvements could be made through Stage 3—post-training with higher-quality data or other training methods (e.g., preference optimization)—we plan to leave this for the future.
3 Progressive Scaling Strategy
As shown in Figure 4, we propose a progressive scaling strategy to efficiently align the vision encoder (e.g., InternViT) with LLMs. While we previously adopted a similar strategy in the training of InternVL 1.5 and 2.0, this is the first time the approach has been formalized into a clear methodology. This strategy adopts a staged training approach, starting with smaller, resource-efficient LLMs and progressively scaling up to larger LLMs. This approach stems from our observation that even when the ViT and LLM are jointly trained using NTP loss, the resulting visual features are generalizable representations that can be easily understood by other LLMs.
Specifically, in Stage 1.5, the InternViT is trained alongside a smaller LLM (e.g., 20B), focusing on optimizing fundamental visual capabilities and cross-modal alignment. This phase avoids the high computational costs associated with training directly with a large LLM. Using a shared-weight mechanism, the trained InternViT can be easily transferred to a larger LLM (e.g., 72B) without requiring retraining. Consequently, when training a larger model, Stage 1.5 can be skipped (see Table 3), as the optimized InternViT module from earlier stages is reused. This not only accelerates training but also ensures that the vision encoder’s learned representations are preserved and effectively integrated into the larger model.
By employing this progressive scaling strategy, we achieve scalable model updates at a fraction of the cost typically associated with large-scale MLLM training. For example, Qwen2-VL processes a cumulative total of 1.4 trillion tokens, whereas our InternVL2.5-78B is trained on only about 120 billion tokens—less than one-tenth of Qwen2-VL. This approach proves particularly advantageous in resource-constrained settings by maximizing the reuse of pre-trained components, minimizing redundant computations, and enabling the efficient training of models capable of addressing complex vision-language tasks.
4 Training Enhancements
To enhance the model’s adaptability to real-world scenarios and overall performance, two key techniques are introduced. These optimizations are essential in improving the user experience and the model’s benchmark performance.
Random JPEG Compression. To avoid overfitting during training and enhance the model’s real-world performance, we apply a data augmentation technique that preserves spatial information: JPEG compression. Specifically, random JPEG compression with quality levels between 75 and 100 is applied to simulate the degradation commonly found in internet-sourced images. This augmentation improves the model’s robustness to noisy, compressed images and enhances the user experience by ensuring more consistent performance across varied image qualities.
Loss Reweighting. Token averaging and sample averaging are two widely applied strategies for weighting the NTP loss. Token averaging computes the average NTP loss across all tokens, whereas sample averaging first averages the NTP loss within each sample (across tokens) and then averages across the number of samples. These strategies can be expressed in a unified format:
where and denote the loss and weight of token , respectively, and denotes the number of tokens in the response to which token belongs.
When using token averaging, each token contributes equally to the final loss, which can result in gradients biased toward responses with more tokens, leading to a drop in benchmark performance. In contrast, sample averaging ensures that each sample contributes equally, but it can cause the model to favor shorter responses, negatively impacting the user experience. To mitigate bias toward either longer or shorter responses during training, we apply a reweighting strategy where . This approach, named square averaging, balances the contribution of responses with different lengths.
Data Organization
In InternVL 2.0 and 2.5, the organization of the training data is controlled by several key parameters to optimize the balance and distribution of datasets during training, as shown in Figure 5.
Data Augmentation. Firstly, data augmentation (i.e., JPEG compression introduced in Section 3.4) is applied conditionally, allowing for enhanced robustness by enabling or disabling augmentation techniques based on dataset characteristics. Specifically, we enable this augmentation for all image datasets, while disabling it for all video datasets, to ensure that different video frames have the same image quality.
Maximum Tile Number. The parameter defines the maximum number of tiles allowed per dataset, effectively controlling the resolution of the image or video frame fed into the model. This ensures flexibility in handling datasets of varying complexity and type. For example, we can set or for multi-image datasets, high-resolution documents, or infographics, use or for most other low-resolution image datasets, and set for video datasets. This adjustment was first introduced in InternVL 2.0, whereas in InternVL 1.5, a uniform value of was applied across all datasets.
Repeat Factor. Finally, the repeat factor determines the sampling frequency of each dataset. With , this parameter enables down-sampling when , reducing the dataset’s weight during training, or up-sampling when , effectively increasing the number of epochs for that dataset. This mechanism finely adjusts the relative proportions of datasets, ensuring a balanced distribution across training data. By adjusting , especially in multi-task learning, the data from each domain or task receives appropriate training, preventing overfitting or underfitting of any single dataset, leading to more balanced model performance.
2 Multimodal Data Packing
In InternVL 2.0 and 2.5, we implement a data-packing strategy to enhance GPU utilization and improve training efficiency. This approach reduces padding by concatenating multiple samples into longer sequences, thereby maximizing the utilization of the model’s input sequence capacity. Specifically, for multimodal models like InternVL, data packing should account for two dimensions: (a) Sequence length for the LLM, which corresponds to the standard input sequence length used in language models. This remains essential in multimodal tasks; (b) Image tile number for the ViT, which denotes the number of image tiles processed by the vision encoder. Efficient management of this dimension is crucial for optimizing training efficiency.
To handle these dimensions efficiently, our data-packing strategy comprises the following steps:
(1) Select: During the selection phase, the algorithm operates similarly to a standard dataset without data-packing, directly sampling independent data. Each sampled item is truncated into multiple smaller items and treated as separate samples. This ensures that the sequence length and image tile count of each sample are within the predefined thresholds (context length) and (image tile limit), respectively.
(2) Search: For a given independent sample, the algorithm searches for another sample from the buffer list to pack them together. The resulting packed sample must have a sequence length shorter than and include fewer than image tiles. If multiple buffers satisfy these requirements, the one with the longest sequence length and the maximum number of image tiles is selected. In practice, the buffer list is maintained in descending order and a binary search is performed to accelerate the search process.
(3) Pack: The sampled data and the selected buffer are packed into a single sequence. If no buffer is selected in the previous step, the sample remains unchanged and proceeds directly to the next phase. Notably, tokens in the packed data can only attend to the context within their respective samples and cannot attend to tokens from other packed samples. Furthermore, the positional index of each sample is maintained independently.
(4) Maintain: In the maintenance phase, if a packed sample exceeds or contains more than image tiles, it is immediately yielded for training. Otherwise, the packed sample is inserted into the buffer list. If the buffer list exceeds its capacity, the sample with the longest sequence length and the highest number of image tiles is yielded to maintain buffer efficiency.
3 Data Filtering Pipeline
During model development, we observed that LLMs are significantly more sensitive to data noise than vision encoders. As shown in Figure 4, during Stage 2, when all model weights are fully trainable, even a small fraction of anomalous samples—such as outliers or repetitive data, numbering only a few thousand—can lead to aberrant model behavior during inference. While conventional wisdom assumes that minor noise in large-scale datasets can be ignored, our findings indicate otherwise: even a tiny fraction of noisy samples can degrade MLLM performance and user experience.
Among these anomalies, we identify repetitive generation as one of the most detrimental issues. In many open-source or synthetic datasets, a small number of samples with repetitive patterns—comprising merely thousands of examples in our fine-tuning data mixture—can cause the model to spiral into repetitive loops, particularly in long-form outputs or CoT reasoning tasks. This phenomenon undermines the effectiveness of test-time scaling strategies. To address this challenge and support future research, we designed an efficient data filtering pipeline to remove low-quality samples, thereby minimizing the risk of repetitive generation.
As shown in Figure 8, our data filtering pipeline consists of two modules. For pure-text data, we implemented three key strategies: (1) LLM-Based Quality Scoring: We begin by categorizing datasets into distinct domains (e.g., disciplinary, programming, mathematics, general). Next, we assign a quality score, ranging from 0 to 10, to each sample using a pre-trained LLM with a domain-specific prompt. Samples with scores below a specified threshold (e.g., 7) are then removed to ensure data quality. (2) Repetition Detection: We use an LLM combined with a specialized prompt to identify repetitive patterns. These samples are then subjected to manual review, and those scoring below a threshold (e.g., 3) are removed to maintain data quality. (3) Heuristic Rule-Based Filtering: We apply specific rules, such as filtering out sentences with abnormal lengths, excessively long sequences of zeros, text with an excessive number of duplicate lines, etc, to identify anomalies in the data. Although this approach may occasionally produce false positives, it improves the detection of anomalous samples. All flagged samples are manually reviewed before final removal.
For multimodal data, given the limitations of open-source MLLMs in scoring such data, we focused on mitigating repetitive patterns through two strategies: (1) Repetition Detection: We exempted high-quality academic datasets and used a specific prompt to identify repetitive patterns in the remaining data. These samples were removed following the same manual review process we applied to textual data. (2) Heuristic Rule-Based Filtering: Similar heuristic rules are applied, followed by manual verification to ensure dataset integrity.
This rigorous data-filtering pipeline significantly reduced the occurrence of anomalous behaviors, particularly repetitive generation, with notable improvements in CoT reasoning tasks. However, we recognize that data filtering alone cannot completely eliminate such issues. This may be due to the inherent noise introduced during the LLM’s pre-training process, which our multimodal post-training efforts can only partially mitigate without fundamentally resolving the issue of repetitive outputs. Future work will explore preference optimization and other strategies to further suppress anomalies and enhance both model performance and user experience.
4 Pre-training Data Mixture
To comprehensively enhance the model’s performance and strengthen its ability to handle complex tasks in real-world scenarios, we collect a broader range of domain-specific data compared to the training corpus of InternVL 1.5 and 2.0. As shown in Table 4, our training corpus is sourced from captioning, general QA, mathematics, charts, OCR, knowledge, grounding, documents, conversation, medical, and GUI tasks.
Notably, during the development of our models, we utilized conversation-format instruction data. For non-conversational datasets, such as image captioning, OCR, and object detection datasets, we construct questions to transform the data into a conversational format. At this stage, since only the parameters of MLP (i.e., Stage 1) or MLP and ViT (i.e., Stage 1.5) are trainable, both low-quality and high-quality data are incorporated. The goal is to enrich the model’s world knowledge as much as possible by exposing it to diverse domain data, thereby improving its generalization capabilities.
In our view, the ideal scenario is for the fine-tuning data mixture to be a subset of the pre-training data mixture. This ensures that the data in this subset can be adequately trained within the vision encoder. However, in practice, due to the high training costs of Stage 1.5, achieving this is often difficult. Therefore, in the training of InternVL 2.5, only a subset of the datasets from the fine-tuning data mixture was included in the pre-training data mixture.
5 Fine-tuning Data Mixture
As shown in Figure 7, from InternVL 1.5 to 2.0 and then to 2.5, the dataset has undergone iterative improvements in scale, quality, and diversity. In terms of data scale, the number of samples grows from 5.1M in InternVL 1.5 to 7.3M in InternVL 2.0, and further doubles to 16.3M in InternVL 2.5. For diversity, our training data spans multiple domains, including general QA, charts, documents, OCR, science, medical, GUI, code, mathematics, et al., while covering multiple modalities such as single-image, multi-image, video, and text.
In InternVL 2.5, single-image data constituted the majority with 45.92% of tokens, while multi-image data accounted for 9.37%, video data contributed 39.79%, and pure-text data made up 4.92%. Compared to earlier versions, multi-image and video data achieved the most notable increases, leading to the enhanced multi-image and long video comprehension abilities of InternVL 2.5. Quality improvements were achieved through unifying conversation templates, using language models to score and refine data, removing repetitive patterns, applying heuristic rules to filter low-quality samples, and rewriting short responses into high-quality and longer interactions. This ensured a robust dataset for model training.
Evaluation on Multimodal Capability
To comprehensively evaluate InternVL’s performance on multimodal tasks, we employ a diverse set of benchmarks, including both well-established classic datasets and newly introduced ones provided by VLMEvalKit . These benchmarks span a wide range of categories, aiming to provide a thorough and balanced assessment of InternVL’s capabilities across various multimodal tasks.
We evaluate InternVL’s multimodal mathematical and reasoning capabilities through a comprehensive assessment across various discipline-related benchmarks.
MMMU : MMMU is a benchmark evaluating MLLMs on college-level tasks across six disciplines, testing expert-level reasoning and advanced perception in specific fields. We report the maximum accuracy achieved across both direct-answer and CoT reasoning approaches on the MMMU validation and test sets.
MMMU-Pro : MMMU-Pro is an upgraded version of the MMMU benchmark, designed to more accurately and rigorously evaluate the multimodal understanding and reasoning capabilities of models in a wide range of academic disciplines. We report three metrics: standard (10 options), vision, and overall (the average of standard and vision). Here, “standard” and “vision” are the maximum scores from the CoT and direct-answer settings, consistent with the original paper.
MathVista : MathVista is a benchmark for evaluating MLLMs’ mathematical reasoning in visual contexts, encompassing reasoning types such as algebra, geometry, and statistics. We report the scores on the testmini set.
MATH-Vision : MATH-Vision is a high-quality dataset of 3,040 visually contextualized math problems sourced from real competitions. We report performance on both the testmini and full sets.
MathVerse : MathVerse is a visual math benchmark for evaluating MLLMs in solving diagram-based math problems. It comprises 2,612 high-quality, multi-subject math problems, each transformed into six distinct versions with varying degrees of visual and textual information. We report performance on the testmini set.
OlympiadBench : OlympiadBench is a bilingual, multimodal benchmark with high-difficulty math and physics problems from Olympiad competitions and Gaokao. Each problem is annotated with expert-level step-by-step reasoning, enabling detailed assessment of logical deduction and problem-solving abilities. This benchmark is challenging, and a well-defined CoT prompt can significantly improve performance.
1.2 Evaluation Results
Multidisciplinary reasoning ability reflects a model’s capacity to comprehend, process, and manipulate abstract concepts, which is crucial for complex problem-solving and decision-making tasks. In the left section of Table 6, we provide a comparison of InternVL 2.5’s performance on multidisciplinary reasoning-related benchmarks, including MMMU and MMMU-Pro .
Here, we test both direct-answer and CoT reasoning performance, reporting the higher score. The results suggest that our model achieves encouraging improvements over existing open-source models, such as LLaVA-OneVision , NVLM , VILA 1.5 , and Qwen2-VL , as well as notable progress compared to earlier versions of the InternVL2 series. Specifically, InternVL2.5-78B achieves a score exceeding 70 on the MMMU validation set, representing a 7.4-point improvement over InternVL2-Llama3-76B. These results indicate that our model’s performance is moving closer to that of some advanced closed-source models, such as GPT-4o , Claude-3.5-Sonnet , and Gemini-1.5-Pro . Additionally, through majority voting, the score of InternVL2-Llama3-76B on the MMMU benchmark is improved from 62.7 to 65.3 when using CoT. We observe a similar phenomenon in InternVL 2.5 as well, which demonstrates that test-time scaling can improve the CoT reasoning of MLLMs.
Mathematical reasoning reflects a higher-level reasoning capability and enhances the potential of MLLMs in scientific and engineering applications. In the right-hand section of Table 6, we present InternVL 2.5’s performance across four multimodal mathematical benchmarks. These results demonstrate significant progress over InternVL 2.0. Notably, InternVL2.5-78B achieved an accuracy of 72.3% on the MathVista test-mini set . Additionally, on the challenging OlympiadBench , the InternVL 2.5 series showed an overall improvement compared to the 2.0 series. We attribute part of this advancement to our data filtering pipeline. Specifically, we observed that the 2.0 models frequently encountered deadlocks during CoT reasoning, failing to reach correct final answers, while this issue has been mitigated in the 2.5 series.
2 OCR, Chart, and Document Understanding
We assess InternVL’s OCR, chart, and document understanding capabilities through a comprehensive evaluation on a variety of OCR-related datasets.
AI2D : AI2D is a dataset of over 5,000 elementary school science diagrams, each with detailed annotations and corresponding multiple-choice questions. For a fair comparison, we report results for both “mask” and “no mask” settings on the test set.
ChartQA : ChartQA is a dataset focused on assessing models’ abilities to interpret and reason with data visualizations such as charts and graphs. Our evaluation metric is the average relaxed accuracy across both human and augmented test sets in ChartQA.
TextVQA : TextVQA is a dataset designed to benchmark visual reasoning based on text within images. It requires models to read and interpret text in images to accurately answer related questions. We report the VQA accuracy on the TextVQA validation set.
DocVQA : DocVQA is a dataset aimed at evaluating models’ ability to comprehend and retrieve information from text within document images. Performance is reported on the test set using the ANLS metric, which captures answer accuracy by measuring text similarity.
InfoVQA : InfographicVQA is a dataset aimed at evaluating models’ ability to interpret and reason with complex infographics that combine text, graphics, and visual elements. Performance is measured using the ANLS metric on the test set.
OCRBench : OCRBench evaluates the OCR capabilities of MLLMs across five tasks: text recognition, scene text VQA, document VQA, key information extraction, and handwritten math expression recognition, with a maximum score of 1000.
SEEDBench-2-Plus : SEED-Bench-2-Plus evaluates MLLMs on text-rich visual tasks, with 2,300 human-annotated questions across charts, maps, and webs. We report the average accuracy on this dataset.
CharXiv : CharXiv is a comprehensive evaluation suite featuring 2,323 charts from scientific papers. It includes two types of questions: reasoning questions (RQ) requiring synthesis of complex visual information, and descriptive questions (DQ) assessing basic chart element understanding.
VCR : Visual Caption Restoration (VCR) is a task that involves restoring partially hidden text within images by understanding both the visual content and the text. We report the Exact Match (EM) score and Jaccard similarity on the VCR-EN-Easy subset.
2.2 Evaluation Results
Table 7 provides a detailed comparison of InternVL 2.5 with its predecessor InternVL 2.0, other representative open-source models (e.g., Qwen2-VL , LLaVA-OneVision ), and closed-source models (e.g., GPT-4o , Claude-3.5-Sonnet ) on OCR-related tasks. Across most benchmarks, InternVL 2.5 achieves significant improvements over InternVL 2.0 at all model scales and demonstrates performance comparable to the current state-of-the-art model, Qwen2-VL-72B , reflecting the effectiveness of the improvements in training strategies and data quality.
However, at the 2B scale, InternVL2.5-2B underperforms compared to Qwen2-VL-2B on benchmarks such as TextVQA , DocVQA , and InfoVQA . We suspect that, in addition to differences in data and training strategies, model architecture may also play a significant role. Specifically, Qwen2-VL-2B features a 600M vision encoder and a 1.5B language model, whereas InternVL2.5-2B employs a smaller 300M vision encoder paired with a 1.8B language model. It appears that, for a smaller-scale MLLM (e.g., 2B), the size of the vision encoder plays a relatively important role in OCR performance, given the same total parameter budget.
Additionally, InternVL 2.5 demonstrates exceptional performance on the visual caption restoration (VCR) task . The 2.5 series achieves a significant improvement over InternVL 2.0 on this task, with the 2B model reaching EM/Jaccard scores of 93.2/97.6, far surpassing the previous generation’s 32.9/59.2. This improvement can be attributed to the introduction of a small portion of the VCR training set (approximately 22K samples). We find that the model’s poor performance on VCR tasks was not due to inadequate OCR capabilities but rather to its insufficient instruction-following ability for task-specific directives. By leveraging these few but focused samples, InternVL 2.5 exhibits a remarkable enhancement in its instruction-following ability for the VCR task, resulting in a substantial performance boost.
3 Multi-Image Understanding
We assess InternVL’s capabilities in multi-image relation perception and understanding across various multi-image benchmarks.
BLINK : The BLINK benchmark evaluates the core visual perception capabilities of MLLMs through 14 tasks inspired by classic computer vision challenges. Over half of the questions involve multiple images. Our results are reported on the validation set.
Mantis-Eval : Mantis-Eval is a meticulously curated small-scale benchmark for evaluating MLLMs’ reasoning capabilities across multiple images. It comprises 217 challenging, human-annotated problems covering topics such as size perception and weight comparison.
MMIU : MMIU is an extensive benchmark suite developed to rigorously assess the performance of MLLMs in multi-image tasks. It encompasses 7 distinct types of multi-image relationships and spans 52 diverse tasks, providing a comprehensive framework for evaluation.
MuirBench : MuirBench is a comprehensive benchmark for evaluating MLLMs capabilities in multi-image understanding. It spans 12 tasks and 10 types of multi-image relations and enhances model assessment with unanswerable instance variants.
MMT-Bench : MMT-Bench evaluates MLLMs on multimodal tasks like driving and navigation, focusing on recognition, reasoning, and planning, with many sub-tasks requiring multi-image understanding. To speed up testing, results are reported on the validation set.
MIRB : MIRB is a benchmark designed to evaluate the ability of MLLMs to understand and reason across multiple images. It contains four task categories: perception, visual world knowledge, reasoning, and multi-hop reasoning. The reported performance is the average score across these four categories.
3.2 Evaluation Results
As multi-image content becomes an increasingly common form of information exchange on the internet, it is essential for models to possess the ability to simultaneously understand and analyze relationships between multiple images. In the left part of Table 8, we evaluate the multi-image understanding capabilities of InternVL 2.5 across six diverse benchmarks: BLINK , Mantis-Eval , MMIU , MuirBench , MMT-Bench , and MIRB . These benchmarks test a range of skills, including reasoning across images, integrating information, and addressing task-specific requirements.
InternVL 2.5 achieves consistent improvements over InternVL 2.0 across all model scales, reflecting enhanced reasoning ability and better integration of multi-image information. For instance, at the 2B scale, InternVL2.5-2B delivers significant gains on Mantis-Eval (54.8 vs. 48.4) and MuirBench (40.6 vs. 32.5). These advancements can be largely attributed to the inclusion of additional multi-image datasets, as detailed in Section 4.5. These datasets, which were carefully curated and of high quality, played a critical role in improving the model’s ability to understand and reason across multiple visual inputs.
At larger scales, InternVL 2.5 demonstrates substantial progress and achieves competitive performance with advanced closed-source models. For example, InternVL2.5-78B scores 55.8 on MMIU, closely matching GPT-4o’s 55.7, and achieves a score of 70.8 on MMT-Bench, surpassing GPT-4o’s 65.4. These results highlight the importance of scaling model size and incorporating high-quality training data specifically tailored for multi-image tasks. However, on BLINK and MuirBench, our model still exhibits a performance gap of around 5 points compared to GPT-4o , suggesting that further improvements are needed, potentially through the inclusion of additional high-quality multi-image training data.
4 Real-World Comprehension
We assess InternVL’s performance on a suite of real-world benchmarks designed to evaluate its capabilities on realistic and complex tasks.
RealWorldQA : RealWorldQA is a benchmark designed to evaluate the real-world spatial understanding capabilities of MLLMs. It contains more than 700 images, each accompanied by a question and a verifiable answer, from various real-world scenarios.
MME-RealWorld : MME-RealWorld is a benchmark for evaluating MLLMs on complex, high-resolution image tasks across 43 real-world scenarios in 5 domains. Here, we test the English full set of the dataset.
WildVision : WildVision-Bench is a benchmark designed to evaluate MLLMs in the wild with human preferences. It comprises 500 high-quality samples meticulously curated from real-world user QA interactions. The benchmark uses a win rate metric to quantify the performance of models, providing insights into their ability to meet human expectations in practical applications.
R-Bench : R-Bench is a benchmark designed to evaluate the robustness of MLLMs against real-world image distortions, measuring their resilience in handling corrupted images in practical scenarios. We report the absolute robustness overall score for the MCQ task, which is the average score across low, mid, and high difficulty levels, corresponding to “R-Bench-Dis” in VLMEvalKit.
4.2 Evaluation Results
Given the complexity and dynamic nature of real-world environments, models must be robust enough to handle a wide range of challenging conditions. As shown in the right part of Table 8, InternVL 2.5 achieves leading performance across four real-world understanding benchmarks, including RealWorldQA , MME-RealWorld , WildVision , and R-Bench , and significantly outperforms the previous version, InternVL 2.0. This indicates that InternVL 2.5 has a stronger potential for practical application in complex and ever-changing real-world scenarios.
In benchmarks like RealWorldQA, MME-RealWorld, and R-Bench, which involve multiple-choice questions, InternVL 2.5 demonstrates strong real-world perceptual and understanding abilities. Differently, the WildVision benchmark uses GPT-4o as the judge model to evaluate the performance of various MLLM against the reference model, Claude-3-Sonnet . In this benchmark, the model’s output quality and user experience are key metrics. Although InternVL2.5-78B performs well in providing concise answers, it still shows a gap when generating longer responses to match human preferences. Specifically, InternVL2.5-78B scores 71.4, while GPT-4o scores 80.6, indicating a notable difference in user experience.
These results indicate that, while InternVL 2.5 delivers accurate and concise responses across most tasks, there is potential for improvement in generating more personalized and detailed answers. Future work will focus on enhancing the model’s performance in open-ended tasks and complex interactions, aiming to better align with human preferences, bridge the gap in user experience with GPT-4o.
5 Comprehensive Multimodal Evaluation
We evaluate InternVL’s comprehensive multimodal capabilities through a range of benchmarks, including:
MME : MME is the first comprehensive evaluation benchmark designed for MLLMs. It assesses models’ perception and cognitive abilities across 14 subtasks, including object presence, counting, position, color recognition, as well as commonsense reasoning, numerical computation, text translation, and code reasoning. We report the overall score across all tasks.
MMBench : MMBench evaluates the multimodal understanding of MLLMs through nearly 3,000 multiple-choice questions spanning 20 dimensions. It supports both English and Chinese versions, and we present the model’s performance scores on the test set.
MMBench v1.1 : Compared to MMBench, MMBench v1.1 features a refined dataset with a small number of noisy or low-quality questions removed, resulting in a subtle improvement in overall data quality. We report the model’s performance on the English version of the test set.
MMVet : MMVet is a benchmark designed to assess the integrated capabilities of MLLMs on complex tasks. It evaluates six core competencies: recognition, knowledge, spatial awareness, language generation, OCR, and mathematics, across 16 integrated tasks. Note that VLMEvalKit uses GPT-4-Turbo as the scoring model for this benchmark, which yields slightly lower scores compared to the official evaluation server.
MMVet v2 : Expanding on MMVet, MMVet v2 introduces an enhanced benchmark with a new capability: image-text sequence understanding, allowing for the assessment of models’ ability to process interleaved content. Here, we utilize the official evaluation server for scoring, which employs GPT-4-0613 as the scoring model.
MMStar : MMStar is a benchmark for evaluating the multimodal capabilities of MLLMs. It includes 1,500 carefully curated samples focusing on advanced visual and language understanding, minimizing data leakage, and emphasizing visual dependency.
5.2 Evaluation Results
Comprehensive multimodal evaluation benchmarks, such as MME , the MMBench series , the MMVet series , and MMStar , provide valuable and widely adopted frameworks for assessing model performance across a diverse set of multimodal tasks.
As shown in the left section of Table 9, the InternVL 2.5 models consistently outperform the previous InternVL 2.0 series across various model sizes, especially for smaller models with 1B-8B parameters. For example, in the MMBench v1.0 benchmark, which evaluates tasks in both English and Chinese, the InternVL 2.5 models show significant improvements. The InternVL2.5-4B achieves a score of 81.1/79.3, surpassing the InternVL2-4B’s 78.6/73.9, while the InternVL2.5-8B reaches 84.6/82.6, outperforming the InternVL2-8B’s 81.7/81.2.
It is also noteworthy that, while we have significantly improved the performance of smaller models on the MMVet series benchmarks, our largest model, InternVL2.5-78B, still does not surpass the Qwen2-VL-72B . Currently, the state-of-the-art models on MMVet v2 remain closed-source models like GPT-4o and Claude-3.5-Sonnet . This highlights the gap between open-source models and closed-source ones in multimodal integrated capability. We recognize this as an important direction for future development.
6 Multimodal Hallucination Evaluation
We evaluate InternVL’s tendency toward hallucinations across four different benchmarks, including:
HallusionBench : HallusionBench is a benchmark for evaluating image-context reasoning in MLLMs through a Yes/No judgment question format, focusing on challenges such as language hallucination and visual illusion. We report performance using the average scores of its three metrics: aAcc, fAcc, and qAcc.
MMHal-Bench : MMHal-Bench is a benchmark designed to evaluate hallucinations in MLLMs. It includes 96 challenging questions derived from images in the OpenImages dataset, along with their corresponding ground-truth answers and image content. Scoring is conducted using GPT-4o, with scores ranging from 0 to 6.
CRPE : CRPE is a benchmark that measures the hallucination level of the relation between objects using multiple-choice questions. We report accuracy on the relation subset for this benchmark.
POPE : POPE is a benchmark for evaluating object hallucination in MLLMs, utilizing binary questions to quantify and analyze hallucination tendencies. We report the average F1 score across three categories: random, popular, and adversarial.
6.2 Evaluation Results
We evaluate the performance of InternVL on four key hallucination evaluation benchmarks: HallusionBench , MMHal , CRPE , and POPE . These benchmarks assess the frequency of hallucinations, or factual inaccuracies, across multimodal tasks, providing a measure of model reliability in handling complex inputs like text and images.
The InternVL 2.5 models show significant progress over the InternVL 2.0 series, particularly in smaller models (e.g., 1B-8B parameters). For instance, InternVL2.5-1B and InternVL2.5-2B demonstrate improved scores on all hallucination benchmarks, with the 1B model achieving a 39.0 score on HallusionBench, up from 34.0 in the earlier version. Similarly, the 2B model improved to 42.6, outperforming the previous 2B model by nearly 5 points. These results indicate substantial gains in reducing hallucinations while handling multimodal data.
The largest model, InternVL2.5-78B, also shows improvements, reducing hallucinations compared to both prior versions and other leading models. It scores 57.4 on HallusionBench, competing with top models like Qwen2-VL-72B (58.1) and GPT-4o (55.0). Although InternVL2.5-78B demonstrates relatively low hallucination rates on these hallucination evaluation benchmarks, some hallucinations are still inevitably present when generating long responses in practical use. This is a challenge we plan to tackle in future work.
7 Visual Grounding
We evaluate InternVL’s visual grounding capability via referring expression comprehension (REC) on the RefCOCO, RefCOCO+, and RefCOCOg datasets, where the model identifies target objects in images from given descriptions.
RefCOCO : Built on COCO, this dataset contains 19,994 images with 142,210 referring expressions for 50,000 objects, split into subsets like test A (people-focused) and test B (other objects) for REC tasks.
RefCOCO+ : Similar to RefCOCO but emphasizing attribute-based descriptions without absolute location cues. It includes 19,992 images and 141,564 expressions, requiring models to focus on descriptive attributes.
RefCOCOg : With 25,799 images and 95,010 expressions, this dataset features longer, more complex expressions, and challenging models to manage intricate language in REC tasks.
7.2 Evaluation Results
Visual grounding is critical for connecting textual descriptions with visual content, enabling accurate multimodal interaction. Table 10 compares InternVL 2.5 with its predecessor, InternVL 2.0, at the 8B and 78B scales, alongside other leading MLLMs (e.g., CogVLM-Grounding-17B , Qwen2-VL ) and specialized grounding models (e.g., Grounding-DINO-L , UNINEXT-H , ONE-PEACE ), evaluated on the RefCOCO , RefCOCO+ , and RefCOCOg datasets.
InternVL2.5-8B improves its predecessor’s performance, with the average score rising from 82.9 to 87.6, achieving comparable results to Qwen2-VL-7B (87.6 vs. 87.9), though slightly behind Ferret-v2-13B and CogVLM-Grounding-17B , which benefit from fine-tuning for grounding and larger model sizes. At the larger scale, InternVL2.5-78B achieves state-of-the-art performance with an average score of 92.3, a 2.3-point improvement over InternVL2-Llama3-76B, surpassing Qwen2-VL-72B . These gains highlight the effectiveness of our data and training optimizations, significantly enhancing localization capabilities.
8 Multimodal Multilingual Understanding
We assess InternVL’s multimodal multilingual understanding capabilities using three representative benchmarks:
MMMB and Multilingual MMBench : MMMB is a large-scale multilingual multimodal benchmark with 6 languages, 15 categories, and 12,000 questions. The languages evaluated are English (en), Chinese (zh), Portuguese (pt), Arabic (ar), Turkish (tr), and Russian (ru). Multilingual MMBench extends MMBench to these 6 languages using GPT-4 translation for multilingual understanding evaluation.
MTVQA : MTVQA is a multilingual benchmark tailored for text-centric visual question answering. It includes high-quality, expert human annotations across nine languages, specifically addressing the “visual-text misalignment” challenge in multilingual contexts. We report the average score of MTVQA.
8.2 Evaluation Results
Multilingual ability is critical for MLLMs as it expands their application and improves cross-language communication. For global deployment, MLLMs must effectively handle both high-resource and low-resource languages. As shown in Table 11, we evaluated our model’s performance on three multilingual benchmarks: MMMB , Multilingual MMBench , and MTVQA .
A comparison between InternVL2.5-78B and Qwen2-VL-72B reveals that, despite differences in training data, model architecture, and training strategies, their multilingual performance is quite similar. This suggests that the multilingual capabilities of MLLMs are largely inherited from the underlying language model. Both models share the same LLM, indicating that a strong multilingual LLM forms the foundation for effective multilingual performance in MLLMs.
9 Video Understanding
Video-MME : Video-MME is a benchmark for evaluating MLLMs in full-spectrum video analysis. It features a wide variety of video types across multiple domains and durations, with multimodal inputs including video, subtitles, and audio. For this benchmark, we test with four settings: 16, 32, 48, and 64 frames, and report the maximum results. We report results for both “with subtitle” and “without subtitle” settings.
MVBench : MVBench is a video understanding benchmark designed to comprehensively evaluate the temporal awareness of MLLMs in the open world. It covers 20 challenging video tasks, ranging from perception to cognition, which cannot be effectively solved using a single frame. We test this benchmark using 16 frames.
MMBench-Video : MMBench-Video is a quantitative benchmark for evaluating MLLMs’ video understanding and temporal reasoning skills, covering diverse domains, multi-shot long videos, and features like hallucination, commonsense reasoning, and temporal reasoning. For this benchmark, we test with four different settings: 16, 32, 48, and 64 frames, and report the maximum scores.
MLVU : MLVU is a comprehensive benchmark designed to evaluate MLLMs in long video understanding tasks, featuring videos ranging from 3 minutes to 2 hours. It includes nine different evaluation tasks divided into three categories: holistic understanding, single-detail understanding, and multi-detail understanding. We evaluate four settings: 16, 32, 48, and 64 frames, and report the highest “M-Avg” results.
LongVideoBench : LongVideoBench focuses on referring reasoning tasks that involve long-frame inputs, requiring the model to accurately retrieve and reason about detailed multimodal information based on referring queries. We test four settings—16, 32, 48, and 64 frames—and report the best results on the validation set.
CG-Bench : CG-Bench is a benchmark for evaluating long video understanding in MLLMs. Unlike existing benchmarks, it focuses on models’ ability to retrieve relevant clues for answering questions. It includes 1,219 curated videos and over 12,000 question-answer pairs. Two novel clue-based evaluation methods are introduced to assess genuine video understanding. We test this benchmark using 32 frames.
9.2 Evaluation Results
Video understanding is vital for assessing MLLMs’ ability to process temporal and multimodal information. To evaluate this comprehensively, we tested six benchmarks: Video-MME , MVBench , MMBench-Video , MLVU , LongVideoBench , and CG-Bench , covering diverse tasks from short video comprehension to long video reasoning.
As shown in Table 12, InternVL 2.5 achieves consistent improvements over InternVL 2.0 across all benchmarks. For example, our smallest model, InternVL2.5-1B improves Video-MME scores from 42.9/45.4 to 50.3/52.3 and MVBench from 57.5 to 64.3. Moreover, we find that InternVL 2.5 demonstrates better scalability when handling increasing input frames compared to its predecessor, as shown in Figure 10. We attribute these improvements to two key enhancements: (1) The inclusion of more high-quality video data, which has significantly enhanced the model’s video understanding capabilities. (2) Adjusting the training frame sampling strategy from 4–24 to 8–32 frames (as shown in Figure 5(c)) enhanced the model’s ability to process richer video information. Consequently, while InternVL 2.0 models typically perform best at 16 or 32 frames but degrade with more input frames, InternVL 2.5 could benefit from increasing input frames, showing better scalability for long video understanding.
Our largest model, InternVL2.5-78B, achieves leading performance among open-source models and approaches the performance of closed-source systems. Compared to open-source models, InternVL2.5-78B surpasses Qwen2-VL-72B on MVBench (76.4 vs. 73.6) and MMBench-Video (1.97 vs. 1.70), though its Video-MME score with subtitles is slightly lower (74.0 vs. 77.8). Against closed-source models like GPT-4o and Gemini-1.5-Pro , InternVL2.5-78B demonstrates competitive performance. On Video-MME, it scores 72.1/74.0, closely matching GPT-4o (71.9/77.2) and Gemini-1.5-Pro (75.0/81.3). However, on LongVideoBench, it achieves 63.6, slightly trailing Gemini-1.5-Pro (64.0) and GPT-4o (66.7). This highlights the remaining challenges in long video understanding for open-source models, indicating room for further improvement.
Evaluation on Language Capability
To thoroughly assess the language capabilities of LLMs and MLLMs, we evaluate their performance across five core dimensions using a diverse set of datasets. These benchmarks encompass tasks like comprehensive examination, language and knowledge, reasoning, mathematics, and coding.
Comprehensive Examination. We conduct a thorough evaluation of LLMs and MLLMs using various exam-related datasets: (1) MMLU includes 57 subtasks covering diverse topics such as humanities, social sciences, and STEM, evaluated with a 5-shot approach. (2) CMMLU , focused on a Chinese context, features 67 subtasks spanning general and Chinese-specific domains, also tested in a 5-shot setting. (3) C-Eval contains 52 subtasks across four difficulty levels, evaluated in a 5-shot setting. (4) GAOKAO-Bench , derived from Chinese college entrance exams, offers comprehensive coverage of both subjective and objective question types, with objective questions evaluated in a 0-shot setting.
Language and Knowledge. For language and knowledge-based assessments, we use a range of datasets designed to test the capabilities: (1) TriviaQA , which includes both reading comprehension and open-domain QA tasks with multiple answers per question, evaluated in a 0-shot setting. (2) NaturalQuestions , featuring user-generated questions validated by experts, also evaluated in a 0-shot manner. (3) C3 , a free-form multiple-choice Chinese machine reading comprehension dataset, with 0-shot results reported. (4) RACE , a reading comprehension dataset containing English exam questions for Chinese middle and high school students aged 12 to 18, with results reported for the high school subset in a 0-shot setting.
Reasoning. To measure reasoning capabilities, we use datasets like (1) WinoGrande , which tests commonsense reasoning through 44,000 multiple-choice questions requiring pronoun disambiguation, evaluated in a 0-shot setting. (2) HellaSwag challenges models with natural language inference scenarios and four outcome options, demanding selection of the most logical conclusion, also evaluated in a 0-shot manner. (3) BigBench Hard (BBH) comprises 23 tasks specifically chosen for their difficulty in surpassing human performance, further evaluating reasoning depth, with 0-shot results reported.
Mathematics. In the domain of mathematics, (1) GSM8K-Test offers approximately 1,300 elementary-level situational problems, evaluated in a 4-shot setting. (2) MATH presents 12,500 high school competition-level problems across subjects like algebra and calculus, each with detailed solutions, also evaluated in a 4-shot manner. (3) TheoremQA introduces 800 STEM-focused problems requiring theorem application in fields like mathematics, physics, and finance, with 0-shot results reported.
Coding. To evaluate coding capabilities, we employ the following benchmarks: (1) HumanEval : This benchmark includes 164 Python programming tasks, each paired with detailed specifications, serving as a standard for assessing coding performance. It is evaluated in a 4-shot setting. (2) MBPP : Comprising 974 entry-level programming tasks, MBPP covers a wide range of challenges, from simple arithmetic problems to more complex sequence definitions, evaluated in a 3-shot setting. (3) MBPP-CN : A Chinese adaptation of MBPP designed to assess multilingual programming capabilities. This extension broadens the evaluation scope to include linguistic and contextual diversity, with 0-shot results reported.
2 Evaluation Results
In the development of MLLMs, maintaining strong pure language capabilities remains critically important. Following the approach of InternLM2 , we conducted a comprehensive evaluation of our models’ performance across 17 pure language benchmarks using the OpenCompass toolkit . These benchmarks are categorized into five major groups, providing a thorough assessment of the models’ pure language abilities.
The results show that InternVL 2.0 demonstrates a slight decline in pure language performance compared to its foundational LLM counterparts. For example, InternVL2-2B achieved an average score of 39.2, a decrease of 2.1 points compared to InternLM2-1.8B-Chat. Similarly, InternVL2-8B scored an average of 67.2, 2.3 points lower than InternLM2.5-7B-Chat.
To address this issue, we curated a large collection of high-quality open-source pure language instruction data and applied rigorous filtering pipelines to eliminate low-quality samples, thereby enhancing the overall data quality. These improvements in InternVL 2.5 have effectively mitigated the decline in language performance, enabling the model to match or even surpass the original LLM in several tasks. This demonstrates that supplementing and optimizing with high-quality language data can not only preserve MLLM’s pure language capabilities but also establish a stronger foundation for multimodal tasks.
Evaluation on Vision Capability
In this section, we present a comprehensive evaluation of the vision encoder’s performance across various domains and tasks. The evaluation is divided into two key categories: (1) image classification, representing global-view semantic quality, and (2) semantic segmentation, capturing local-view semantic quality. This approach allows us to assess the representation quality of InternViT across its successive version updates.
We assess the global-view semantic quality of InternViT through a comprehensive evaluation on diverse image classification datasets.
ImageNet-1K : A widely-used large-scale dataset containing over 1 million images across 1,000 classes, commonly used for benchmarking image classification models.
ImageNet-ReaL : A re-labeled version of ImageNet’s validation set, providing multi-label annotations that are more accurate and robust, following an enhanced labeling protocol.
ImageNet-V2 : A dataset designed to evaluate the robustness of models trained on ImageNet-1K, featuring new test images collected using the original ImageNet methodology.
ImageNet-A : A challenging dataset of naturally occurring, unmodified images that are often misclassified by ResNet models. It highlights the limitations of models when exposed to adversarially difficult examples in real-world settings.
ImageNet-R : A rendition dataset with 30K images across 200 ImageNet classes, composed of art, sketches, toys, sculptures, and other creative representations. It assesses the robustness of models in recognizing abstract renditions of common objects.
ImageNet-Sketch : This dataset contains 51K sketch images, with approximately 50 sketches per ImageNet class. It is constructed via Google Image queries using the class name followed by “sketch of,” testing a model’s ability to generalize to abstract, hand-drawn representations.
1.2 Settings
In this study, two evaluation methods, linear probing and attention pooling probing, are employed to assess the performance of the InternViT models:
Linear Probing : This method involves freezing the pre-trained model and training only a linear classifier on top. It evaluates the quality of the learned features without updating the backbone, providing insights into how effectively the pre-trained model captures semantic information usable by a simple linear classifier in downstream tasks like image classification.
Attention Pooling Probing: In contrast, attention pooling probing evaluates the model by adding an attention pooling layer on top of the frozen features. This approach allows the vision encoder to retain richer information in the final layer, as attention pooling can dynamically select task-relevant features for classification without interference from unrelated information.
For both experiments, we use ImageNet-1K as the training set and evaluate the models on the ImageNet-1K validation set along with several ImageNet variants (i.e., ImageNet-ReaL , ImageNet-V2 , ImageNet-A , ImageNet-R , and ImageNet-Sketch ) to benchmark their domain generalization capabilities.
The models are trained using SGD as the optimizer, with a peak learning rate of 0.2, a momentum of 0.9, and no weight decay. A cosine learning rate decay schedule is applied over 10 training epochs, with 1 warmup epoch. We use input resolutions of 448448, with a patch size of 14 and a total batch size of 1024. Data augmentation techniques, such as random resized cropping and horizontal flipping, are employed during training. The code and logs of these classification experiments will be released on our GitHub repositoryhttps://github.com/OpenGVLab/InternVL/tree/main/classification.
1.3 Evaluation Results
As shown in Table 14, the results reveal an interesting trend across the version updates of InternViT: as the model progresses, the performance of linear probing declines substantially, with all versions showing an average below the gray baseline. In contrast, attention pooling probing consistently outperforms the gray baseline despite some fluctuations. This results in a growing trend in the average score difference (from 3.5 to 6.7), denoted as , across successive InternViT versions.
This suggests that features in the model’s final layer become less linearly separable, likely as representations evolve to capture more complex, open-ended semantic information. The attention pooling mechanism effectively selects relevant features from this enriched representation space, offsetting challenges from reduced linear separability. Additionally, these findings imply that InternViT maintains key pre-training attributes through iterative updates without catastrophic forgetting. With each version, its representations grow more diverse, capturing open-set semantics and enhancing generalization—an advantage particularly valuable for MLLMs requiring high abstraction for real-world tasks.
2 Semantic Segmentation
We evaluate the local-view semantic quality of InternViT using two representative semantic segmentation datasets, ADE20K and COCO-Stuff-164K.
ADE20K : A comprehensive dataset containing over 20,000 images with annotations across 150 object and background categories, widely used for scene parsing. It provides detailed pixel-level labels for both objects and parts, facilitating a range of fine-grained segmentation tasks.
COCO-Stuff-164K : An extension of the original COCO images with pixel-level annotations, adding 91 “stuff” classes (like grass and sky) to 80 “thing” categories (like people and cars), covering a total of 172 classes. With these comprehensive labels, the dataset supports tasks in scene parsing and semantic segmentation, enabling richer context understanding in image analysis.
2.2 Settings
In this study, three evaluation methods—linear probing, head tuning, and full tuning—are employed to assess the performance of the InternViT models on semantic segmentation tasks:
Linear Probing: Linear probing applies a frozen backbone with a linear segmentation head, offering insight into the linear separability of learned features. This method provides a baseline for evaluating pixel-level semantic information with minimal adaptation, though it may not fully capture the encoder’s capacity for complex features.
Head Tuning: In head tuning, the InternViT is frozen while the UperNet head remains trainable, allowing the model to utilize a stronger head to reduce its dependence on linear separability. This setup mitigates the decline in linear separability caused by the complex, open-ended features, enabling a more precise evaluation of the vision encoder’s capabilities.
Full Tuning: Full tuning involves making both the InternViT backbone and the UperNet segmentation head trainable, allowing the model to adapt all layers for the target task and minimizing reliance on pre-existing linear separability. This setup provides an alternative perspective for evaluating the vision encoder’s capacity to extract visual features.
We use AdamW with a peak learning rate of 4e-5 and a polynomial decay schedule. Layer-wise learning rate decay (0.95) is applied in full tuning. Weight decay is set to 0.05 for both head and full tuning, and none for linear probing. The input resolution is 504504, with a patch size of 14 and a batch size of 16. Training consists of 1.5K warmup iterations and 80K total iterations. A drop path rate of 0.4 is applied in full tuning. We utilize default data augmentation from MMSegmentation . All the code and logs related to these experiments will be released on GitHubhttps://github.com/OpenGVLab/InternVL-MMDetSeg.
2.3 Evaluation Results
As shown in Table 15, the semantic segmentation performance of InternViT models is evaluated across three configurations—linear probing, head tuning, and full tuning—on ADE20K and COCO-Stuff-164K . The results reveal distinct trends in how the models’ feature representations evolve across version updates.
Linear probing results show a decline in mIoU scores as the model versions progress, with average scores dropping from 45.0 in InternViT-6B-224px to 37.5 in InternViT-6B-448px-V2.5. This indicates that as InternViT updates, the features become less linearly separable, reflecting a shift toward capturing more complex and open-ended information.
In head tuning, the models display a different trend compared to linear probing. All other versions of InternViT surpass the baseline InternViT-6B-224px’s mIoU score of 51.9, showing no performance decline. This leads to increasing values, growing from 6.9 in InternViT-6B-224px to 15.1 in InternViT-6B-448px-V2.5. The rise in suggests that while the features become less linearly separable, their quality remains intact, effectively capturing complex information. Similarly, full tuning yields consistent results, as seen in the values. The increase in from 10.2 in InternViT-6B-224px to 17.7 in InternViT-6B-448px-V2.5 further supports this trend.
Overall, the increasing values of and across model versions highlight the shift from simple, linearly separable features to more complex, nonlinear representations. This evolution aligns with InternViT’s growing capability to extract visual information as its versions progress within the development of InternVL. It demonstrates the effectiveness of our ViT incremental learning strategy in enhancing the vision encoder’s ability to extract open-ended features.
Conclusion
In this work, we introduce InternVL 2.5, an advanced open-source multimodal large language model (MLLM) series that builds upon the architecture of InternVL 2.0 with significant improvements in training, testing strategies, and data quality. We systematically explore the relationship between model scaling and performance, analyzing vision encoders, language models, dataset sizes, and test-time configurations. Extensive evaluations on diverse benchmarks demonstrate that InternVL 2.5 achieves competitive performance across tasks such as multi-discipline reasoning, document understanding, video understanding, multilingual processing, etc. Notably, it is the first open-source MLLM to surpass 70% on the MMMU benchmark, narrowing the gap between open-source and commercial models like OpenAI o1. By sharing InternVL 2.5 with the community, we hope to contribute a powerful tool for advancing multimodal AI research and applications, and we look forward to seeing future developments building upon this work.
Acknowledgement
This work is supported by the National Key R&D Program of China (No. 2022ZD0160102, 2022ZD0161300), the National Natural Science Foundation of China (No. 62376134, 62372223, U24A20330), the China Mobile Zijin Innovation Institute (No. NR2310J7M), and the Youth PhD Student Research Project under the National Natural Science Foundation (No. 623B2050).