Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
Zhangwei Gao, Zhe Chen, Erfei Cui, Yiming Ren, Weiyun Wang, Jinguo Zhu, Hao Tian, Shenglong Ye, Junjun He, Xizhou Zhu, Lewei Lu, Tong Lu, Yu Qiao, Jifeng Dai, Wenhai Wang
Introduction
In recent years, there have been significant advancements in multimodal large language models (MLLMs) , which leverages the powerful capabilities of pre-trained large language models (LLMs) alongside vision foundation models (VFMs) . These models undergo multi-stage training on extensive image-text data, which effectively aligns visual representations from VFMs with the latent space of LLMs, leading to promising performance in general vision-language understanding, reasoning, and interaction tasks. However, the large computational burden and the poor performance on long-tail domain-specific tasks hindered the widespread application of MLLMs in practical scenarios.
The emergence of lightweight MLLMs has provided a good balance between parameter size and performance, alleviating the reliance on expensive computing devices and fostering the development of various downstream applications. However, there are still several challenges: (1) Most existing MLLMs use vision encoders like CLIP , which are trained on Internet-domain image-text data and are aligned with BERT . As a result, these vision encoders are not capable of covering the extensive range of visual domains and are misaligned with LLMs’ representations. (2) To adapt MLLMs to specialized domains, existing methods mainly focus on modifying the model architectures, gathering extensive related training data, or customizing the training process for the target domain. There is still no consensus framework for LLMs’ downstream adaptation. Different domains have different model designs, data formats, and training schedules.
To address these issues, there is a need for a strong vision encoder with comprehensive visual knowledge as well as a general transfer learning paradigm that allows for efficient application across downstream tasks in various domains at a low marginal cost.
In this work, we introduce Mini-InternVL, a series of powerful pocket-sized MLLMs that can be easily transferred to various specialized domains. To this end, we first enhance the representational capabilities of a lightweight vision encoder. We initialize a 300M vision encoder using the weights from CLIP and apply knowledge distillation using InternViT-6B as the teacher model. Subsequently, we develop Mini-InternVL series with 1 billion, 2 billion, and 4 billion parameters, by integrating the vision encoder with the pre-trained LLMs such as Qwen2-0.5B , InternLM2-1.8B , and Phi-3-Mini , respectively. Benefiting from the robust vision encoder, Mini-InternVL exhibit excellent multimodal performance on general multimodal benchmarks like MMBench , ChartQA , and MathVista . Remarkably, compared with InternVL2-76B, the proposed Mini-InternVL-4B achieves 90% of the performance of larger counterparts while using 5% fewer parameters, significantly reducing computational overhead.
To further adapt our models to specific-domain downstream tasks, we introduce a straightforward yet effective transfer learning paradigm. Within this paradigm, we developed a unified transfer approach applicable to various downstream tasks, including autonomous driving, medical images, and remote sensing. This approach standardizes the model architecture, data format, and training schedule. The results demonstrate the effectiveness of this method in enhancing the model’s visual understanding and reasoning capabilities in domain-specific scenarios, enabling it to match the performance of proprietary commercial models within the target domains.
In summary, our contribution has three folds:
(1) We propose Mini-InternVL, a powerful pocket multimodal model, that not only achieves robust multimodal performance with fewer than 4 billion parameters but also easily transfers to downstream tasks across various domains at low marginal cost.
(2) We develop several design features for Mini-InternVL, including a lightweight vision encoder—InternViT-300M, that is robust for various visual domains. Additionally, we introduce a simple but effective paradigm that standardizes model architecture, data format, and training schedule for effective downstream task transfer.
(3) We comprehensively evaluate our models through extensive experiments on general and domain-specific benchmarks. These results show that our multimodal models achieve 90% of the performance using significantly fewer parameters on general multimodal benchmarks. For specific domain tasks, with minimal computational cost for fine-tuning, they can rival closed-source commercial models. We conduct a series of ablation studies to explore the impact of data sample size on domain adaptation, hoping to provide insights into the application of MLLMs in specialized domains.
Related Works
Benefiting from the advancement of LLMs, the MLLMs have also achieved great progress. Early works consider multi-modal understanding as one of the tool usage tasks and prompt the LLMs to ask other models to write a caption about the corresponding input modality so that LLMs could understand the multi-modal input. To effectively utilize the ability of pre-trained LLMs and VFMs, a series of works propose to use a connector to align the embedding space between them, which achieve promising performance under a controllable cost. Another series of work extend pre-trained LLMs with extra layers to fuse the vision features, which reduce the number of required visual tokens inputted into LLMs while introducing extra training cost.
Recently, some works, such as Fuyu , MoMa and Chameleon , propose a vision encoder-free architecture. This type of architecture consists of a single transformer model, which is used to process both visual and textual information simultaneously without requiring an additional encoder, making it more deployment-friendly. Despite these advancements, the heavy inference cost of these MLLMs hinders their application in downstream tasks. To address such issue, a series of lightweight MLLMs, such as MiniCPM-V , are proposed. However, since most of them use CLIP-L as the vision encoder, which is only trained on the natural image domain, these models are limited to the general domain and fail to generalize to other domains. In this work, we propose InternViT-300M, which is distilled from InternViT-6B and trained on a diverse image domain.
From a vision-centric perspective, most MLLMs utilize vision models such as CLIP and SigLIP , which are trained on large-scale web image-text data. However, such vision encoders face significant limitations in terms of parameter scale and representational ability. Several studies have explored this issue. For instance, Tong et al. identified significant differences in the visual patterns of CLIP and DINOv2 , leading to the development of a mixture-of-features module that integrates these two VFMs. LLaVA-HR introduced a dual-branch vision encoder that employs CLIP-ViT for low-resolution pathways and CLIP-ConvNext for high-resolution pathways. Similarly, DeepSeek-VL utilized a dual vision encoder design, incorporating SigLIP-L for low-resolution images and SAM-B for high-resolution images.
However, these methods involve excessively complex pathways, which complicates the practical application of the models. Moreover, such approaches do not resolve the issue of vision encoders lacking comprehensive visual knowledge across various domains. In contrast, InternViT implements progressive image-text alignment, and acquires representational capabilities across multiple domains by performing generative training on datasets spanning various fields. We propose injecting visual knowledge from the capable vision encoder into lightweight visual models, thus avoiding the computational expense associated with iterative generative pre-training.
Several methods have been explored to apply MLLMs to specific domains, such as GeoChat and EarthGPT for remote sensing, LLaVA-Med and Qilin-Med-VL for the medical images, ChemVLM for chemistry, DriveVLM , DriveMLM and DriveGPT4 for autonomous driving. Although these methods have achieved promising results, they involve modifications to model architectures, the collection of extensive domain-specific training data, or customization of the training process for the target domain. Nonetheless, there is still no universally accepted framework for the downstream adaptation of MLLMs. we propose a straightforward yet effective transfer learning paradigm, aiming to prevent significant disparities among MLLMs in different fields that hinder interoperability.
Method
In this section, we introduce Mini-InternVL, a series of lightweight multimodal large language models (MLLMs). Section 3.1 provides a comprehensive overview of Mini-InternVL. Then, Section 3.2 details InternViT-300M, a lightweight vision model developed through knowledge distillation, which inherits the strengths of a powerful vision encoder. Finally, Section 3.3 describes a transfer learning framework designed to enhance the model’s adaptation to downstream tasks.
As shown in Figure 1, Mini-InternVL consists of three main components: InternViT, MLP projector, and LLM. We employ InternViT-300M as our vision encoder, a lightweight vision model that inherits the capabilities of a powerful vision encoder. Based on InternViT-300M, we develop three versions of Mini-InternVL: Mini-InternVL-1B, Mini-InternVL-2B, and Mini-InternVL-4B. Each version is respectively connected to the pre-trained Qwen2-0.5B , InternLM2-1.8B , and Phi-3-mini . Similar to other open-source MLLMs , Mini-InternVL employs an MLP projector to connect the vision encoder and the LLMs.
We adopt a dynamic resolution input strategy similar to that of InternVL 1.5 , which improves the model’s ability to capture fine-grained details. We also apply a pixel unshuffle operation to reduce the number of visual tokens to one-quarter of the original. Consequently, in our model, a 448448 image is represented by 256 visual tokens, enabling it to process up to 40 image tiles (i.e., 4K resolution).
The training of Mini-InternVL consists of two stages: (1) Language-Image Alignment: We keep only the MLP component unfrozen during this stage. Following InternVL 1.5 , we use a diverse range of training datasets that encompass various tasks, including captioning, detection, grounding, and OCR. The diversity of these datasets ensures robust pre-training of Mini-InternVL, enabling the model to handle a variety of linguistic and visual elements across different tasks. (2) Visual Instruction Tuning: We carefully select datasets to enhance the model’s performance across a broad spectrum of multimodal tasks, similar to InternVL 1.5. These tasks include image captioning, chart interpretation, OCR, and cross-disciplinary reasoning. We conduct full-parameter fine-tuning with these datasets, further injecting world knowledge and teaching models to follow user instructions.
2 InternViT-300M
Most existing MLLMs employ vision encoders that are trained on web-scale image-text paired data, such as CLIP, to obtain their representations. These encoders lack comprehensive knowledge of the visual world, which needs to be acquired through iterative generative pre-training in conjunction with LLMs. Unlike other approaches that enhance the visual foundation models by using auxiliary pathways , our method directly leverages a powerful vision model that has undergone generative training on diverse datasets to transfer knowledge to a lightweight vision model. Specifically, we use InternViT-6B as the teacher model and initialize the student model’s weights using CLIP-ViT-L-336px. We align the representations of the student model with those of the teacher model by computing the negative cosine similarity loss between the hidden states of the last transformer layers. The resulting model is named InternViT-300M.
The primary goal of this knowledge transfer is to inherit the pre-training knowledge embedded in InternViT-6B. To achieve this, we curate a dataset sourced from a diverse range of publicly accessible resources, as detailed in Table 1. This dataset comprises four main types of data: natural images, OCR images, charts, and multi-disciplinary images. All images are resized to a resolution of 448 448, and dynamic resolution is disabled for training efficiency. Ultimately, we develop a vision encoder, termed InternViT-300M, which is infused with diverse knowledge and is adaptable to various language models.
3 Domain Adaptation
Although many studies have successfully applied MLLMs to downstream tasks, a universally accepted framework for adapting MLLMs to these applications has yet to be established. Differences in model design, data formats, and training strategies across various domains result in significant heterogeneity among MLLMs, making standardization challenging. To address this issue, we propose a straightforward yet effective transfer learning framework.
Instruction tuning is a crucial training stage to teach models to follow user instructions, of which the training data is formulated as visual question answering (VQA) and conversation format. VQA datasets of the downstream tasks, such as RSVQA and PMC-VQA , are directly utilized as instruction-following data. For other conventional tasks, as shown in Figure 2, we formulate them into VQA format according to the following approaches separately:
(1) Image Classification Tasks. In most traditional classification tasks within specialized domains, a wide range of technical terms are involved. In the majority of cases, we can easily format the classification task as a multiple-choice question. Given an image
USER: [Image][Prompt_Prefix][Candidate Labels][Prompt_Suffix] ASSISTANT: [Ground Truth]
A direct example can be seen in our approach to remote sensing image classification, where we utilize prompts such as “Classify the image within one of the given classes: dense residential area, ..., school. Answer with one word or short phrase.”, as shown in Figure 2. This method transforms image classification tasks into multiple-choice questions. For behavior prediction of the ego vehicle in autonomous driving data, we draw inspiration from DriveLM by employing templates like “Predict the behavior of the ego vehicle. Please select the correct answer from the following options: A. The ego vehicle is going straight. The ego vehicle is not moving. B. ...”.
(2) Visual Grounding Tasks. The native support for the visual grounding task in Mini-InternVL allows the use of a special token, , to enclose the name of the object to be detected. With this token, the model can be directed to provide the object’s location enclosed within
(3) Region Perception Tasks. Region-level conversation tasks are prevalent in specialized domains. These tasks involve supplying the model with spatial location information, in addition to the question input. The model is required to focus on objects within the specified attention region to generate a response. Specifically, there are two implementation methods. The first method involves directly annotating the location on the image using bounding boxes, masks, or contours, as illustrated in the “Region Perception2” of Figure 2. The second method denotes the object within the question by
For instance, in remote sensing applications, the goal is to train the model to identify objects within specific coordinates [x1, y1, x2, y2]. To achieve this, we use a prompt such as “What object is in this location
(4) Multi-View Images. In autonomous driving, the images are captured from six different viewpoints. As shown in Figure 3, we effectively utilize dynamic resolution to accommodate this type of data. Specifically, InternVL supports splitting images into 448448-sized tiles based on their aspect ratio. Consequently, we resize each image to 896448 pixels and then, as illustrated, combine these images in a fixed sequence, resulting in a final resolution of 2688896. This means that the images are automatically processed into 12 tiles, and an additional thumbnail is added to provide the model with global context. Furthermore, we labeled each viewpoint image with text indicating its camera position, such as “CAM_FRON”.
(5) Video Frames. InternVL supports video frames in an interleaved image format. We represent the frame sequence using a template such as “Frame1:
During the domain adaptation phase, we perform full-parameter fine-tuning on Mini-InternVL. For a domain-specific application scenario, we convert corresponding data into the required format and incorporate it into our training dataset. Adding a certain proportion of general multimodal data during the domain adaptation phase will not affect the performance in the specific domain, while retaining the model’s general multimodal capability. In our experiments, we find that adding general data can improve the generalization ability of the model on other tasks. Therefore, when performing domain adaptation, we can choose the appropriate general data ratio on the premise of balancing computational overhead and performance.
Experiments
In this section, we begin by conducting a comprehensive comparison of our Mini-InternVL with leading multi-modal large language models (MLLMs) on representative vision-language benchmarks (Section 4.1). Following this, in Section 4.2, we apply the domain adaptation framework introduced in Section 3.3 to transfer our models to three specialized domains: autonomous driving (Section 4.2.1 and Section 4.2.2), medical images (Section 4.2.3), and remote sensing (Section 4.2.4). Additionally, we perform an extensive ablation study to explore the impact of data sample size and model size on domain adaptation (Section 4.3).
In this section, we present a comprehensive evaluation of our model’s multimodal understanding and reasoning capabilities across a variety of benchmarks. The benchmarks used in our study are categorized into four distinct types: OCR-related tasks, including DocVQA , OCRBench and InfographicVQA ; chart and diagram understanding, including AI2D and ChartQA ; general multimodal tasks, such as MMBench ; and multimodal reasoning, including MMMU and MathVista . Additionally, we report the average score of our model on the OpenCompass .
As shown in Table 2, Mini-InternVL demonstrates strong performance across the majority of benchmarks. Our smallest model contains only 1 billion parameters, yet it demonstrates performance comparable to 2 billion parameter models, such as DeepSeek-VL-1.3B and MiniCPM-V 2.0. Compared to other lightweight models, our Mini-InternVL-4B excels across most benchmarks, particularly in MMbench, ChartQA, DocVQA, and MathVista, where its performance is on par with commercial models like Gemini-Pro-1.5. Notably, compared to InternVL2-Llama3-76B, which utilizes the larger InternViT-6B, Mini-InternVL achieves approximately 90% of its performance while using 5% parameters. This highlights the effectiveness of our knowledge distillation strategy.
2 Transfer to Various Specialized Domains
We select DriveLM-nuScenes version 1.1 as our training dataset, which contains 317K training samples and encompasses various aspects of the driving process. This dataset includes data for perception, prediction, and planning, offering a comprehensive understanding of autonomous driving scenarios.
In DriveLM-nuScenes, the images are captured from six different viewpoints. We effectively utilize dynamic resolution features to accommodate this type of data. Specifically, Mini-InternVL supports splitting images into 448448-sized tiles based on their aspect ratio. As illustrated in Figure 3, we resize the image of each view to 896448 pixels and then combine these images in a fixed sequence, resulting in a final resolution of 2688896. This means that the images are automatically processed into 12 tiles, and an additional thumbnail is added to provide the model with global context. Furthermore, we mark the image of each view with text indicating its camera position, such as “CAM_FRON”.
As shown in Figure 3, DriveLM-nuScenes contains QA pairs with coordinates, thus we need to normalize them to a range of 0 to 1000 to align with the output of Mini-InternVL. In the dataset, objects are represented by c tags. We use a tailored prompt: “Objects are encoded using
Finally, we conduct full-parameter fine-tuning of Mini-InternVL using 8 A100 GPUs, training the model for 1 epoch with a learning rate of 1e-5. We report the performance of our model after transfer learning on the CVPR 2024 Autonomous Driving Challenge . Furthermore, we evaluate our model on autonomous driving scenarios from MME-Realworld , where we separately assess its performance on Perception and Reasoning tasks.
We test our model using DriveLM-nuScenes-version-1.1-val , and the results are presented in Table 4. Our final score of Mini-InternVL-2B is 0.5958, which is comparable to the best result on the CVPR 2024 Autonomous Driving Challenge Leaderboardhttps://opendrivelab.com/challenge2024/#driving_with_language, InternVL4Drive-v2 . Notably, our model uses only one-tenth of the parameters of InternVL4Drive-v2.
Our model scores slightly lower in the Match metric, which might be due to Mini-InternVL’s lack of proficiency in predicting object center points. InternVL4Drive-v2 offers a viable solution by using Segment Anything to convert object center points into object bounding boxes. In Figure 3, we show the predicted answers of our model in perception, prediction, and planning tasks, demonstrating alignment with human driving behavior. Furthermore, our 4B-parameter model performs similarly to our 2B-parameter model. We attribute this potentially to the limitations of the existing training data and evaluation criteria, which might constrain larger models from achieving significant performance gains.
As shown in Table 5, in the autonomous driving scenarios of MME-Realworld, we observe that even when using only DriveLM as domain-specific training data, our model achieves an improvement of over 10 points. The transferred model surpasses the best-performing model on this task, LLaVA-OneVision-7B , as well as several commercial closed-source models such as GPT-4o and Claude 3.5 Sonnet , demonstrating the strong generalization capability of our model.
2.2 Autonomous Driving with Temporal Information
Using single-frame images alone is insufficient for accurate perception and prediction of vehicle behavior. Therefore, we explore temporal expansion. Specifically, we utilize instruction-following data constructed by DriveGPT4 from the BDD-X dataset as our training set, which comprises 26K video clips. Multiple video frames are organized as described in Section 3.3. Each training sample contains four aspects of question-answer data: Action Description, Action Justification, Speed Signal Prediction, and Turning Angle Signal Prediction. We set the proportion of general to domain-specific data at 1:1.
Following DriveGPT4 , we report several metric scores widely used in the NLP community, including CIDEr, BLEU4, and ROUGE-L, to evaluate the action descriptions and justifications. For open-loop control signal prediction, we use root mean squared error (RMSE) and threshold accuracies () for evaluation.
We report our scores on the BDD-X testing set in Table 6 and Table 7. Although our model has not undergone pre-training on large amounts of proprietary domain data like DriveGPT, it still performs comparably to DriveGPT4 across the four tasks and surpasses ADAPT . Note that the performance of the three models on this dataset is similar, which aligns with the observations discussed in Section 4.2.1.
2.3 Medical Image Question Answering
We utilize several publicly available medical image-text datasets to improve the model’s understanding of medical images. These datasets include a wide range of medical images, such as photos, X-rays, and pathology images. The datasets include PMC-OA , MedICaT , PMC-Image , Open-i , MedPix , Quilt-1M , RP3D , MIMIC-CXR , and Retina Image Bank , which together provide a substantial collection of medical images. From these datasets, we sampled 500K image-text pairs for the model’s training set. Finally, We add general data in a 1:1 ratio to the domain-specific data and conduct full-parameter training of the model for one epoch.
In this section, we present the performance of Mini-InternVL and its fine-tuned variant, Mini-InternVL-DA, on a comprehensive medical AI benchmark, GMAI-MMBench . Our evaluation is conducted using the VLMEvalKithttps://github.com/open-compass/VLMEvalKit framework. Table 9 illustrates the performance of various models on the medical VQA tasks.
After supervised fine-tuning, our model shows significant improvement across most evaluation metrics. Specifically, it excels in 2D classification (2D Cls), 2D detection (2D Det), and 2D multi-class accuracy (2D Mcls_acc). These results highlight its strong multimodal understanding capabilities in complex medical visual question-answering tasks.
Furthermore, our model of 4B size outperforms several medical-specialized models (e.g., LLaVA-Med , RadFM ) and some commercial closed-source models (e.g., Claude3-Opus ) on most metrics. However, there is no improvement in multiple-choice questions after SFT, which we attribute to the lack of multiple-choice question data in the training dataset.
2.4 Remote Sensing
The training data is summarized in Table 10. The GeoChat instruction set serves as the primary component of our training dataset. To enrich the dataset with high-resolution imagery, we also incorporate the RSVQA-HR dataset . Additionally, we include 100K VQA instances sampled from FIT-RS to further expand our training set. For the visual grounding task, we integrate the DIOR-RSVG dataset into our training process. All data are reformatted according to the methods outlined in Section 3.3.
A single epoch of training on the visual grounding data is found to be insufficient, so we repeat the DIOR-RSVG training samples multiple times. Finally, we incorporate 20% of the total general domain training samples into the training data and conduct training following the settings described in Section 4.2.1.
We assess the performance of our transferred model using the RSVQA dataset for the VQA task and the DIOR-RSVG dataset for the visual grounding task. Following the methodology outlined in , we chose the Presence, Comparison, and Rural/Urban subsets of the RSVQA-LR and RSVQA-HR datasets for assessment.
Table 11 presents the performance of our models on remote sensing VQA and visual grounding tasks. On the RSVQA task, our model demonstrates strong performance under both high-resolution and low-resolution conditions. Unlike existing remote sensing MLLMs such as GeoChat and SkySenseGPT , which support only single-resolution images, our model leverages dynamic resolution to effectively benefit from high-resolution training data. Compared to traditional models in the remote sensing domain–such as RSVQA , Bi-Modal , and EasyToHard –our model achieves superior scores on both RSVQA-HR-Test1 and RSVQA-HR-Test2, showcasing its generalization ability. Furthermore, our models of three different sizes outperform SkyEyeGPT on DIOR-RSVG, indicating that our framework can effectively model visual grounding tasks.
3 Ablation Study
We demonstrate the effectiveness of knowledge distillation in this study. We train a new 2B-sized MLLM, which we refer to as Mini-InternVL-CLIP-2B. Specifically, we maintain the structure of the MLLM while replacing the vision encoder from InternViT with CLIP-ViT-L-336px , employing the same training recipe as described in Section 3.1. As shown in Table 12, the results across multiple benchmarks indicate that Mini-InternVL-2B significantly outperforms Mini-InternVL-CLIP-2B on general benchmarks, particularly in document-related tasks. This clearly illustrates that our knowledge distillation in our method effectively enables InternViT to acquire visual world knowledge. Furthermore, we adapt both models to autonomous driving tasks, and the results demonstrate the advantages of our model in proprietary domain transfer.
In this study, we investigate the effect of the proportion between general data and domain-specific data on model transferability. We conducted experiments in the autonomous driving domain, utilizing all samples from DriveLM and supplementing them with general training data at multiples of times the number of DriveLM samples. The results are shown in Figure 4(a). Our findings indicate that relying exclusively on domain-specific training data does not yield optimal performance on downstream tasks. Introducing a specific ratio of general data not only enhances performance on domain-specific tasks but also reduces performance degradation on general multimodal benchmarks. This demonstrates that incorporating an appropriate proportion of general data is crucial for improving the model’s generalization ability and maintaining its general capabilities. For our autonomous driving scenario, we observe that performance peaks at ; beyond this point, performance slightly declines as increases. This suggests that we can achieve benefits without significantly increasing computational load.
We investigate how varying the quantity of training samples affects performance on downstream tasks. In this experiment, we use different amounts of training data while maintaining a 1:1 ratio between general data and domain-specific data, as shown in Figure 4(b). Training the model with only one-quarter of the full dataset significantly reduces the computational load during training while resulting in only a minor loss in performance. Notably, the model’s score on general benchmarks remains largely unchanged when the proportion of different training data is kept constant.
In this study, we examine the effects of different adaptation methods–LoRA, freezing the vision encoder, and full-parameter fine-tuning–on model performance. Using the dataset described in Section 4.2.4, we apply these three methods to train the model across varying numbers of steps and evaluate its performance on four benchmarks: general multimodal VQA, remote sensing VQA, and visual grounding. We record memory consumption and training speed during the process, as shown in Table 13. As illustrated in Figure 5, model performance converges after 1500 steps, with each method exhibiting distinct performance ceilings. Notably, full-parameter fine-tuning achieves the highest scores on domain-specific tasks. Additionally, we find that LoRA underperforms on the visual grounding task, and freezing the vision encoder strikes a balance between performance and computational efficiency. The scores for all three adaptation methods remain stable across different training steps on the general multimodal benchmark, maintaining strong performance in the general domain even with extended training.
Conclusion
In this work, we introduce Mini-InternVL, a series of lightweight, open-source MLLMs designed to tackle the challenges of deploying MLLMs in resource-constrained environments. Mini-InternVL utilizes InternViT-300M as a compact vision encoder, integrating world knowledge across multiple domains through knowledge distillation from a more capable teacher model, thereby addressing the limitations of encoders like CLIP-ViT. Mini-InternVL achieves approximately 90% of the performance of larger models using significantly fewer parameters, excelling particularly in tasks such as OCR and domain-specific image understanding. To facilitate the application of small-scale multimodal models in specialized fields, we employ a unified transfer format, enabling our models to be effectively adapted to multiple specific domains, where they achieve comparable performance to other domain-specific approaches. We hope that this work provides valuable insights into the application of MLLM.