MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, Qianyu Chen, Huarong Zhou, Zhensheng Zou, Haoye Zhang, Shengding Hu, Zhi Zheng, Jie Zhou, Jie Cai, Xu Han, Guoyang Zeng, Dahai Li, Zhiyuan Liu, Maosong Sun

Introduction

The rapid development of Multimodal Large Language Models (MLLMs) have brought an impressive surge in multimodal capabilities in understanding, reasoning and interaction. This has not only fundamentally reshaped the landscape of AI research and industry, but also shed light on a promising path towards the next AI milestone. However, current MLLMs are still far from being practical in real-world applications. One of the most predominant challenges is that current MLLMs typically entail a massive number of parameters and impose heavy computational burdens. As a result, most MLLMs can only be deployed on high-performing cloud servers, leading to significant energy consumption and carbon emissions. This limitation significantly constrains the potential application scopes such as on mobile devices, energy-sensitive scenarios, offline scenarios without stable network connections, and privacy/security protective scenarios for both personal and industrial users.

In light of these limitations, there is a growing interest in exploring more efficient lightweight MLLMs that can run on end-side devices. End-side scenarios encompass a broader scope of equipment, including mobile phones, personal computers, vehicles and robotics, etc., which are ubiquitous in users’ daily lives and experiencing rapid advancements in computation capacities. End-side MLLMs provide a promising solution towards more practical applications due to their broader usage scope, better computation efficiency, more robust offline behaviors, and better privacy/security protection.

However, developing capable end-side MLLMs is challenging due to significantly constrained parameter and inference computation budgets. As a result, more careful architecture designs and training recipes are required to fully unleash the potential of end-side MLLMs. In this work, we present MiniCPM-V, a series of efficient MLLMs deployable on end-side devices. The philosophy of MiniCPM-V is to achieve a good balance between performance and efficiency, a more important objective in real-world applications. To date in 2024, we have unveiled three models: (1) In February, we launched MiniCPM-V 1.0 2B, one of the first MLLMs designed for mobile phones. (2) In April, MiniCPM-V 2.0 2B was introduced, outperforming strong larger MLLMs such as Qwen-VL 9B , CogVLM 17B , and Yi-VL 34B . This iteration also introduces support for high-resolution image input and exhibits promising OCR capabilities. (3) Most recently in May, we released MiniCPM-Llama3-V 2.5 8B, which outperforms strong GPT-4V-1106, Gemini Pro and Claude 3 on OpenCompass evaluation. Noteworthy features of this model include strong OCR capability, high-resolution image perception, trustworthy behavior, multilingual support, and efficient end-side deployment optimization.

More importantly, MiniCPM-V can be viewed as a representative example of a promising trend. Fig. 1 summarizes the recent development of MLLMs in terms of performance, parameters and release time. We observe an interesting trend akin to Moore’s Law indicated by the red line: the sizes of models reaching the GPT-4V level performance are rapidly decreasing over time. This phenomenon could perhaps be called the Moore’s Law of MLLMs. Simultaneously, the computational capacity of end-side devices such as phones and personal computers is steadily increasing (qualitatively depicted by the blue line). The convergence of these two trends indicates usable (e.g., GPT-4V level) MLLMs deployable on end-side devices are soon within reach, opening up broader possibilities and benefiting more application scenarios in the near future. From a historical perspective of human technology development, this trend can also be viewed as human pursuit of miniaturization of state-of-the-art technologies, which have been repeatedly witnessed in other science and technology fields. For example, in aerospace, the latest SpaceX Raptor 2 rocket engine can achieve a strong thrust of 2,256 kN with a mass of 1.6 tons, whereas 20 years ago, the RD-0750 rocket engine could only achieve a thrust of 1,413 kN with a mass exceeding 4 tons .

In this paper, we will take the latest MiniCPM-Llama3-V 2.5 as an example, and systematically introduce the notable features of MiniCPM-V series and the key techniques behind them:

Leading Performance. MiniCPM-Llama3-V 2.5 achieves better performance than GPT-4V-1106, Gemini Pro and Claude 3 on OpenCompass collection, a comprehensive evaluation over 11 popular benchmarks. This is jointly contributed by its careful design in architecture, data and training recipes, which we will detail in the following.

Strong OCR Capability. MiniCPM-Llama3-V 2.5 outperforms GPT-4V, Gemini Pro and Qwen-VL-Max on OCRBench. It also supports high-utility functions such as table-to-markdown conversion and full OCR content transcribtion. These are largely attributed to the 1.8M pixel high-resolution (e.g., 1344 ×\times 1344) image perception technique across any aspect ratios .

Trustworthy Behavior. Based on the RLAIF-V and RLHF-V techniques that align MLLM behaviors from AI/human feedback, MiniCPM-Llama3-V 2.5 exhibits more trustworthy behaviors, achieving lower hallucination rates than GPT-4V-1106 on Object HalBench.

Multilingual Support. Inspired by the findings from VisCPM , the integration of multilingual LLM significantly alleviates the heavy reliance on multimodal training data in low-resource languages. Based on the foundation, a high-quality multilingual multimodal instruction tuning helps MiniCPM-Llama3-V 2.5 generalize its multimodal capabilities to more than 30 languages.

Efficient End-side Deployment. We systematically integrate a suite of end-side optimization techniques, encompassing quantization, memory optimization, compilation optimization and NPU acceleration, enabling efficient deployment on end-side devices.

We hope MiniCPM-V series can serve as an example for unveiling the potential of end-side MLLMs, and help draw more attention to promote the research in this direction. Following Moore’s Law for MLLM, we believe there will be increasingly powerful end-side MLLMs with reduced sizes, bringing efficient, safe, and trustworthy AI services on devices soon.

The contribution of this work is summarized as follows: (1) We introduce and open-source MiniCPM-V, a series of efficient end-side MLLMs achieving a good balance between performance and efficiency. (2) We investigate key techniques driving MLLMs towards the performance-efficiency balance at scale, unveiling the potential of these techniques. (3) We summarize the trend of MLLM development in its Moore’s Law, and empirically instantiate the trend with representative examples of MiniCPM-V.

Related Works

The development of LLMs has significantly advanced the progress in MLLMs. Flamingo first proposes to connect a pre-trained visual encoder with the Chinchilla 70B LLM and demonstrate the MLLM’s zero-shot and few-shot ability across a series visual language tasks. After the appearance of ChatGPT, many open-source models including BLIP-2 , Kosmos-1 , MiniGPT-4 , LLaVA , and VPGTrans are proposed. Among them, most are built upon existing pre-trained LLMs like Llama and Vicuna , while Kosmos-1 tries to train the LLM from scratch. Later, researchers continue to extend the function scope of MLLMs and improve the visual perception capabilities. Kosmos-2 , CogVLM , Shikra , and NExT-Chat further incorporate the localization capabilities to the MLLMs with either pix2seq paradigm or connecting with detection/segmentation models. Qwen-VL-Chat , Yi-VL , DeepSeek-VL , InternVL and Intern-XComposer pay more attention to improving the models’ capability with different techniques like high-resolution input, more training data, and better data ratio.

End-side Multimodal Large Language Models.

The huge number of parameters of MLLMs incurs prohibitively high computation costs in both training and deployment, greatly limiting the widespread applications. Recently, there has been a trend of building smaller LLMs with fewer parameters. The representative models are Phi , Gemma , MobileLLM , MiniCPM , etc. The moderate size of these models makes them applicable on end-side devices such as personal computers and even mobile phones. With optimized training strategies, end-side LLMs like MiniCPM 2B can achieve comparable performance with strong 7B models like Llama2-7B . Similar trends have also been witnessed in MLLMs. For example, Mini-Gemini and PaliGemma are built based on Gemma 2B and MobileVLM V2 is built based on MobileLlama . However, the fewer-parameter nature of end-side MLLMs presents significant challenges to building a capable model. MiniCPM-V series aims to push forward the potential of end-side MLLMs by addressing the key bottleneck problems through careful designs in architecture, training, inference and deployment.

Model Architecture

In this section, we present the model architecture of MiniCPM-V, outlining the overall structure and the adaptive high-resolution visual encoding approach. The design philosophy of MiniCPM-V series is to achieve a good balance between performance and efficiency, a more practical objective for a broader scope of real-world applications, which is implemented in architecture design, training, inference, and deployment.

The model comprises three key modules: the visual encoder, compression layer, and LLM. The input image is first encoded by a visual encoder, utilizing the adaptive visual encoding approach. Specifically, we employ SigLIP SoViT-400m/14 as the visual encoder. The visual tokens are then compressed by the compression layer, which adopts a perceiver resampler structure with one layer cross-attention. Finally, the compressed visual tokens, along with the text input, are fed into the LLM for conditional text generation.

2 Adaptive Visual Encoding

Recently, there has been growing consensus on the fundamental role of visual encoding in MLLM performance , especially for fine-grained capabilities such as OCR. For effectiveness, a good visual encoding strategy should both respect the raw aspect ratio of the input and preserve sufficient visual details (high resolution). For efficiency, the number of visual tokens from image encoding should be moderate to be affordable on end-side devices. To this end, we take advantage of the adaptive visual encoding method proposed by LLaVA-UHD .

We select the partition with the highest score from all possible candidates:

Slice Encoding.

Token Compression.

After visual encoding, each slice is encoded into 1,024 tokens, where 10 slices can yield over 10k tokens collectively. To manage this high token count, we employ a compression module comprising of one-layer cross-attention and a moderate number of queries, with 2D positions informed . In practice, the visual tokens of each slice are compressed into 64 queries for MiniCPM V1&2 and 96 tokens for MiniCPM-Llama3-V 2.5 through this layer. Compared with other MLLMs with competitive performance, the significantly smaller number of visual tokens in MiniCPM-V series enables superior efficiency in terms of GPU memory consumption, inference speed, first-token latency and power consumption, making it more friendly to wider application scopes and communities.

Spatial Schema.

To indicate each slice’s position relative to the whole image, inspired by , we additionally introduce a spatial schema. We first wrap tokens of each slice by two special tokens and <\slice>, and then employ a special token “\n” to separate slices from different rows.

Training

The model training consists of 3 phases: the pre-training phase, the supervised fine-tuning phase, and the RLAIF-V phase. We will introduce the training recipe in the following sections.

In this phase, we utilize large-scale image-text pairs for MLLM pre-training. The primary goal of this phase is to align the visual modules (i.e., visual encoder and compression layer) with the input space of the LLM and learn foundational multimodal knowledge. The pre-training phase is further divided into 3 stages.

The role of stage-1 is to warm up the compression layer, primarily connecting the visual encoder and LLMs. (1) Trainable Modules. We randomly initialize the compression layer and train this module in stage-1, keeping other parameters frozen. The visual encoder’s resolution is set to 224×\times224, which is the same as the visual encoder’s pre-training setting. (2) Data. To warm up the compression layer, we randomly select 200M data from the Image Captioning data in Table 1. Data cleaning is performed to remove image-text pairs with poor correlation and ill-formatted text data, ensuring the data quality.

Stage-2.

After the warm-up training of the compression layer, the role of stage-2 is to extend the input resolution of the pre-trained visual encoder. (1) Trainable Modules. In stage-2, we extend the image resolution from 224×\times224 to 448×\times448. The whole visual encoder is trained, leaving other parameters frozen. (2) Data. To extend the pre-trained resolution, we additionally select 200M data from the Image Captioning data in Table 1.

Stage-3.

After extending the primary input resolution of the visual encoder, we finally train the visual modules using the adaptive visual encoding strategy, which can further accommodate high-resolution inputs with any aspect ratio. (1) Trainable Modules. During the stage-3 training, both the compression layer and the visual encoder are trained to adapt to the language model embedding space. The LLM is kept frozen to avoid disruption from the relatively low-quality pre-training data. (2) Data. Different from the previous stages with only image captioning data, during the high-resolution pre-training stage, we additionally introduce OCR data to enhance the visual encoders’ OCR capability.

Caption Rewriting.

Image-text pairs sourced from the Web can suffer from quality issues in the caption data, including non-fluent content, grammatical errors, and duplicated words. Such low-quality data can lead to unstable training dynamics. To address the issue, we introduce an auxiliary model for low-quality caption rewriting. The rewriting model takes the raw caption as input and is asked to convert it into a question-answer pair. The answer from this process is adopted as the updated caption. In practice, we leverage GPT-4 to annotate a small number of seed samples, which are then used to fine-tune an LLM for the rewriting task.

Data Packing.

Samples from different data sources usually have different lengths. The high variance of sample lengths across batches will lead to inefficiency in memory usage and the risk of out-of-memory (OOM) errors. To address the issue, we pack multiple samples into a single sequence with a fixed length. By truncating the last sample in the sequence, we ensure uniformity in sequence lengths, facilitating more consistent memory consumption and computational efficiency. Meanwhile, we modify the position ids and attention masks to avoid interference between different samples. In our experiments, the data packing strategy can bring 2~3 times acceleration in the pre-training phase.

Multilingual Generalization.

Multimodal capability across multiple languages is essential for serving users from broader communities. Traditional solutions involve extensive multimodal data collection and cleaning, and training for the target languages. Fortunately, recent findings from VisCPM have shown that the multimodal capabilities can be efficiently generalized across languages via a strong multilingual LLM pivot. This solution largely alleviates the heavy reliance on multimodal data in low-resource languages. In practice, we only pre-train our model on English and Chinese multimodal data, and then perform a lightweight but high-quality multilingual supervised fine-tuning to align to the target languages. Despite its simplicity, we find the resultant MiniCPM-Llama3-V 2.5 can achieve good performance in over 30 languages as compared with significantly larger MLLMs.

2 Supervised Fine-tuning

After learning foundational capabilities from pre-training, we perform supervised fine-tuning (SFT) on high-quality visual question answering datasets to further learn knowledge and interaction capability from human annotations.

Compared with the pre-training phase which mainly uses crawled data from the Web, the SFT phase mainly utilizes high-quality datasets annotated by either human lablers or strong models such as GPT-4. Therefore, we unlock all model parameters to better exploit the data and learn rich knowledge during SFT phase.

Data.

Recent works show that data near the end of training plays a more important role in shaping the models’ capabilities and response styles. We categorize the SFT data into two parts. Part-1 focuses on bolstering the models’ basic recognition capabilities, while part-2 is tailored to enhance their capabilities in generating detailed responses and following human instructions. Specifically, part-1 data consists of the traditional QA/captioning datasets with relatively short response lengths, which helps enhance the model’s basic recognition capabilities. In comparison, part-2 encompasses datasets featuring long responses with complex interactions, either in text or multimodal context. During SFT, these two parts of data are concatenated and sequentially fed into the model. For MiniCPM-Llama3-V 2.5, we integrate 2M data from the recent Cauldron dataset for multimodal knowledge augmentation, and 90K multilingual data over 36 languages for boosting the multilingual conversation capability.

3 RLAIF-V

MLLMs are typically prone to hallucination problems, generating responses that are not factually grounded in the input image . The issue greatly limits the wide application of MLLMs, especially in high-stakes scenarios, such as autonomous driving and assistance for visually impaired groups. To address the hallucination problem, we employ the recent RLAIF-V approach (Fig. 4), where the key is to obtain scalable high-quality feedback from open-source models for direct preference optimization (DPO) .

The first step of RLAIF-V is to generate multiple responses for a given instruction using the policy model. Specifically, given a model MM waiting for alignment, we sample 10 responses Y={y1,y2,⋯ ,yn}Y=\{y_{1},y_{2},\cdots,y_{n}\} from MM using sampling decoding with high temperatures. There are several benefits of using the policy model MM for response generation: (1) Feedback collection and learning can better focus on trustworthiness, since different text styles from multiple MLLMs are avoided. (2) Feedback learning is more efficient since preference is directly collected on the distribution of the policy model.

Feedback Collection.

Collecting high-quality feedback from open-source MLLMs can be challenging due to their typically weaker capabilities compared with proprietary models. To address the issue, RLAIF-V uses a divide-and-conquer strategy for response scoring. Specifically, each response yiy_{i} is divided into atomic claims Ci={c1,c2,⋯ ,cm}C_{i}=\{c_{1},c_{2},\cdots,c_{m}\} using Llama-3 8B, where the correctness of atomic claims is much easier to evaluate. Then, we verify the claims by converting each claim to a yes/no question and employing an open-source MLLM to score each claim. In practice, we adopt OmniLMM 12B for MiniCPM-V 2.0 scoring and LLaVA-NeXT-Yi 34B for MiniCPM-Llama3-V 2.5 scoring. The final score sis_{i} of the response yiy_{i} is given by −nrej-n_{rej}, where nrejn_{rej} is the number of invalid atomic claims.

Direct Preference Optimization.

After collecting the high-quality AI feedback, we perform preference learning via DPO method. The DPO algorithm requires training on preference pairs, where one sample ywy_{w} is preferred to the other one yly_{l}. To compose the preference dataset, we randomly sample pairs from each response set Y={y1,y2,⋯ ,yn}Y=\{y_{1},y_{2},\cdots,y_{n}\}, and determine (yw,yl)(y_{w},y_{l}) based on their relative scores. Finally, we construct a preference dataset consisting of 6K preference pairs from 3K unique images for preference learning.

End-side Deployment

In this section, we investigate the deployment of MiniCPM-V on end-side devices. We first introduce the challenges, and then present the basic and advanced practices for end-side deployment. Finally, we analyze and discuss the evaluation results across different devices.

End-side devices, such as smartphones and computers, often face resource limitations due to factors like heat dissipation, size constraints, and power consumption. We identify several key challenges of end-side deployment for MLLMs by comparing end-side devices with high-performance servers:

High-performance servers typically boast extensive memory capacities, often exceeding 100GB or even 1TB. In contrast, the memory available on mobile phones typically ranges from 12GB to 16GB, which can be insufficient for MLLM deployment.

CPU/GPU Speed Restriction.

The overall processing speeds of CPUs in smartphones are notably slower. For instance, the Snapdragon 8 Gen3 features 8 CPU cores https://docs.qualcomm.com/bundle/publicresource/87-71408-1_REV_E_Snapdragon_8_gen_3_Mobile_Platform_Product_Brief.pdf, whereas high-performance server like Intel Xeon Platinum 8580 has 60 CPU cores https://www.intel.com/content/www/us/en/products/sku/237250/intel-xeon-platinum-8580-processor-300m-cache-2-00-ghz/specifications.html. Similarly, mobile phone GPUs are not as powerful as server GPUs. For example, Qualcomm Adreno 750 only has 6 TFLOPS, while NVIDIA 4090 can reach 83 TFLOPS.

2 Basic Practice

To deploy the MLLM on end-side devices, we first employ quantization for reduced memory cost, and empirically investigate the deployment results on different frameworks.

Quantization is a widely used technique to reduce memory consumption. The main idea of model quantization is to use a unified scaling factor to compress multiple weights into a narrower range, followed by discretization. This process is mathematically represented as:

where w′w^{\prime} denotes the quantized parameter and ss signifies the calculated scale factor. The round function discretizes the quantized value.

For MiniCPM-Llama3-V 2.5, the fp16 version model typically demands 16~17G memory. We opt for the Q4_K_M mode 4-bit quantization strategy within GGMLhttps://github.com/ggerganov/ggml framework. This reduces the memory requirement to around 5G, which is friendly to mobile phone usage.

Deployment Framework.

Several frameworks have been proposed for end-side deployment. Illustrated in Fig. 5, we make a thorough investigation of different frameworks for different chip types including CPU, GPU, and NPU.

Given the ubiquity of CPU usage across devices, we prioritize this chip type and opt for the llama.cpp framework. Combining quantization and llama.cpp on Xiaomi 14 Pro (Snapdragon 8 Gen 3), the model achieves a text encoding latency of 64.2s and a text decoding speed of 1.3 tokens/s (as depicted in Fig. 6), which is still far from acceptable for users.

3 Advanced Practice

To enhance user experience, we investigate a series of advanced techniques including memory usage optimization, compilation optimization, configuration optimization, and NPU acceleration.

Experimental results show that, without specific optimizations, image processing can be the bottleneck of the inference speed due to limited memory resources on mobile phones. To address the issue, we explore memory usage optimization strategies. Instead of loading both ViT and LLM simultaneously into memory, we adopt a sequential loading approach. Specifically, we first load ViT for visual encoding, followed by the LLM for visual and text token encoding. By releasing the large amount of memory occupied by LLM, we can prevent frequent paging (swapping in and out) during ViT encoding, thereby improving the program efficiency. This optimization technique, as illustrated in Fig. 6 (a), results in a notable reduction of image processing time from 45.2s to 31.5s.

Compilation Optimization.

We find that directly compiling the models on the target devices can significantly improve the encoding latency and the decoding throughput. This can be attributed to better consistency between the compilation and target device instruction set architecture. As depicted in Fig. 6, this optimization endeavor yields promising results. Encoding latency shows a notable reduction from 50.5s to 17.0s, while decoding throughput experiences a significant boost from 1.3 tokens/s to 3.2 tokens/s.

Configuration Optimization.

We observe that a single default configuration of the llama.cpp framework may not be optimal for diverse end-side devices. To maximize the inference speed, we devise an automatic parameter search algorithm that dynamically determines the most suitable configurations (e.g., computation allocation on different CPU cores). Through configuration optimization, we can achieve good improvements. Specifically, decoding throughput surged from 3.2 tokens/s to an impressive 8.2 tokens/s, surpassing the typical human reading speed.

NPU Acceleration.

The above techniques are mostly tailored for CPU deployment. Another promising avenue involves leveraging alternative chip types such as GPUs and NPUs. Despite the potential of GPU, we find in our experiments that current frameworks for mobile phone GPU are not optimized or compatible enough to exceed the results on CPU. As an alternative, we turn to NPUs (Neural Processing Units), which represent a novel class of specialized hardware introduced in recent years, specifically designed for accelerating AI applications. Some smartphones are already equipped with NPUs, which are recognized as better suited for addressing computation bottlenecks.

In practice, we primarily leverage NPUs to accelerate visual encoding. Specifically, we replace the backend framework of ViT to QNN, while retaining the llama.cpp backend for the LLM component. For mobile phones equipped with Qualcomm NPUs, this optimization results in a notable reduction in visual encoding time, decreasing from 3.7s to 1.3s, as illustrated in Fig. 6 (a).

4 Results.

For a comprehensive assessment of MiniCPM-Llama3-V 2.5’s performance across various end-side devices, we present test results on Xiaomi 14 Pro (Snapdragon 8 Gen 3), vivo X00 Pro (Mediatek Dimensity 9300), and Macbook Pro (M1) in Fig. 7. Thanks to the deployment optimization techniques, MiniCPM-Llama3-V 2.5 can operate efficiently on both mobile phones and personal computers, delivering acceptable latency and throughput. For instance, leveraging NPU on Xiaomi 14 Pro enables it to achieve a similar encoding speed as the Mac M1. Furthermore, nearly all devices exhibit comparable or higher throughput compared with human reading speed.

Discussion.

Upon analyzing Fig. 7, it becomes evident that the current computation bottleneck primarily stems from LLM prefilling, which mainly involves encoding image and text tokens for LLM inference. Promising research directions involve developing more efficient visual encoding methods with fewer visual tokens, and better leveraging GPU/NPU acceleration for LLM encoding. With increasing attention to end-side MLLMs and the rapid advancement of GPU/NPU acceleration techniques, we believe that real-time interaction with end-side MLLMs can be reached soon.

Experiments

In this section, we perform a comprehensive evaluation of MiniCPM-V series.

We have released 3 models in the MiniCPM-V series, including MiniCPM-V 1.0, MiniCPM-V 2.0, and MiniCPM-Llama3-V 2.5. As shown in Table 3, MiniCPM-V 1.0 is trained with the pre-training stage1&2 and SFT without using the adaptive visual encoding and RLAIF-V. For MiniCPM-V 2.0, we include all of the training stages and the adaptive visual encoding strategy to further improve performance. In MiniCPM-Llama3-V 2.5, Llama3-Instruct 8B is adopted as the base LLM.

2 Experiment Settings

We perform a comprehensive evaluation on popular benchmarks covering visual question answering, multimodal conversation, knowledge and reasoning, OCR, and hallucination. (1) General benchmarks. We adopt OpenCompass as the general evaluation indicator, which is a comprehensive collection over 11 popular multimodal benchmarks, including MME , MMBench , MMMU , MathVista , LLaVA Bench , etc. We also report the results on RealWorldQA for real-world spatial understanding capabilities. (2) OCR benchmarks. We adopt three widely used benchmarks for OCR capability evaluation, including including OCRBench , TextVQA and DocVQA . (3) Hallucination benchmarks. We also include Object HalBench to evaluate the trustworthiness of the models.

Baselines.

We compare with strong baselines in different series: For open-source models, we compare with strong models including Yi-VL-6B/34B , Qwen-VL-Chat , DeepSeek-VL-7B , TextMonkey , CogVLM-Chat-17B , CogVLM2-Llama3-19B , Idefics2-8B , Bunny-Llama-3-8B , XTuner-Llama-3-8B-v1.1 , LLaVA-NeXT-Llama-3-8B , Cambrian-8B/34B , LLaVA-NeXT-Yi-34B , DeepSeek-VL-1.3B , MobileVLM V2 , Mini-Gemini and Phi-3-Vision-128k-instruct . For proprietary models, we compare with GPT-4V-1106 , Gemini-Pro and Claude 3 Opus .

3 Experimental Results

From the experimental results in Table 4, we have the following observations: (1) MiniCPM-Llama3-V 2.5 outperforms strong open-source models by a notable margin. For instance, MiniCPM-Llama3-V 2.5 surpasses the recent strong Idefics2-8B by 7.9 points on the OpenCompass benchmark, with similar model sizes. It also achieves better results than significantly larger models such as Cambrian-34B, LLaVA-NeXT-Yi-34B, Yi-VL-34B and CogVLM2-Llama3-19B. (2) Compared with powerful proprietary models, such as GPT-4V-1106 and Gemini Pro, MiniCPM-Llama3-V 2.5 achieves better performance on the OpenCompass benchmark with significantly fewer parameters. In addition, MiniCPM-Llama3-V 2.5 also achieves lower hallucination rates than GPT-4V-1106 on Object HalBench, indicating its trustworthiness for real-world applications. (3) The smaller MiniCPM-V 2.0 with 2B parameters achieves significantly better performance compared with other 2B~3B models, and is even comparable with Llama3-based 8B MLLMs such as Bunny-Llama-3-8B. In summary, the results show that MiniCPM-V series achieves a good balance between performance and efficiency, making it more friendly for broader communities and applications.

Results on OCR Benchmarks.

MiniCPM-V models also show strong OCR capabilities, including scene-text, document and screenshot understanding. As shown in Table 5, MiniCPM-Llama3-V 2.5 outperforms open-source MLLMs ranging 1.7B~34B on OCRBench, TextVQA, and DocVQA, and even performs comparably to proprietary models such as GPT-4V-1106 and Gemini Pro.

Multilingual Multimodal Capability.

Based on the multilingual multimodal generalization approach from VisCPM, MiniCPM-Llama3-V 2.5 extends its multimodal capability to over 30 languages. As shown in Fig. 8, MiniCPM-Llama3-V 2.5 can outperform Yi-VL 34B and Phi-3-vision-128k-instruct on the multilingual LLaVA Bench. The promising multilingual multimodal capability makes MiniCPM-Llama3-V 2.5 useful in serving larger groups with various languages.

Comparison with Other Llama-3 based Models.

From experimental results in Table 4, we can observe that: (1) MiniCPM-Llama3-V 2.5 outperforms other Llama-3 based models by a large margin. For example, compared with the strong LLaVA-NeXT-Llama-3-8B, MiniCPM-Llama3-V 2.5 consistently achieves better results on all benchmarks. (2) Moreover, it is worth noting that MiniCPM-Llama3-V 2.5 requires significantly less inference computation. For example, the visual token number range of MiniCPM-Llama3-V 2.5 is (96, 960), which is lower than LLaVA-NeXT-Llama-3-8B’s (1728, 2880). This can be important especially for real-world end-side applications in terms of inference speed, first-token latency, memory usage, and power consumption.

4 Ablation Study

We perform an ablation study on components of MiniCPM-Llama3-V 2.5, including RLAIF-V and multilingual training.

From the results in Table 6, we can observe that RLAIF-V effectively reduces the hallucination rates of the base model on both response level and mention level. This makes the model more trustworthy in behaviors. Importantly, the hallucination reduction does not sacrifice the general capabilities. In contrast, RLAIF-V further improves the overall performance on OpenCompass by 0.6 points on an average of 11 benchmarks.

Multilingual Generalization.

We investigate the necessity and effectiveness of the multilingual generalization technique. As shown in Table 8, we can see over 25 point improvement in all languages when using less than 0.5% multilingual SFT data. The results show that the multilingual generalization method can effectively improve multilingual capability with good data and computation efficiency. In addition, we notice that the performance improvement is uneven across languages. We hypothesize that the improvement extent might be related to multiple factors like the base LLM’s ability of the given language. We leave more systematical exploration for future works.

5 Case Study

We provide a more intuitive understanding of MiniCPM-Llama3-V 2.5 capabilities in the case study.

MiniCPM-Llama3-V 2.5 shows strong OCR capabilities in real-world scenarios. Illustrated in Fig. 2, the model accurately transcribes English articles from screenshots into plain text, converts tables containing both English and Chinese content into Markdown format, comprehends code logic, and provides reasonable plans based on image content.

Any Aspect-ratio High-resolution Input.

A standout feature of the MiniCPM-Llama3-V 2.5 is its capability to handle high-resolution input with extreme aspect ratios. As depicted in Fig. 9, the model well processes input with an aspect ratio of 10:1, accurately recognizing fine-grained article contents. Interestingly, the model can also interpret images within images, correctly describing the central image as a man “smiling and holding a chocolate mousse."

Multilingual Multimodal Capability.

Benefiting from the multilingual multimodal generalization approach from VisCPM , MiniCPM-Llama3-V 2.5 exhibits multilingual proficiency, generalizing across more than 30 languages. Fig. 11 showcases multimodal conversations in German, French, Japanese, Korean, and Spanish, showing good knowledge of language-specific cultures.

Trustworthy Behavior.

Based on RLAIF-V, MiniCPM-Llama3-V 2.5 ensures more trustworthy responses with lower hallucination rates. As demonstrated in Fig. 10, the model’s responses exhibit fewer hallucinations as compared with powerful GPT-4V, showing its promising reliability and trustworthiness in real-world scenarios.

Conclusion

In this work, we introduce the MiniCPM-V series models as a primary exploration into powerful end-side MLLMs. Thanks to techniques such as adaptive visual encoding, multilingual generalization, and the RLAIF-V method, MiniCPM-Llama3-V 2.5 can achieve GPT-4V level performance with significantly fewer parameters. With various end-side optimization techniques, this model ensures an acceptable user experience on mobile phones.

Limitations.

Despite promising performance, there remain several limitations with current MiniCPM-V models. (1) Capability Depth. there is still plenty of room for improvement in enhancing multimodal understanding capability and inference efficiency. (2) Capability Width. In addition to image modality, it’s promising to expand MLLM capabilities to encompass other modalities, such as video and audio, etc., where GPT-4o and Google Astra have given good examples.

In addition to MLLM capabilities, end-side deployment also presents unique challenges. The inference speed and latency are still far from good enough and the model service can be limited by the battery capacity. In addition, previous efforts on chips and deployment frameworks mainly target CNNs and LSTMs, which can be sub-optimal for MLLMs. Tailored efforts to MLLMs can bring ample room for improvement.

Future Works.

Considering the current limitations and the promising future of end-side MLLMs, we anticipate increasing efforts from both academia and industry in enhancing model capabilities in terms of depth and width, and improving smartphone chips and deployment frameworks. We believe that simultaneous advancements in model capability and end-side device capacity will lead to end-side applications providing a satisfying user experience in the near future.

References