Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Shuhao Gu, Jialing Zhang, Siyuan Zhou, Kevin Yu, Zhaohu Xing, Liangdong Wang, Zhou Cao, Jintao Jia, Zhuoyi Zhang, Yixuan Wang, Zhenchong Hu, Bo-Wen Zhang, Jijie Li, Dong Liang, Yingli Zhao, Songjing Wang, Yulong Ao, Yiming Ju, Huanhuan Ma, Xiaotong Li, Haiwen Diao, Yufeng Cui, Xinlong Wang, Yaoqi Liu, Fangxiang Feng, Guang Liu

Introduction

Recently, Vision-Language Models (VLMs) Li et al. (2023); Liu et al. (2024b); Dai et al. (2023); Zhu et al. (2024); Bai et al. (2023b); Wang et al. (2023b); Xiao et al. (2024); OpenAI (2024); Yao et al. (2024); Wang et al. (2024a); Chen et al. (2024b); Li et al. (2024a) have made significant progresses, drawing increasing attention. With the ongoing advancements in foundational language models, multimodal architectures, multimodal training data, and evaluation benchmarks, the capabilities of multimodal models have greatly improved. Among these developments, the expansion of training data scale, the enhancement of data quality, and the optimization of training strategies have emerged as key factors in boosting model performance Liu et al. (2024b, 2023a); Tong et al. (2024); Li et al. (2024a, b). Currently, the two primary methods for data acquisition are manual data collection and annotation, as well as using models to synthesize instructions.

Many works have focused on exploring more effective ways to generate and utilize training data. For instance, Liu et al. (2023a) leverages GPT-4 to generate various types of instructions, including dialogues, detailed descriptions, and complex reasoning, based on textual descriptions of images. Building on this, Li et al. (2024b) further expands the data scale, leading to performance improvements. Tong et al. (2024) enhances model performance by increasing the dataset size and adjusting the data type ratios, while Li et al. (2024a) introduces a "high-quality knowledge learning" phase to further enrich the model’s knowledge base. In addition, several works explore using closed-source commercial models to generate synthetic instruction data, such as generating captions with GPT-4o or GPT-4v models Chen et al. (2023, 2024a) or OCR data Carter (2024), as well as conversation data Wang et al. (2023a). Despite these advancements, existing open-source data and instruction datasets remain insufficient to support models in achieving optimal performance. Models trained solely on open-source data still significantly lag behind SOTA closed-source models or open-source models trained on proprietary data. The limitations in both the quantity and quality of open-source data are key factors constraining model performance.

To further enhance the performance of open-source models, this work explores improving model effectiveness by expanding the scale of instruction data and increasing the diversity of instruction types. We have extensively collected existing open-source multimodal instruction data, constructing a dataset of approximately 40 million samples, and applied rigorous quality filtering and deduplication processes. The model trained on this dataset demonstrated excellent performance, achieving a very high level of accuracy. Building on this, we propose a multimodal instruction synthesis method based on open-source VLM models. By providing highly detailed annotations for images and generating diverse questions for each image to ensure comprehensive coverage of the information, we can produce higher-quality instruction data, further improving the model’s ability to understand and follow instructions. Ultimately, we successfully trained a 2-billion-parameter VLM model based on open-source data and synthetic data generated by open-source models, achieving SOTA performance comparable to SOTA open-source models of similar scale.

The key contributions of this research include:

We collected, organized, and open-sourced a large-scale multimodal instruction dataset, Infinity-MM, consisting of tens of millions of samples. Through quality filtering and deduplication, we ensured the dataset’s high quality and diversity.

We proposed a synthetic data generation method based on open-source models and a labeling system, capable of producing high-quality instruction data and effectively expanding the scale of instruction datasets.

Based on Infinity-MM, we successfully trained a 2-billion-parameter VLM model, Aquila-VL-2B, achieving state-of-the-art performance among models of the same scale.

Related Work

Vision-Language Model VLMs can be categorized into three types based on their capabilities. The first type focuses on understanding multimodal information, such as videos and images Radford et al. (2021); Alayrac et al. (2022); Liu et al. (2024b); Li et al. (2023); Diao et al. (2024). These models typically take multimodal data as input and produce natural language output, characterized by their ability to integrate and process information from different modalities in a unified manner. The second type emphasizes visual generation, primarily aimed at producing high-resolution images and videos Shi et al. (2020); Peebles and Xie (2023); Ramesh et al. (2021); Ding et al. (2021). The third type combines both visual understanding and generation capabilities Sun et al. (2024b, a); Wang et al. (2024b); Zhou et al. (2024); Xie et al. (2024). In this work, we focus on enhancing the model’s ability to comprehend multimodal information.

Multi-modal Instruction Data Currently, a considerable amount of research has explored leveraging closed-source commercial models (mainly the GPT-4 series) to generate synthetic instruction data. The first category of work primarily utilizes GPT-4o or GPT-4v to generate specific types of data, such as captions Chen et al. (2023, 2024a), OCR Carter (2024) and conversations Wang et al. (2023a). Another category of work attempts to generate more complex dialogue or other types of instructions. For example, Liu et al. (2023a) uses GPT-4 to generate various types of instructions based on textual descriptions of images. Wang et al. (2023a) directly uses GPT-4V to generate instructions from images, though it generates only one instruction per image except for the text description for the image. In this work, we focus on how to leverage open-source models to generate high-quality multimodal instruction data.

Data

First, we extensively collect existing open-source multimodal datasets and categorize them based on task and quality. Subsequently, we will introduce the process of synthesizing data. Finally, we performed a unified deduplication and filtering of all collected data. The next three subsections will each address specific aspects in detail.

We systematically gathered available open-source multimodal datasets and categorized them. These datasets were classified into four categories, as outlined in Table 1.

Image-Caption Data We collected the Image-Caption dataset generated by Emu2 (Sun et al., 2024a). Caption generation is a relatively fundamental task, making it well-suited for the initial training of large multimodal models.

General Visual Instruction Data We collected various general task data encompassing OCR, mathematical reasoning, chart comprehension, and other tasks. Training large multimodal models with this data equips them with the fundamental capabilities to tackle multimodal tasks effectively.

Selective Visual Instruction Data The sources of this data are Llava-OneVision Li et al. (2024a), Docmatix Laurençon et al. (2024) and the subjective components of Infinity-Instruct BAAI (2024b). Verifications on these data have shown that the quality of this data is superior to that of general task instruction data.

GPT4 & Synthetic Data This part mainly includes data generated by GPT-4 and the synthetic instruction data introduced in Section 3.2, along with a small amount of data specifically tailored for targeted tasks. Experimental results indicate that training with datasets containing synthetic data can further enhance model performance.

2 Synthetic Data Generation

In this study, we propose a multimodal instruction data synthesis method based on open-sourced VLMs. Our goal is to ensure that the generated instructions are closely aligned with the content of the images, while maintaining diversity in instruction types and ensuring the accuracy of instruction responses. The overall process of the method is shown in Figure 1. The images of the synthetic data are extracted from the instruction dataset synthesized using the GPT-4 series models, which is of high quality. However, due to budget constraints, the scope and quantity of the synthetic data are limited. Therefore, we aim to leverage open-source models to synthesize more high-quality data, combining it with the original data to further enhance model performance.

We first utilize the RAM++ model Huang et al. (2023) to automatically annotate images by extracting key information such as objects, actions, and scenes. These tags form the semantic foundation of the images, providing a critical basis for subsequent instruction generation. The RAM++ model demonstrates excellent performance when processing large-scale image datasets, accurately capturing essential details in multimodal scenes. This lays a solid foundation for generating precise and contextually relevant multimodal instructions.

To systematize the instruction generation process, we designed a three-level instruction tagging system that covers different types of instructions. Following Liu et al. (2023b), the first-level tags of the instruction tagging system are divided into six categories, which are:

Fine-grained Perception (single-instance)

The middle level further refines task characteristics, while the bottom level provides a detailed classification based on specific task requirements. We employed a commercial closed-source model to extend and enhance this system, ensuring its comprehensiveness and rationality. The complete tagging system can be found in Appendix D.

2.2 Question Generation

We randomly selected a portion of the open-source data we collected as seed data and annotated both the images and instructions using the method described in the previous section. We then established a set of mapping rules by analyzing the correlations between image tags and instruction tags. Specifically, we calculated the TF-IDF values for the image tags corresponding to each instruction type and ranked the results. Higher TF-IDF values indicate that images with those tags are more suitable for generating that particular type of instruction. Using these rules, we can automatically determine the appropriate instruction type to generate when processing new images. This approach significantly enhances the alignment between generated instructions and image content.

During question generation, we input both the images and the target instruction type into the VLM model, prompting the model to generate questions based on the image. Additionally, we randomly select two examples from the seed data to input into the model alongside the image for reference, enabling few-shot generation. For the questions generated by the VLM, we further input both the image and the question back into the question VLM to evaluate the relevance of the question to the image, filtering out lower-quality questions.

2.3 Answer Generation

After generating the questions, we proceeded to generate the corresponding instruction answers. The goal at this stage was not only to ensure the accuracy of the generated answers but also to account for the diversity of different instruction types. To achieve this, we introduced various prompts to increase answer diversity. Specifically, we employed three different types of prompts: one instructed the model to provide short answers using single words or phrases; another prompted the model to first generate a simple explanation before giving the answer; and the third prompted the model to provide a detailed explanation followed by the answer. We then input the image, question, and generated answer into the VLM model to filter out instructions and answers that did not align with the image content or task.

Finally, we obtained approximately 10M question-answer pairs. To further ensure the quality of the generated data, we input the images, questions, and answers into the Qwen2-VL-2B model to compute the data loss and filtered out about 3M samples. We combined multiple QA pairs corresponding to the same image into multi-turn instruction data, resulting in approximately 800K training instructions. The distribution of instruction types in the final synthetic data is shown in Figure 2.

3 Data Processing

After collecting all the data, we proceeded with data processing. First, to facilitate large-scale training, we standardized the format of data from various sources. Then, to improve training efficiency and enhance model performance, we conducted a series of data-cleaning steps. Specifically, we removed duplicate Image-Text pairs and filtered out images with high similarity based on their pHash values. Additionally, we used Qwen2-VL-2B to calculate the loss for each sample and excluded the top 5% with the highest loss, as high loss in well-trained multimodal models often indicates noisy data or outliers. To prevent data contamination, we also deduplicated the training set against images found in the test set. We have included more detailed data information in Appendix C.

Architecture and Training

Aquila-VL builds upon LLaVA-OneVision architecture(Li et al., 2024a), comprising a language tower, a vision tower, and a projector.

Language Tower We chose Qwen-2.5 (Bai et al., 2023a) as the language tower for its outstanding performance among open-source models and its availability in various sizes.

Vision Tower We utilized SigLIP (Zhai et al., 2023), with approximately 400 million parameters, as the vision tower to extract visual features from input images and videos.

Projector We utilized a two-layer MLP (Liu et al., 2024b) with a GELU (Hendrycks and Gimpel, 2023) activation to project visual features into the word embedding space.

2 Training Details

We implemented a curriculum learning approach to train Aquila-VL-2B in phases following Wang et al. (2024a); Li et al. (2024b). Our training is divided into four stages, progressively increasing the task difficulty, image resolution, and data quality. The training setup is presented in Table 2.

Stage 1: We train the projector using 10M image-caption data to align the visual feature space with the word embedding space. Both the vision tower and language tower are frozen during this phase.

Stage 2: We utilized general visual instruction data for further training to equip the model with fundamental capabilities for solving multimodal tasks. The data was divided into three subsets, and during each stage of training, the maximum visual resolution was progressively increased to enhance the model’s comprehension of visual information.

Stage 3: We employed selective visual instruction data for training and further increased the maximum resolution to improve performance.

Stage 4: We fine-tuned the model using training data from GPT-4 and synthetic data. Experiments demonstrate that scaling with synthetic data can further enhance model performance.

We used the official codes of LLaVA-OneVisionhttps://github.com/LLaVA-VL/LLaVA-NeXT/tree/main/scripts/train. Besides, we also supported the training of Aquila-VL-2B in FlagScale BAAI (2024a), which is a comprehensive toolkit tailored for large models based on open-source projects. Under identical configurations, end-to-end training with FlagScale achieves a 1.7x acceleration compared to DeepSpeed. To efficiently handle large-scale training data without encountering out-of-memory (OOM) errors, we have developed an optimized multi-modal data loader, leveraging Megatron Energon Nvidia (2024). This loader processes datasets in a pre-formatted structure, prepared offline, to ensure efficient data consumption during training. The model was trained on both Nvidia A100 and chips from a Chinese manufacturer.

Evaluation

In this section, we first evaluate the performance of the model through a comparative analysis of multiple benchmarks, demonstrating the advantages of our approach. Subsequently, we conduct a detailed examination of the model’s specific capabilities, including general visual perception, document understanding and mathematical reasoning. Finally, we carry out an ablation study to investigate several key components of our approach.

We assessed the visual capabilities of Aquila-VL-2B using a range of visual benchmarks provided by the VLMEvalKit Duan et al. (2024). Aquila-VL-2B demonstrates highly competitive performance at the same scale, achieving new state-of-the-art results. Specifically, we evaluated the capabilities of Aquila-VL-2B across three task categories.

General Visual Question Answering We conducted extensive evaluations across a diverse array of general visual question answering benchmarks: MMStar, HallusionBench, MMVet Yu et al. (2023), and MMBench-1.1 Liu et al. (2023b). Aquila-VL-2B demonstrated strong performance across these benchmarks, achieving or surpassing state-of-the-art results in most cases at the same scale. On MMStar, which evaluates multimodal capabilities by integrating visual and textual information, Aquila-VL-2B achieved a remarkable score of 54.9, surpassing previous state-of-the-art results and demonstrating its strong proficiency in handling diverse multimodal tasks. On HallusionBench, which evaluates image-context reasoning, Aquila-VL-2B achieved a score of 43.0, surpassing both the previous state-of-the-art and strong baselines, demonstrating its superior ability in understanding and reasoning within complex visual contexts. On MMVet, which evaluates large multimodal models for integrated vision-language capabilities across 16 complex multimodal tasks, Aquila-VL-2B achieved a score of 44.3, demonstrating its ability to address diverse multimodal challenges. On MMBench, which evaluates fine-grained abilities across 20 dimensions, Aquila-VL-2B exhibited strong performance, achieving a score of 76.3 on the English test set, matching the state-of-the-art, and 74.1 on the Chinese test set, demonstrating its robust capabilities in this benchmark.

Knowledge and Mathematical Reasoning We conducted experiments on the AI2D Kembhavi et al. (2016), MMMU Yue et al. (2023), and MathVista datasets to evaluate the model’s capabilities in knowledge and mathematical reasoning. The MMMU dataset is a new benchmark designed to assess multimodal models on extensive multi-disciplinary tasks that require college-level subject knowledge and deliberate reasoning. Aquila-VL-2B achieved a strong score of 47.4, surpassing state-of-the-art results at the same scale and demonstrating its proficiency in addressing complex multimodal challenges. MathVista is a comprehensive benchmark comprising 6,141 diverse examples of mathematical and visual tasks. The Aquila-VL-2B series exhibited exceptional performance on the MathVista benchmark, achieving a score of 59.0, thereby outperforming other large vision language models (LVLMs). AI2D focuses on multiple-choice questions related to scientific diagrams containing text. Aquila-VL-2B exhibited outstanding performance at a comparable scale, achieving a score of 75.0, which represents the highest performance in this benchmark, highlighting its competitive strengths.

Text Reading We assessed Aquila-VL-2B’s capabilities in text reading and diagram comprehension using the OCRBench dataset. OCRBench is a mixed-task dataset that emphasizes mathematical formula parsing and information extraction, in addition to text-based visual question answering (VQA). The performance results highlight potential areas for further optimization.

2 Effects of Sythesis Data

To assess the impact of synthetic data on model performance, we conducted an ablation study. In this experiment, we removed all synthetic data and trained the model using only the original GPT-generated data. The results, as shown in Table 4, revealed a significant decline in overall model performance after removing the synthetic data. This demonstrates that the synthetic data played a crucial role in enhancing the model’s performance, further validating the effectiveness of our approach in data augmentation and diversity.

3 Data Scaling

To further analyze the impact of data size scaling on model performance, we conducted a detailed study on how model performance varies with the amount of training data. The results, shown in Figure 3, indicate a consistent improvement in performance as the training data increases. This trend clearly demonstrates that expanding the scale of instruction data has a significant positive effect on model performance. This observation suggests that as more diverse instruction data is introduced, the model’s ability to handle complex tasks is enhanced. Therefore, scaling up the instruction data is an effective strategy for improving overall model performance.

Conclusion

In this work, to enhance the performance of open-source models, we built the Infinity-MM multimodal instruction dataset with tens of millions of samples, increasing the data volume to improve model efficacy. Besides, we proposed a method for synthesizing instruction data based on open-source models, which further generated high-quality instruction data and expanded the dataset size. Ultimately, we trained the Aquila-VL-2B model using Infinity-MM, achieving state-of-the-art performance for models of comparable size.

References

Appendix A Comprehensive Benchmark Comparisons to SOTAs

In Table 5, we provide a more comprehensive assessment of our model’s performance through extensive comparisons across multiple benchmarks against SOTA models. These benchmarks encompass a diverse range of tasks, including visual recognition, reasoning, and document understanding, offering a robust evaluation of the model’s capabilities.

Appendix B Video Understanding

To enhance Aquila-VL-2B’s ability to process multi-image and video data, we extracted a total of 937K multi-image and video samples from the LLaVA-OneVision dataset, and combined them with 1M single-image samples drawn from Stage 4 for further training. The results, as shown in Table 6, demonstrate that even prior to incorporating the multi-image and video data, our model already exhibited a solid ability to handle video imagery with satisfactory performance. After introducing the additional multi-image and video data for further training, the model’s capacity to process such data was significantly improved.

Appendix C Data composition

We listed the sources, sizes, and types of all our data in Table 7.

Appendix D Instruction Label System

D.2 Fine-grained Perception (single-instance)

D.3 Fine-grained Perception (cross-instance)

D.4 Relation Reasoning

Identify family/ friendship/ professional/ hostile relationships

Identify spatial/ mechanical/ cause-effect relationships

D.5 Attribute Reasoning

D.6 Logic Reasoning

Other Structuralized Image-Text Understanding

Predict trend/ social interaction/ physical movement/environmental changes