Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, Jiaya Jia

Introduction

With the rapid evolution in Large Language Models (LLMs) , empowering the impressive capabilities for multi-modality inputs is becoming an essential part of current Vision Language Models (VLMs) . To bridge the modality gap, several studies are conducted to marry vision with LLMs from images to videos . Despite these advancements, a significant gap remains between academic initiatives and the prowess of well-established models like GPT-4 and Gemini , which are trained with huge amounts of data and resources.

For vision itself, image resolution is a core part of explicitly despite the surrounding environment with minimal visual hallucination. To this end, more attempts are performed to further improve the visual understanding in current VLMs. For instance, LLaVA-Next and Otter-HD are proposed to enhance the ability based on previous work by improving the image resolution. Increasing the number of visual tokens with higher resolution images undeniably enriches visual embeddings in LLMs. However, this improvement comes with escalated computational demands and associated costs, particularly when processing multiple images. Moreover, the existing data quality, model capabilities, and application scopes remain inadequate for accelerated training and development processes. This scenario prompts a critical inquiry: how to push forward the VLMs approaching well-developed models with acceptable cost in an academic setting?

To answer this question, we explore the potential of VLMs from three strategic aspects, i.e., efficient high-resolution solution, high-quality data, and expanded applications. Firstly, we utilize ConvNet to efficiently generate higher-resolution candidates, thus enhancing visual detail while maintaining the visual token count for LLMs. To bolster data quality, we amalgamate high-quality datasets from diverse public sources, ensuring a rich and varied data foundation. Furthermore, our approach integrates these enhancements with cutting-edge LLMs and generative models, aiming to elevate VLM performance and user experience. This multifaceted strategy enables us to delve deeper into the capabilities of VLMs, achieving significant advancements within manageable resource constraints.

In general, our method employs an any-to-any paradigm, which is adept at handling both image and text as input and output. In particular, we introduce an efficient visual token enhancement pipeline for input images, featuring a dual-encoder system. It comprises twin encoders, one for high-resolution images and the other for low-resolution visual embedding, mirroring the cooperative functionality of the Gemini constellation. During inference, they work in an attention mechanism, where the low-resolution one generates visual queries, and the high-resolution counterpart provides candidate keys and values for reference. To augment the data quality, we collect and produce more data based on public resources, including high-quality responses , task-oriented instructions , and generation-related data . The increased amount and quality improve the overall performance and extend the capability of model. Additionally, our model supports concurrent image and text generation, facilitated by the seamless integration of our VLM with advanced generative models . It leverages VLM guidance for image generation by providing the generated text from LLMs.

The Mini-Gemini framework, can be easily instantiated with a range of LLMs from 2B to 34B parameter scales, as detailed elaborated in Section 3. Extensive empirical studies are conducted in Section 4 to reveal the effectiveness of the proposed method. Remarkably, our approach attains leading performance in various settings and even surpasses the well-developed Gemini Pro , Qwen-VL-Plus , and GPT 4V in the complex MMB and MMU dataset, respectively. These results underscore Mini-Gemini’s potential to set new benchmarks in the realm of VLMs, highlighting its advanced capabilities in handling complex multi-modal tasks.

Related Work

Recent progress in Natural Language Processing (NLP) has been dramatically accelerated by advancements in large language models (LLMs). The seminal introduction of the Transformer framework served as a cornerstone, enabling a new wave of language models such as BERT and OPT that exhibit profound linguistic understanding. The inception of the Generative Pre-trained Transformer (GPT) introduced a novel paradigm through auto-regressive language modeling, establishing a robust method for language prediction and generation. The emergence of models like ChatGPT , GPT-4 , LLaMA , and Mixtral further exemplified the field’s rapid evolution, each demonstrating enhanced performance on complex language processing tasks, attributable to their training on extensive textual datasets. Instruction tuning has emerged as a key technique for refining the output of pre-trained LLMs, as evidenced by its application in the development of open-source models such as Alpaca and Vicuna . They iterate on the LLaMA with custom instruction sets. Additionally, the integration of LLMs with specific tools for visual tasks highlights their adaptability and potential for broad application, underscoring the utility of LLMs in extending beyond traditional text-based processing to include multimodal interactions. In this work, we take several pre-trained LLMs as benchmarks and build multi-modality frameworks upon them to further extend the impressive reasoning ability.

Vision Language Models.

The convergence of computer vision and natural language processing has given rise to VLMs, which marry visual and linguistic models to achieve cross-modal comprehension and reasoning capabilities. This integration has been pivotal in advancing tasks that require both visual understanding and language processing, as evidenced by models trained on diverse datasets for understanding and reasoning . Groundbreaking models such as CLIP have further bridged the gap between language models and vision tasks, showcasing the feasibility of cross-modal applications. Recent developments underscore a growing trend toward leveraging the robust capabilities of LLMs within the realm of VLMs. Innovations like Flamingo and BLIP-2 have capitalized on massive collections of image-text pairs to fine-tune cross-modal alignment, significantly boosting learning efficiency. Building upon these advancements, several models have focused on generating high-quality instructional data based on BLIP-2, leading to marked improvements in performance. Furthermore, LLaVA adopts a simple linear projector to facilitate image-text space alignment with minimal learnable parameters. It leverages tailored instruction data and exemplifies an efficient strategy that demonstrates the model’s potent capabilities. Different from them, we aim to explore the potential for both comprehension and generation.

LLM as Generation Assistant.

Combining LLMs with image outputs has emerged as a pivotal area in recent multimodal research. Methods like InternLM-XComposer utilize image retrieval to produce interleaved text and image outputs, bypassing direct generation. Conversely, auto-regressive token prediction approaches, exemplified by EMU and SEED , enable LLMs to decode images through massive image-text data directly. These methods require enormous training resources, and their auto-regressive nature leads to undesirable latency. Recent studies strive to align with latent diffusion models to streamline image generation. They typically require designing text embeddings and additional optimization to achieve the desired generation effect. This joint training can compromise the performance of VLMs in text generation. Mini-Gemini distinguishes itself by adopting a text-data-driven approach to enable the model to generate high-quality images. We leverage a mere 13K pure text data to activate the LLM’s ability as a high-quality re-captioner without undermining the fundamental performance of VLMs.

Mini-Gemini

The framework of Mini-Gemini is conceptually simple: dual vision encoders are utilized to provide low-resolution visual embedding and high-resolution candidates; patch info mining is proposed to conduct patch-level mining between high-resolution regions and low-resolution visual queries; LLM is utilized to marry text with images for both comprehension and generation at the same time.

2 Patch Info Mining

3 Text and Image Generation

With the mined visual tokens TVT_{V} and input text tokens TTT_{T}, we concatenate them as the input to LLMs for auto-regressive generation, as presented in Figure 2. Distinguished from traditional VLMs , the proposed Mini-Gemini supports both text-only and text-image generation as input and output, i.e., any-to-any inference. Despite the image comprehension, we anchor Mini-Gemini’s ability to generate images on its outstanding image-text understanding and reasoning capabilities. Unlike recent works , which address the domain gap between text embeddings of LLMs and generation models, we choose to optimize the gap in the domain of language prompts. Precisely, Mini-Gemini translates user instructions into high-quality prompts that produce context-relevant images in latent diffusion models . This approach is reflected in subsequent high-quality image generation frameworks, such as DALLE 3 and SORA , which leverage the generation and understanding capabilities of VLMs to obtain higher-quality text conditions for generation tasks.

For better cross-modality alignment and instruction finetuning, we collect high-quality datasets from publicly available sources. In particular, for cross-modality alignment, we utilize 558K image-caption pairs from the LLaVA-filtered CC3M dataset and 695K sampled GPT-4V-responded captions from the ALLaVA dataset . It brings about 1.2M image captions in total for projector pretraining. As for instruction finetuning, we sample 643K single- and multi-turn conversations (excluding 21K TextCaps data) from the LLaVA dataset, 100K QA pairs from ShareGPT4V , 10K LAION-GPT-4V captions, 700K GPT-4V-responded instruction pairs from ALLaVA dataset , and 6K text-only multi-turn conversations from LIMA and OpenAssistant2 . To bolster the OCR-related abilities, we further collect 28K QA pairs that comprise 10K DocVQA , 4K ChartQA , 10K DVQA , and 4K AI2D data. In general, there are about 1.5M instruction-related conversations for image comprehension. Moreover, we also collect 13K pairs for image-related generation that will be elaborated on subsequently.

Generation-related Instructions.

To support image generation, we further construct a 13K instruction-following dataset using GPT-4 Turbo. As depicted in Figure 4, the training data encompasses two tasks: (a) Simple instruction re-caption: we adopt 8K descriptive image captions from LAION-GPT-4V and let GPT-4 inversely infer the corresponding user’s short input and the target caption in the Stable Diffusion (SD) domain. (b) In-context prompt generation: based on a few high-quality real-world conversation contexts in LIMA and OpenAssistant2 , we generate prompts that produce images suitable for the conversation context, bringing 5K instructions in total. For both kinds of data, in each query to GPT-4, we randomly sample 5 high-quality SD text-to-image prompts from GigaSheet as in-context examples to obtain target prompts for generation. We format our data to use as a trigger to initiate the generation process and wrap the target caption within .... Following text generation, Mini-Gemini extracts target captions and utilizes SDXL to generate the corresponding image. More details are discussed in Appendix A.

Experiments

In this section, we first outline our experimental framework, commencing with the experimental setup. Subsequently, we compare Mini-Gemini with leading methods on various benchmarks. Component-wise analysis and qualitative results are given at the end of this section.

In this study, we instantiate Mini-Gemini with the CLIP-pretrained ViT-L for LR vision encoder and the LAION-pretrained ConvNext-L for HR vision encoder. For efficient training, we keep two vision encoders fixed and optimize the projectors of patch info mining in all stages. Meanwhile, we optimize the LLM during the instruction tuning stage only. Regarding the training scheme, we optimize all the models for 1 epoch with the AdamW optimizer and a Cosine learning schedule. In most cases, the initial learning rates for modality alignment and instruction tuning are respectively set at 1e−31e^{-3} and 2e−52e^{-5}, with an adjusted rate of 1e−51e^{-5} for the Mixtral-8×\times7B and Hermes-2-Yi-34B to ensure stable instruction tuning. The framework involves training on 8×\timesA800 GPUs for standard machine configurations. For the largest model with Hermes-2-Yi-34B, we leverage 4 machines and complete the optimization within 2 days with DeepSpeed Zero3 strategy. For the HD version, the total cost is enlarged to about 4 days because of the extended visual tokens in LLMs.

Datasets.

For model optimization, we construct high-quality data for cross-modality understanding and generation. It mainly includes 1.2M caption pairs for modality alignment and 1.5M single- or multi-round conversations for instruction tuning, as elaborated in Section 3.3. Moreover, we report results on widely-adopted zero-shot image-based benchmarks, including VQAT{}^{\text{T}} (TextVQA) , MMB (MMBench) , MME , MM-Vet , MMMU , and MathVista datasets.

2 Main Results

In Table 1, we compare with previous leading approaches across several settings, including normal and high resolution, and also consider private models. At normal resolution, Mini-Gemini consistently outperforms existing models across a wide range of LLMs. In the efficient model category, Mini-Gemini, when configured with Gemma-2B , demonstrates superior performance compared to the efficient MobileVLM and even surpasses InstructBLIP equipped with Vicuna-7B and even 13B. The scalability of Mini-Gemini is evident when larger LLMs are employed. Given the same LLM, the proposed Mini-Gemini is validated to surpass LLaVA-1.5 with a large margin across all benchmarks. Notably, with the Hermes-2-Yi-34B LLM, Mini-Gemini achieves exceptional results, outpacing high-resource private models like Qwen-VL-Plus and Gemini Pro in some challenging benchmarks like MMMU and MMB .

High Resolution.

To validate the framework for extended visual tokens, we perform experiments with an input size of 672 for LR visual encoder and 1536 for HR visual encoder in Table 1. As discussed above, the HR visual encoder primarily serves to offer high-resolution candidate information. Importantly, despite the increased resolution, the effective number of visual tokens processed by the LLM remains consistent with the LR input size of 672, ensuring computational efficiency. The benefits of this approach are particularly evident in detail-oriented tasks. For example, in the TextVQA benchmark, our method achieved a performance rate of 74.1% with the Hermes-2-Yi-34B configuration, closely matching the performance of the well-established Gemini Pro . Detailed results in Table 1 show that Mini-Gemini excels in more challenging benchmarks as well. For instance, the proposed method is on par with Qwen-VL-Plus on the MathVista and MMMU benchmark and even surpasses Gemini Pro and GPT-4V on the widely-adopted MMB benchmark.

3 Component-wise Analysis

We first delve into the proposed patch info mining and report results in Table 2. It is clear that the model achieves significant gains with the ConvNeXt-L integrated as the vision encoder for HR images. For example, when the LR and HR are respectively set to 224 and 512, the model increases 4.0% and 18.1 in TextVQA and MME datasets. Elevating the HR resolution to 768 further widens the performance margin, achieving a 5.7% uplift in TextVQA compared to the baseline. These results underscore the substantial impact of patch info mining in harnessing more detailed visual cues. When we further extend the LR resolution to 336, patch info mining still contributes consistent gains. For instance, with the default ConvNeXt-L as vision encoder, it surpasses the baseline with 3.3%, 6.3, and 3.5% in TextVQA , MME , and MM-Vet dataset, respectively. This proves the capability of designed modules with input resolution scaled up.

Vision Encoder.

To investigate the effect brought by mining candidates, we conduct experiments with various HR vision encoders in Table 2. Compared with the default ConvNeXt-L, we add two encoders for contrast trials, i.e., ConvNeXt-B, and ConvNeXt-XXL. With the basic ConvNeXt-B, the model performs better in TextVQA and MM-Vet . However, the ConvNeXt-L encoder consistently delivers peak results, especially in the MME and MM-Vet datasets, indicating a superior balance in handling detailed visual information. We can conclude from the table that a larger vision encoder for HR images contributes more to the candidate quality, but the model converges with a too large encoder like ConvNeXt-XXL. Hence, considering the balance between effectiveness and computational efficiency, ConvNeXt-L is chosen as the default HR vision encoder. This decision is based on its ability to provide high-quality visual information mining while maintaining reasonable computational demands, as evidenced by the comparative performance across the benchmarks.

High-quality Data.

In this era, the significance of high-quality data for enhancing the capabilities of LLMs and VLMs cannot be overstated. In our comprehensive analysis of data combination effects, presented in Table 3, we begin with a baseline model incorporating patch info mining. The integration of high-quality captions from ShareGPT4V yields improved visual alignment and performance gains. We validate the zero-shot performance on the TextVQA benchmark, notably removing TextCaps data from the training set in line with previous studies . This modification led to a notable performance decrease, underscoring the value of specific data types in training. To counteract this decline, we incorporate additional high-quality captions from LAION-GPT-4V and OCR-specific data, thus enhancing the model’s OCR reasoning capabilities. More details are provided in the appendix. As elaborated in Section 3.3, we utilize generation-related instructions to expand the application. It is interesting to find that such data also benefits the image understanding ability and brings 3.3% gains in MM-Vet dataset. Moreover, with the high-quality GPT4V responses from ALLaVA dataset, the framework respectively pushes the baseline over 7% and 9% in TextVQA and MM-Vet datasets. This comprehensive evaluation underscores the pivotal role of strategic high-quality data integration in amplifying the potential of the Mini-Gemini framework.

Visual Token Extension.

As depicted in Figure 3(b), the proposed patch info mining is adeptly designed to accommodate extended visual tokens, thereby generalizing its utility across different input resolutions. We validate the effectiveness of the token extension in Table 3. When increasing LR and HR input resolution, the model achieves significant gain in all benchmarks. Notably, in detail-oriented tasks such as TextVQA, we observe a performance uplift of over 3%, indicating a significant enhancement in the model’s ability to handle complex visual data. Our empirical observations suggest that the increase in resolution significantly diminishes visual hallucinations, leading to more accurate and reliable image comprehension. Generally, with the increased visual token number, Mini-Gemini can be scaled up towards better capability. We can also draw the same conclusion from high-resolution results in Table 1.

4 Qualitative Results

To ascertain the visual comprehension prowess of Mini-Gemini in real-world settings, we apply it to a variety of understanding and reasoning tasks in Figure 5. Thanks to the patch info mining and high-quality data, Mini-Gemini can well solve several complex cases. For example, it is capable of recognizing plotted curves in graphical data and directly translating them into Python code for immediate application. Beyond mere recognition, it exhibits a keen attention to detail, accurately describing intricate elements within complex indoor scenes, and demonstrating a nuanced understanding of character associations in memes. Moreover, Mini-Gemini’s analytical capabilities extend to chart analysis and practical problem-solving, such as intelligence tests.

Image Generation.

In Figure 6, we provide a comprehensive evaluation of Mini-Gemini’s generation capabilities. Compared with recent studies such as AnyGPT and ChatIllusion , our stronger multi-modal understanding ability allows us to generate text-to-image captions that better align with the given instructions, resulting in more contextually appropriate image-text answers. A noteworthy point, as shown in Figures 1 and 6, is its proficiency in generating high-quality content based on multi-modal human instructions, with text-only training data. This capability underscores Mini-Gemini’s robust image-text alignment and semantic interpretation skills, which come into play effectively in the inference stage. By leveraging the powerful reasoning ability of the LLM, it can produce reasonable image-text outputs in single or multi-round conversations.

Conclusion and Discussion

We presented Mini-Gemini, a streamlined and potent framework for multi-modality VLMs. The essence of Mini-Gemini is to harness the latent capabilities of VLMs through strategic framework design, enriched data quality, and expanded functional scope. At its core, patch info mining enables efficient extraction of detailed visual cues by engaging with high-resolution candidates. From the data perspective, our meticulously compiled high-quality dataset ensures accurate vision-language alignment and bolsters strong instruction-following ability. Furthermore, we support reasoning-based generation in Mini-Gemini and empower current VLMs with any-to-any workflow. Extensive experiments on several zero-shot benchmarks prove the superiority of the proposed method, which surpasses previous leading approaches and even private models. We hope the Mini-Gemini can serve as a strong benchmark for image understanding and VLM-guided generation.

Although Mini-Gemini achieves good results, it still has great potential to be further explored. For visual comprehension, the counting ability and complex visual reasoning ability are still far from satisfactory. This could be attributed to the lack of corresponding training data especially in the pretraining stage. Meanwhile, for reasoning-based generation, we use text to bridge the VLM and diffusion model in this work because we do not find apparent gain with embedding-based approaches. We will try to find a more advanced manner for visual understanding, reasoning, and generation.

Appendix

Appendix A Data Collection Details

In this section, we delve into the specifics of OCR-related data collection. Natural images can be easily annotated with detailed captions, but text-rich figures or diagrams, such as documents , charts , and scientific diagrams , present a more intricate challenge for models in complex questions and answers. Therefore, to facilitate the optimization process, we follow the strategy in TextVQA and incorporate OCR tokens for model reference in the training phase. In particular, we utilize the PaddleOCR to initially identify and extract textual elements within each image. Then, we append the text characters to the original conversations in a format of Reference OCR token:Text_1,...,Text_n, where Text_1 to Text_n indicates nn detected text strings. This approach ensures that the model has access to textual representations from the images, enhancing its ability to understand the image content. It is important to note that the OCR detector is utilized solely for generating enriched data and is not employed during testing. This distinction underscores our objective to train Mini-Gemini with a comprehensive understanding of both textual and visual elements, thereby improving capability in complex image-text scenarios.

Generation Data Collection.

For the data generation collection described in Section 3.3, we provide specific examples of query prompts and their corresponding reference data sources for two generation tasks in Figure 7. We commence with a corpus comprising 10K GPT4V caption data and 6K English-only LLM SFT data. After filtering out results that did not meet the format requirements, we ultimately obtained 13K data points. To enhance the contextuality and quality of the queries, we incorporate two distinct types of in-context examples: get_example_captions() and get_example_queries(). The former function randomly selects 5 high-quality Stable Diffusion (SD) Text-to-Image (T2I) prompts, while the latter extracts 3 instances from a repository of simple instructional templates. These in-context examples serve as a foundational guide, providing diverse and representative prompts that significantly enrich the generation process. This strategic approach ensures the production of high-quality, relevant data, effectively supporting the generative capabilities.

Appendix B Extended Showcases

In this section, we further provide more cases to validate the generality and capability of Mini-Gemini in various environments. As presented in Figures 8 and 9, Mini-Gemini can well answer detail-oriented questions and solve OCR-related and scientific problems. For image generation, we present more examples in Figure 10 that include direct T2I generation, multi-round conversation, reasoning-based generation, storytelling, and in-context generation. They further prove the superiority of Mini-Gemini in both visual comprehension and generation.

References