MANTIS: Interleaved Multi-Image Instruction Tuning

Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, Wenhu Chen

Introduction

Large Multimodal Models (LMMs) have advanced significantly in recent years. Both closed-source models like GPT-4V , Gemini , Reka , MM1 and open-source models like LLaVA , LLaVA-NeXT , BLIP , CogVLM , Idefics have shown strong visual-language understanding and generation capabilities in single-image tasks like VQA , TextVQA , GQA , etc. Despite the significant interest in improving LMMs’ performance on single-image vision-language tasks, relatively less work attempts to improve LMMs’ capability to solve multi-image vision-language tasks. To the best of our knowledge, only a few open-source models (e.g. Idefics , OpenFlamingo , Emu2 ) and VILA support multi-image inputs.

We argue that multi-image visual ability is also a crucial in real-world applications. Specifically, we categorize the ability into four skills including (1) Co-reference: understanding the references like “second image” in the natural language expression and grounding it on the referred image to generate a response. (2) Comparison: capturing the nuances and commonalities between several images. (3) Reasoning: capturing the information across multiple images and reasoning over these multiple pieces to derive the response. (4) Temporal understanding: observing multiple frames of a video or a scene to understand the temporal information like actions, behaviors, interactions, etc. These are the four major skills we want the LMMs to encompass.

The existing multi-image LMMs like Idefics , OpenFlamingo , Emu2 , MM1 , VILA heavily rely on pre-training on a massive interleaved image-text web documents like MMC4 , Obelics . For example, MMC4 contains 500 million examples, while Obelics contains 140 million examples. Such a pre-training procedure would require a massive amount of computation, which hampered its adoption in most open-source LMM training pipelines. In this paper, we attempt to build LMMs with academic-level resources to support interleaved multi-image inputs. Specifically, we made a few efforts along this line:

Dataset: We create the first multi-image instruction-tuning dataset Mantis-Instruct. It has a total of 721K instances, consisting of 14 subsets to cover all the multi-image skills. Among the 14 subsets, 10 subsets are from the existing datasets. For example, NLVR2 , IconQA , etc are used to cover ‘reasoning’ skill; DreamSim , Birds-to-Words , etc are used to cover ‘comparison’ skill; NExT-QA , STAR , etc are used to cover ‘temporal understanding’ skill. We additionally curate four new datasets LLaVA-665k-multi, LRV-multi to cover the ‘coref’ skill and Contrast-Caption, Multi-VQA to broaden the ‘reasoning’ skill. Architecture: Similar to LLaVA , we start from a pre-trained language model like LLaMA-3 or Fuyu and a vision transformer encoder from CLIP or SigLIP . We use the same multimodal projector as LLaVA to map the vision embeddings to the text embeddings space to simply concatenate them with the text embeddings. We also define a text-image interleaving format: “…(image {i}: image embeddings)…”, where and are the image delimiters. We explored multiple instantiations and found this optimal design. Evaluation: We manually curate a new challenging dataset Mantis-Eval to cover several multi-image skills. This evaluation can help us better analyze LMMs’ multi-image abilities.

We train Mantis by combining Mantis-Instruct with another 268K single-image vision-language data to balance the multi-image and single-image abilities. To evaluate the multi-image skills, we adopt four existing benchmarks NLVR2 , QBench , BLINK , MVBench and Mantis-Eval to cover all the mentioned skills. We include 15 competitive baseline models including single-image LMMs like LLaVA-Next , Qwen-VL and multi-image LMMs like GPT-4V, Idefics1/2, etc. For single-image LMMs, we merge multiple images horizontally as a single-image input. By training on 16 x A100-40G for 36 hours, Mantis achieves state-of-the-art performance on all five multi-image tasks, excelling on all the multi-image skills. Notably, Mantis outperforms the strongest baseline Idefics2 by an average of 11 absolute points. The held-in and held-out results are equivalently strong, which shows the generalization ability of Mantis. These results are encouraging because we show that the low-cost instruction tuning on 721K high-quality data can lead to a much better generalization performance than other LMMs intensively pre-trained on 100x larger datasets. Mantis also preserves its strong performance of single-image tasks and is on par with CogVLM and Emu2 . We also present insights on the architecture selection and inference approaches based on the results in section 3. We believe Mantis-Instruct can serve as an essential baseline for future studies on multi-image LMMs due to its simplicity and low cost. We will release all the code, data, and models to help with the reproducibility of our results.

Mantis: Interleaved Multi-Image Instruction Tuning

Most existing LMMs do not support multi-image inputs, either due to the architecture design , or the lack of associated code support . To enable multi-image training and inference, we have modified the LLaVA architecture and added code support for it. Idefics2 already supports multi-image as inputs naturally, we thus inherit their structure.

Image Context Length:

How many images can a multi-image supported model accept? This is limited by both the LLM backbone’s max token length and the number of image tokens the vision encoder generated. For the LLaVA architecture, each image consumes a fixed (336/14)2=576(336/14)^{2}=576 image tokens. For Llama3 with a maximum window size of 8K, we can feed at most 14 images. For Idefics2, the number of tokens per image is resampled to be 6464 tokens only due to the existence of perceiver, which is pretty efficient. Given 8K context length, it can accept 128 images at most, even more than some video models.

Text-Image Interleaving Format:

A proper text-image interleaving format can make an LLM easier to acquire multi-image understanding and reasoning ability. We contend that a good text-image interleaving format should: (1) mark boundaries between images clearly, and (2) denote the serial number of images. Following this principle, we designed our interleaving format as follows: "(image {i}: )", where is the begin of image token and is the end of image token. is the placeholder for image patches. This format adds clear separators between images, and gives LMM serial information of the image through "image {i}". In practice, we set and to be and respectively. Previous work also demonstrates the effectiveness of this approach through a comprehensive ablation study .

2 Mantis-Instruct: A large scale multiple image question answering dataset

We construct Mantis-Instruct from multiple publicly available datasets. Detailed statistics are shown in Table 1 with several newly curated subset by us. We use a similar dataset format with LLaVA’s, where each data item contains multiple images and multiple turns of QA pairs are gathered. We demonstrate some examples of our newly curated subset in Figure 2. We report the number of examples, the number of images per item, the average conversation turns, and the average length. Mantis-Instruct is collected based on the 4 multi-image skills. We here describe the gathered subsets for each skill briefly. Construction details can be found in subsection A.1.

requires the model to create a mapping from natural language references, such as "second image," to the actual images in the input. It does not require the model to infer across multiple images. To achieve this, we constructed LLaVA-665k-multi and LRV-multi by concatenating multiple single-image conversations into a multi-image sequence. Deliberately, we included natural language references like "For the second image" for each question.

Comparison

requires the model to not only have good co-reference ability but also be able to compare and understand differences across multiple images. We gather Co-Instruct, Dreamsim, Spot-the-Diff, and Birds-to-Words together to enhance the model’s comparison ability. They cover a wide range of topics such as image quality, visual similarity, and difference description.

Reasoning

requires the model to further make inferences by combining its world knowledge with the information collected using the co-reference and comparison ability. It distinguishes from the comparison skill by introducing more complex human instruction associated with the images. We gather NLVR2, IconQA, Contrast-Caption by reformatting captioning datasets, ImageCoDe, and our self-collected Multi-VQA to further equip the model with reasoning ability. The topics include logical reasoning, counting, image matching, image retrieval, and free-form multi-image QA.

Temporal Understanding

requires the model to understand temporal image sequences such as videos, comics, etc. We include a small number of video understanding datasets (14K), VIST, NExT-QA, and STAR, to activate the model’s temporal understanding ability. We contend that a model with the above 3 multi-image skills shall be easily tuned to understand image sequences.

Experiments

In this section, we aim to evaluate Mantis’s ability on both multi-image and single-image vision-language tasks. For multi-image tasks, depending on whether the LMM can handle multiple images, we either concatenate multiple images (‘merge’) horizontally as a single image or feed multiple images as a ‘sequence’ to the model. We have trained four model variants listed in Table 2, where we list the architecture and dataset details. Mantis-CLIP and Mantis-SigLIP only adopt the CC3M 560K subset for feature alignment. Mantis-Flamingo and Mantis-Idefics2 are initialized from OpenFlamingo and Idefics2, respectively, which have been intensively pre-trained on the multi-image corpus. After this pre-training and preparation, we fine-tune them on Mantis-Instruct to get Mantis. Please refer to subsection A.3 for details of hyper-parameters and settings.

For ‘merge’ evaluation, we include BLIP-2-13B , InstructBLIP-13B , Qwen-VL-Chat , LLaVA-1.5-7B , LLaVA-1.6-7B , Kosmos2 , and Fuyu-8B . For ‘sequence’ evaluation, we include OpenFlamingo , Otter , VideoLLaVA , Emu2-Chat-34B , VILA , Idefics1 , Idefics2 , GPT-4V . Additionally, for single-image tasks, we include InstructBLIP-7B-Vicuna , Yi-VL-6B .

2 Evaluation Benchmarks

For multi-image tasks, we use 2 held-in benchmarks: NLVR2 and Qbench ; and 3 held-out benchmarks: Mantis-Eval, BLINK , and MVBench . We report detailed statistics in Table 3. NLVR2 evaluates the model’s ability to conduct logical reasoning across the contents of images. It asks the model to compare the contents of the two given images and judge whether a given statement is correct or not. All questions are multiple-choice. We use the test-public split for evaluation. Qbench is a benchmark evaluating whether LMMs can properly judge and compare the quality of a benchmark. Q-bench aims to evaluate the low-level visual abilities of LMMs, such as judging the image quality. All questions are multiple-choice. In our experiments, we evaluate the Qbench2-A2-pair dev set, where a low-level visual question is asked based on multiple image contents. Mantis-Eval consists of 217 multiple-image reasoning examples that cover different topics, including size perceptions, weight comparisons, etc. This dataset is carefully curated by our annotators, where images are acquired via Google Search, and then come up with a proper question which requires understanding the contents in the two images well to answer. Mantis-Eval contains both multiple-choice and short-answer questions. We report results in the test split. BLINK is a benchmark on core visual perception abilities, where most of the tasks can be solved by humans “within a blink” (e.g., relative depth estimation, visual correspondence, forensics detection, and multi-view reasoning). Some questions in the benchmark involved multiple images, such as image similarity comparison. We report results in the validation set of the benchmark. MVBench is a comprehensive multimodal video understanding benchmark covering 20 challenging video tasks. The questions cannot be solved with a single frame, which necessitates understanding an image sequence to solve the questions in the benchmark. We report results in test split.

3 Results on Multi-Image Tasks

We report our results on the multi-image task in Table 4. We put the GPT-4V’s results at the top of the table as it’s a closed-source LMM. We can see that our best version Mantis-Idefics2 already matches the overall performance of GPT-4V. We put comparison results between Mantis and other open-source LMMs, along with the insights derived in the following paragraphs.

Mantis performs well on held-in evaluation We found that Mantis-SigLIP without any multi-image pre-training can already achieve best known performance on these two benchmarks. Mantis-Idefics2 has better performance of 89.7189.71 and 75.2075.20 on NLVR2 and Q-Bench respectively, marking it the SoTA for these two tasks. Notably, Mantis-Idefics2 also beats the Idefics2 by 2.842.84 points, which was already trained on NLVR2. It shows the cross-task benefits of our instruction tuning. The remarkable performance on these 2 benchmarks demonstrate that Mantis has gained strong ability in multi-image logical reasoning and low-level vision understanding. Mantis generalizes well on held-out evaluation We show that our models can attain strong performance on other held-out benchmarks, Mantis-Eval, BLINK (val), and MVBench. Mantis-SigLIP achieves 59.4559.45 on Mantis-Eval, while Mantis-Idefics2 achieves 49.0549.05, and 51.3851.38 on BLINK and MVBench respectively, surpassing all other baselines on these held-out tasks. The gain over the existing strongest baseline is particularly tremendous on Mantis-Eval and MVBench. Multi-Image Pre-training is not necessary In these experiments, Mantis-CLIP and Mantis-SigLIP are attaining much better performance than Mantis-Flamingo and similar performance as Mantis-Idefics. These results reveal that the multi-image pre-training (on massive web corpus) is not a necessary path toward strong multi-image performance. ‘Merge’ vs ‘Sequence’ input The ‘merge’ evaluation is an important baseline to evaluate the multi-image understanding ability of those LMMs that can only accept one image at a time. LLaVA-v1.6’s dynamic image resolution feature can divide a single image into multiple tiles, processing the ‘merge’ image like multiple images . But it still performs poorly on multi-image tasks, only getting 45.6245.62 and 39.5539.55 on the Mantis-Eval and BLINK benchmark, which are at least 10 absolute points lower than Mantis. This result indicates the advantages of building models to accept sequence image inputs.

4 Results on Single-Image Tasks

We also evaluate Mantis-CLIP and Mantis-SigLIP on various single-image tasks, including TextVQA , VQA-v2 , MMBench , MMMU , etc. Results are shown in Table 5. All the evaluations are conducted with the help of LMMs-Eval tool.

Compared to a set of popular LMMs, Mantis-SigLIP gets significant improvements on MMBench-English, MMMU, and ScienceQA benchmarks though we did not specifically optimize the single-image abilities. Solving problems in these tasks requires not only the visual perception ability but also the good reasoning ability of the LLM backbone. Considering that we are using the powerful LLaMA-3 as the backbone LLM, these improvements are expected. We also compare the results of Mantis-Idefics2 against Idefics2 on single-image tasks. On several tasks like MMB, MMMU, and OKVQA, Mantis-Idefics maintains a similarly strong performance as Idefics2. However, we do observe some drops in ScienceQA, MathVista, which are mainly due to a lower mix ratio of OCR and Math data in our instruction tuning. Overall, the average drop of 4% is in the tolerable range.

5 Ablation Studies

Is multi-image instruction tuning necessary? To investigate, we further train a Mantis-Flamingo based on OpenFlamingo , which has been trained on MMC4 , a text-image interleaved pre-training dataset. As shown in Table 4, the average performance of Mantis-Flamingo increased by 19.219.2 points after further fine-tuning on Mantis-Instruct. The results show that there is still a huge improvement space, even if a model is fully pre-trained on a multi-image dataset, which demonstrates the necessity of learning multi-image skills in the fine-tuning phase.

Is multi-image pre-training necessary? To investigate, we design an ablation study, where one model is pre-trained on the CC3M 560K subset, which contains only image-text pairs, while the other one is pre-trained additionally on OBELICS-100K subset, which contains interleaved multi-image-text pairs. Due to limited computation resources, we fine-tune both models on a downsampled Mantis-Instruct dataset to compare their performance. Results in Table 6 show that there are little improvements by adding additional pre-training on OBELICS-100K. We thus conclude that multi-image pre-training could be useful, but it is not a necessary path for models to gain multi-image abilities. The poor performance on Mantis-Flamingo also confirm our hypothesis.

Ablation study of Mantis-Instruct components: We conduct a data ablation study to analyze the effects of different subsets in Mantis-Instruct to enhance a model’s multi-image ability. As shown in Table 7, As new subsets of Mantis-Instruct are added continually, the overall performance is continually improving, proving that every subset is contributing positively.

6 Case Studies

To conduct a qualitative analysis of Mantis models on specific multi-image scenarios, we include some case studies in Figure 3, where we include Emu2, which can accept a sequence of images as inputs, and LLaVA-1.5, which can only accept one image at a time, and the powerful GPT-4V, which perform well on both single-image and multi-image scenario.

The first case requires the model to count the number of dice respectively in the two images. While previous LMMs might have learned good counting ability in the single-image scenario, most of them failed in the multi-image counting problem, including GPT-4V. The second case requires the model to involve their world knowledge to infer the mood of the character in three images. However, Emu2, and LLaVA-1.5 output similar contents, and both fall back into the simple captioning behavior. It demonstrates that these models struggle to distinguish two images as separate information. In contrast, Mantis can correctly do the task due to the existing of reasoning subset in Mantis-Instruct.

Related Works

Large Multimodal Models (LMMs) are mostly denoted as large language models that can understand multiple modalities besides the human language. Some works aim to fuse image, audio, video, or more modalities with language , while more works mainly focus on better-combining vision knowledge with the language . BLIP-2 has curated a large-scale image captioning dataset, bootstrapping a language model along with a vision encoder to be a powerful multimodal model . Following this, LLaVA further proposes a cheaper approach to train a powerful LMM through visual instruction tuning . LLaVA-Next further optimizes single-image performance, with a cost of an increasing number of image tokens per token , consuming more than 2k for each image, which is about 4 times of the original LLaVA model. Later LMMs like QwenVL , CogVLM , Yi-VL , etc. all follow a similar architecture of LLaVA. However, real-world visual problems are way more complex and usually require a sequence of images to describe the situation. Although better performance is achieved in the single-image scenario, it’s doubtful whether it’s worth the cost.

2 Multi-Image Supported Large Multimodal Models

Previous works have also noticed the importance of multi-image ability. Deepmind’s close-sourced Flamingo focuses on optimizing in-context learning for a sequence of image/text pairs but lacking free-from image-text interleaved training examples. Kosmos2 , Fuyu both support text-image interleaved format given the model structure, but they did not optimize in multi-image reasoning, and the released inference codes also do not support multi-image as input. Emu2 is a generative multimodal model that supports interleaved text-image inputs, as well as generating both images and texts. However, it’s mainly focused on in-context examples like Flamingo and auto-regressive image generation, instead of the 4 skills mentioned in section 1 Another common variant of LMMs that can accept multiple images as input is video understanding models, such as Video-LLaMA . However, as shown in the MVBench benchmark, it is also shown that Video-LLaMA behaves worse than LMMs that can only accept one image as input , which raises doubts about its multi-image understanding ability. Different from these works, Mantis aims to optimize the multi-image ability of LMMs guided by the 4 defined skills, co-reference, reasoning, comparison, and temporal understanding, thus equipping the model with better multi-image ability.

3 Multi-Image Evaluation

While many benchmarks have been developed for single-image tasks, such as TextVQA , VQA-V2 , MMBench , and so on, the multi-image evaluation did not capture much attention until recently. Difference description is the first task which large-scale data and good evaluation benchmarks are developed for, including works like Spot-the-Diff , Birds-to-Words , etc. They are usually evaluated by n-gram metrics like BLEU, and ROUGE, and only involve two images. The later NLVR2 further includes logical reasoning across 2 images into evaluation, focusing on a higher level of cognition ability. Q-bench and BLINK are both proposed for low-level multi-image vision tasks, where humans easily do but model struggles. Another set of multi-image evaluations comes from video or image sequence evaluation, including Mementos , MVBench , etc. In this case, the model usually needs to take more than 2 images to answer. Recently, there are also attempts to prompt LMMs to evaluate vision-language task , or image-generation tasks , which also demands good multi-image reasoning ability. The accuracies of the these tasks of different LMMs can be seen as another form of benchmark for multi-image ability.

Conclusion

We propose Mantis, a new large multimodal model family designed to be able to accept interleaved text-image inputs. Mantis is fine-tuned on our Mantis-Instruct, a text-image interleaved dataset consisting of 721K examples, covering 4 crucial skills, including co-reference, reasoning, comparison, and temporal understanding, for multi-image tasks. We also propose Mantis-Eval, a high-quality multi-image evaluation benchmark consisting of 217 examples with an average of 2.5 images per example. Comprehensive experiments show that Mantis models effectively acquire the 4 multi-image skills and achieves state-of-the-art performance on 5 multi-image evaluation benchmarks, and preserve decent single-image reasoning performance. We shared insights about the model architecture and inference approaches based on the experiment results. Extending the image sequence context length and compressing the number of tokens per tokens efficiently are regarded as our future work.

Limitations

Our work mainly focuses on improving the multi-image ability of LMMs. Our collected dataset Mantis-Instruct reformat some existing multi-image datasets, which might still contain some noise due to the quality of the original collection. Besides, we found that most of the examples in the Mantis-Instrut tend to be multiple-choice QA questions, where the response tends to be short, resulting in the current version of Mantis models tending to output shorter responses instead of long reasoning texts and falling short in instruction-following ability in the multi-image scenario. Although we have introduced Multi-VQA subset to alleviate this issue, this problem still exists, and we will leave it to future work by introducing more long-form multi-image data. While enhancing the multi-image ability well, we find out that there is a trend of performance degeneration in the single-image performance compared to the original model Table 5. Some single-image QA datasets have been introduced to alleviate this issue, but the conflicts between multi-image and single-image ability seem to exist continually. How to balance these two abilities will also be our future direction.

Societal Impacts

Our proposed model, Mantis, and released dataset, Mantis-Instruct, can have substantial impacts on society through the application of Large Multimodal Models. LMMs with multi-image ability can analyze, and reason in real-world scenarios, helping machines, and robots, to make decisions automatically, assisting people in website browsing, travel planning based on multiple maps and pictures, etc. We believe Mantis can serve as one of the first useful open-source LMMs that make progress in these applications to make these scenarios possible and efficient.

However, there is a potential that our Mantis can generate hallucinations, make the wrong decisions, and fail to reason across complex real-world scenarios. While we have tried to improve Mantis’s accuracy and performance in instruction-following and multi-image understanding by carefully training on high-quality Mantis-Instruct, Mantis can still sometimes make mistakes, thus harming its faithfulness as a useful LMM tool. There is also the potential for misuse of our model and dataset after we make them open source. Though we have added proper licenses and terms of usage to our assets, misuse cannot be safely eliminated.

References

Appendix A Appendix / supplemental material

We construct Mantis-Instruct from multiple publicly available datasets. Detailed statistics are shown in Table 1. We use a similar dataset format with LLaVA’s, where each data item contains multiple images and multiple turns of QA pairs are gathered. We report the number of examples, the number of images per item, the average conversation turns, and the average length. We describe the task type of each dataset and the processing techniques used in the following:

LLaVA-665k-multi: Besides the above-mentioned multi-image reasoning tasks, we also manually reformat the LLaVA-655k dataset from LLaVA-1.5 into a "synthetic" multi-image dataset. Since each original data item consists of multiple QA pairs about a single image, we randomly merge 2 to 4 data items into a multi-image data item. Then we add proper image denotation like "For the second image, …" for each question, and shuffle all the QA pairs to form the final multi-image data item. Answering each question in this "synthetic" dataset does not require reasoning across images, but demands high image co-reference ability. Besides, it also helps avoid forgetting single-image ability in multi-image scenarios and increases instruction-following ability. Co-Instruct : Co-Instruct is a dataset curated to compare the quality of 2-4 images in the format of open-ended answer or multi-choice QA format. It’s similar to the tasks in the Qbench . NLVR2 : NLVR2 is a natural language reasoning dataset grounded in images. The task is to determine whether a statement is true or false given the visual contents of 2 images. NLVR2 covers many linguistic phenomena, such as cardinality, existential quantifiers, universal quantifiers, etc. Dreamsim : DreamSim proposes a dataset to train a model to judge image similarity. We reuse their datasets, and the task is to judge which candidate image is more similar to a reference image. ImageCoDe : ImageCoDe is a multimodal challenge called image retrieval from contextual descriptions, where a multimodal model is required to retrieve/select the correct image from a set of 10 minimally contrastive images. For each image, 9 similar images are retrieved from Open Images Dataset V6 using CLIP encodings. Contrast-Caption: Inspired by the paradigm of contrastive learning, we reformat existing image captioning tasks into a multi-image version, where 2 tasks are defined. Given multiple images, the first task we defined is to judge which image matches the provided caption, and the second task is to generate a caption for a denoted image. We randomly select 2 to 8 images along with their captions to form a single item, then apply the template of the 2 tasks defined. We used the images and captions provided by ShareGPT-4V and LAION GPT-4V captions from LVIS . Spot-the-diff : This is a dataset that requires the model to generate the difference description given 2 images. We use the processed version from MIMIC-IT . LRV-multi : This is a single-image dataset used to mitigate hallucinations of LMMs. We process it similarly to LLaVA-665k-multi to get a "synthetic" multi-image dataset. Birds-to-Words : The task of this dataset is to generate different descriptions and give 2 images of different birds. We directly take the training split for training. Besides, we prompt GPT-3.5-turbo to generate a possible question given the reference answer, so we can get a diverse question pool. Multi-VQA: This dataset contains multi-image reasoning QA pairs, where each question is generated by GPT-4 based on the captions of the images. Each question is required to involve at least two images. Images and captions are sampled from ShareGPT-4V . The prompt template is in Table 8. Video QA: We also include some video question answering datasets, including NExTQA , STAR , and visual-story-telling . nextqa and star are the processed version from Flipped-VQA . The visual-story-telling are the processed version from MIMIC-IT . Single-image-VQA: Instead of focusing only on multi-image tasks, we claim that it’s necessary to also include some single-image VQA dataset to avoid catastrophic forgetting on single-image tasks. We included some of the single image data from LLaVA-665k as well as DocVQA , DVQA , and ChartQA to enhance its ability on diagrams and OCR. In practice, since the size of DVQA is too large, we only use 30k out of 200k examples for the training.

A.2 Heuristics of data curation

Similar to visual instruction tuning , data curation is an important step to ensure good performance. During the curation, we have found some heuristics that can improve the performance. We list them below.

Conflicts between short answers and long answers. Many datasets we collected are domain-specific, and some of them are classification tasks like NLVR2 and DreamSim . For these datasets, we convert them into multiple-choice QA format to avoid conflicts between repeated short-answers and some long answers from other datasets.

Addition of image denotation in the text for each question. We manually add image denotations like "the first image", "image 2", "the right image", etc. to each question in our dataset to make sure the question involved with multiple images is clear to answer.

Positions of the "" placeholder. For each item with multiple images, we write some rules to randomly put the image placeholders at the beginning of the first question or the end of the first question. This technique is also used by the curation LLaVA-665k .

A.3 Training Details

For Mantis-CLIP and Mantis-SigLIP, we first pre-train the multimodal projector on LLaVA pre-train data, where the learning rate is set to 1e-3, the batch size is set to 256. After pre-training, we fine-tune the 2 models on the Mantis-Instruct. For Mantis-Flamingo and Mantis-Idefics2, we directly fine-tune them on Mantis-Instruct since they have already been pre-trained on millions of data.

During the fine-tuning, we train each model on the data for 1 epoch, with a batch size of 128. The maximum context length is set to 8192. The learning rate is all set to 1e-5, except for Idefics2, where the learning rate is set to 5e-6 to better preserve its original knowledge. We set the warmup ratio to be 0.03 and use a cosine learning rate scheduler. We speed up our training and inference with Flash-attention2 . We use DeepSpeed Zero-3 for full fine-tuning. We apply QLoRA along with DoRA to more efficiently do comprehensive ablation studies under limited resources. All full fine-tuning ran on 16 A100 GPUs while the ablation study ran on 8 A100 GPUs.

A.4 Prompt Template