MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Jingkang Yang, Chunyuan Li, Ziwei Liu

Introduction

The recent advancements in artificial intelligence have focused on conversational assistants that possess a strong ability to understand user intentions and then execute actions . In addition to the strong generalization ability of large language models (LLMs), the notable achievements of these conversational assistants can be attributed to the practice of instruction tuning . It involves fine-tuning LLMs on a range of tasks specified through diverse and high-quality instructions . By incorporating instruction tuning, LLMs acquire a heightened comprehension of user intentions , enabling them to exhibit improved zero-shot capabilities even in previously unseen tasks . One potential reason for the zero-shot performance gain by instruction tuning is that it internalizes the context , which is preferred in user interactions especially when user input skips commonsense context.

Conversational assistants that excel in language tasks have achieved remarkable success. However, an optimal conversational assistant should be able to address tasks involving multiple modalities. This requires access to a diverse and high-quality multi-modal instruction-following dataset. The LLaVA-Instruct-150K dataset , also known as LLaVA, is the pioneering vision-language instruction-following dataset. It is constructed using COCO images, instructions and responses obtained from GPT-4 based on image captions and object bounding boxes.

Although inspiring, LLaVA-Instruct-150K exhibits three limitations. (1) Limited visual diversity: The dataset’s visual diversity is constrained due to its exclusive reliance on the COCO image. (2) Single image as visual data: it utilizes a single image as visual data, while a multi-modal conversational assistant should possess the capability to process multiple images or even extensive videos. For instance, it should effectively provide answers when a user presents a collection of images (or a sequence of images, such as a video) alongside the instruction: "Help me think of an album title for these images." (3) Language-only in-context information: it depends solely on language for in-context information, whereas a multi-modal conversational assistant should integrate multi-modal in-context information to better comprehend user instructions. For example, an assistant could more accurately align its description of an image with the tone, style, or other aspects if the human user provides a concrete image example of the desired attributes.

Addressing these limitations, we introduce MultI-Modal In-Context Instruction Tuning (MIMIC-IT). MIMIC-IT is characterized by: (1) Diverse visual scenes, incorporating images and videos from general scenes, egocentric view scenes, and indoor RGB-D images across various datasets. (2) Multiple images (or a video) as visual data, supporting instruction-response pairs accompanied by any number of images or videos. (3) Multi-modal in-context information, featuring in-context information formulated in multi-modal formats, including multiple instruction-response pairs and multiple images or videos (see Fig. 2 for data format clarification). To efficiently generate instruction-response pairs, we introduce Sythus, an automated pipeline for instruction-response annotation inspired by the self-instruct method . Sythus employs system message, visual annotation, and in-context examples to direct the language model (GPT-4 or ChatGPT) in generating instruction-response pairs based on visual context, including timestamps, captions, and object information, targeting three fundamental capabilities of vision-language models: perception, reasoning, and planning (refer to Fig. 1). Additionally, instructions and responses are translated from English into seven languages to support multi-lingual usage.

On MIMIC-IT, we train a multi-modal model Otter based on OpenFlamingo . We evaluate Otter’s multi-modal capabilities in two aspects: (1) ChatGPT evaluation on the MMAGIBenchmark , comparing Otter’s perception and reasoning abilities with other recent vision-language models (VLMs), where Otter demonstrates the strongest performance. (2) Human evaluation on the Multi-Modality Arena , where Otter outperforms other VLMs, achieving the highest Elo rating. Furthermore, we assess Otter’s few-shot in-context learning ability using the COCO Caption dataset , with results showing Otter’s superior performance over OpenFlamingo in all few-shot settings. In summary, our contributions include:

MultI-Modal In-Context Instruction Tuning (MIMIC-IT) dataset, a dataset comprising ~ 2.8M multi-modal in-context instruction-response pairs, with 2.2 million unique instructions, across various real-life scenes.

Syphus, an automatic pipeline built with LLMs to generate high-quality and multi-lingual instruction-response pairs based on visual context.

Otter, a multi-modal model demonstrates robust multi-modal perception and reasoning capabilities, effectively following human intent while exhibiting adeptness in-context learning.

Related Work

The notion of instruction tuning in multi-modal models was initially introduced in the work called Multi-Instruct , which encompassed a wide range of multi-modal tasks involving visual understanding and multi-modal reasoning, such as Visual Question Answering . Similarly, Mini-GPT4 created its instruction-based dataset by merging Conceptual Caption , SBU , and LAION with handwritten instruction templates. More recently, LLaVA-Instruct-150K has elevated the quality of instruction tuning datasets by utilizing self-instruct and GPT-4 , along with handwritten seed instructions on COCO images . While these previous works on multi-modal instruction tuning primarily focused on general scene images, our approach categorizes our data sources into indoor scenes, outdoor scenes, conversations, and egocentric videos. Additionally, drawing inspiration from the image-text interleaved structure of the MMC4 dataset , our approach further distinguishes itself by incorporating a multi-modal in-context format into instruction tuning.

2 Multi-modal Foundation Models

With the recent success of ChatGPT , GPT-4 , and other LLMs , recent studies start to explore incorporating information from other modalities into pretrained language models. These studies extend the capabilities of LLM to more tasks and modalities and can be categorized into two classes: (i) Multi-model Aggregation. These approaches take an LLM as a dispatch scheduler and connect different expert models through it to allow for different tasks. Language serves as an interface to call expert visual-language models within their respective task domains. However, this approach is limited that each model cannot be trained individually on new tasks. (ii) End-to-End Trainable Models. These approaches connect models from different modalities into integrated end-to-end trainable models, also known as multi-modal foundation models. Among them, based on large-scale image-text interleaved pretrained model OpenFlamingo , Otter is the first open-sourced model to further demonstrate the power of multi-modal in-context instruction tuning.

Multi-modal In-context Instruction Tuning Dataset

We aim to build MIMIC-IT dataset to support more VLMs in acquiring the ability to comprehend the real world. In this section, we provide an overview of the MIMIC-IT dataset, starting with the data format in Sec. 3.1 and our automatic instruction generation pipeline, Sythus, in Sec. 3.2.

Each instance ii in the MIMIC-IT dataset comprises an instruction-response pair and a set of NN images. We regard it as query example with a tuple: (Iq,Rq,Xq)(I_{q},R_{q},X_{q}), where {xj=1N}∈Xq\left\{x_{j=1}^{N}\right\}\in X_{q}. Here, IqI_{q} denotes the qq-th instruction in our dataset, RqR_{q} represents the response, and XqX_{q} refers to the images or videos Videos can be viewed as ordered sequences of images.. Our primary objective is to develop a visual language model pθ(Rq∣(Iq,Xq))p_{\theta}(R_{q}\mid(I_{q},X_{q})) parametrized by trainable parameters θ\theta, the model generates the response RiR_{i} for each query (Iq,Xq)(I_{q},X_{q}). With above example denotes the standard instruction tuning process of a visual language model. Further, we could define a set of in-context examples as (Ik,Rk,Xk)k=1M(I_{k},R_{k},X_{k})_{k=1}^{M}, where MM is the number of the set.

We then define a context function Cψ:(Iq,Xq)↦{(Ik,Xk)}k=1MC_{\psi}:(I_{q},X_{q})\mapsto\{(I_{k},X_{k})\}_{k=1}^{M} to represent the in-context examples with current query example. In summary, all data in the MIMIC-IT dataset will be represented in the following format, query example with its corresponding in-context examples.

Now the visual language model that incorporates in-context examples can be denoted as pθ(Rq∣(Iq,Xq,Cψ(Iq,Xq)))p_{\theta}(R_{q}\mid(I_{q},X_{q},C_{\psi}(I_{q},X_{q}))). CψC_{\psi} is task-dependent, we apply different approaches to organize the in-context examples with the current query example. The details will be presented in Sec. 3.3 and illustrative examples will be showcased in Fig. 2.

2 Sythus: Automatic Instruction-Response Generation Pipeline

We present Sythus (see Figure 3), an automated pipeline for generating high-quality instruction-response pairs in multiple languages. Building upon the framework proposed by LLaVA , we utilize ChatGPT to generate instruction-response pairs based on visual content. To ensure the quality of the generated instruction-response pairs, our pipeline incorporates system messages, visual annotations, and in-context examples as prompts for ChatGPT. System messages define the desired tone and style of the generated instruction-response pairs, while visual annotations provide essential image information such as bounding boxes and image descriptions. In-context examples assist ChatGPT in learning within the context. Since the quality of coreset impacts subsequent data collection process , we employ a cold-start strategy to enhance in-context examples before the large-scale query. During the cold-start stage, in-context examples are collected by prompting ChatGPT solely through system messages and visual annotations, employing a heuristic approach. This stage concludes only when satisfactory in-context examples are identified. In step 4, once the instruction-response pairs are obtained, the pipeline expands them into Chinese (zh), Japanese (ja), Spanish (es), German (de), French (fr), Korean (ko), and Arabic (ar). For further details, please refer to Appendix C, and task-specific prompts can be found in Appendix D.

3 Visual Data Exploration

Acknowledging the importance of high-quality visual annotations and the need for diverse vision-language instructions that align with the distribution of real-world visual content, we curate a collection of seven image and video datasets spanning a wide spectrum of scenes, from general to specific. Encompassing various topics, the MIMIC-IT dataset includes general scene understanding and reasoning, spoting general and subtle differences, as well as facilitating egocentric view comprehension to assist VLMs in future AR headsets, etc. In the subsequent sections, we will present the application scenarios of our dataset: General Scene Understanding in Sec. 3.3.1 and General Scene Understanding in Sec. 3.3.2. In each sub-task, we elaborate on the process of organizing various data into an in-context instruction tuning format, based on the previously established guidelines.

For understanding the general scenes, we include four tasks: (1) LLaVA-Interleaved. (2) Spot The Difference. (3) Visual Story Telling. (4) Dense Captions.

LLaVA-Interleaved (LA-I). Learning with in-context examples is essential for effective instruction tuning. To achieve this, we refine the LLaVA-Instruct-150K dataset by retrieving ten in-context examples for each instruction-response pair in LLaVA-Instruct-150K, building LLaVA-Interleaved (LA-I). We identify each data’s in-context examples based on instruction text-to-text similarity or image-image similarity. Further details on locating in-context examples and the data sources for LA-I can be found in the Appendix.

Spot The Difference (SD). Learning to discern differences between images is vital for understanding real-world changes. Our study encompasses two interrelated task types in Scene Difference (SD), addressing varying complexity levels in difference identification. The first type, General Scene Difference, involves creating a pair of images by determining the most similar one to the current image, utilizing image-to-image similarity relationships from the COCO2017 . The second type, Subtle Difference, features pairs of similar images with subtle distinctions sourced from the Spot-the-Diff, extracted from surveillance footage. For the first type, we prompt ChatGPT using original image captions and object detection annotations, while for the second type, we employ natural language difference descriptions as annotations. The resulting instruction-response pairs focus on identifying differences between the paired images.

Visual Story Telling (VIST). Beyond traditional scene understanding, the ability to generate coherent and engaging narratives based on visual input expands the context comprehension of Visual Language Models (VLMs). To enable this, we propose a task using the Visual Storytelling datase , which includes event-based image sequences and corresponding inquiry questions. Given that image annotations often contain narratives and timelines not directly observable, we instruct ChatGPT to act as a viewer answering questions about the images. The prompts also incorporate thought-provoking inquiries to promote creativity. Each task instance comprises multiple images and instruction-response pairs, providing in-context examples.

Dense Captions (DC). Expanding the scope of video understanding, DC features dense captions from corresponding to clips within longer videos. The instructions pose a diverse set of questions, addressing the general visual content of the video, human actions, and behaviors, the chronological sequence of events, and causal relationships. This approach encourages VLMs to delve deeper into the intricacies of video content.

TV Show Captions (TVC). The primary purpose of incorporating TV show clips with high-level captions into the training process of VLMs is to enhance their social reasoning abilities and deepen their understanding of complex character dynamics. By organizing drama clips from to analyze character relationships and motivations, we aim to challenge VLMs to move beyond mere perception and demonstrate their reasoning capabilities within the context of TV show narratives. This focused approach is crucial for fostering advanced VLMs capable of effectively handling diverse real-world situations and user queries.

3.2 Egocentric View Understanding

Indoor Event Planning (IEP). Emphasizing the planning capabilities of virtual assistants, we utilize visual inputs consisting of a collection of 2D photos depicting a room. We gather indoor scene RGB-D images from ScanNetv2 and sample them into multiple 2D visual inputs, representing a room’s layout from a first-person perspective. We prompt ChatGPT to generate instructions that direct humans to perform various activities in indoor spaces. Initially, we have ChatGPT create a personality for the room owner. Subsequently, the planning should be intimately related to the room’s layout and the generated room owner, underlining the importance of context awareness in VLMs. This approach ensures that models can effectively support users across diverse indoor scenarios.

Ego4D (E4D) . Utilizing E4D’s egocentric videos, we strive to enable VLMs to function effectively as augmented reality (AR) assistants in real-life scenarios. By prompting ChatGPT to generate instructions based on visual descriptions, our goal is to simulate practical interactions between users and AR assistants. To this end, we devise assistant-related questions and tasks that demand context-aware responses. For instance, Instruction: What should I do now? Response: Based on my observation, you can now proceed to do…. This focused approach underscores the potential of VLMs in providing valuable insights and assistance across a diverse range of daily life situations.

4 Dataset Statistics

Table 1 presents the essential statistics pertaining to the generated data. Our dataset comprises over 2.8 million instruction-response pairs, wherein each pair includes at least one multi-modal in-context example and one language-only in-context example. Among these pairs, there are 2.2M unique instructions. Furthermore, to examine the characteristics and diversity of the instructions (refer to Fig. 4 (a)) and responses (refer to Fig. 4 (b)), we analyze the verb-noun structure present in them, refering to . Specifically, we employ spaCy for parsing the instructions, extracting the verb closest to the root, and retrieving its first direct noun objecthttps://github.com/explosion/spacy-models/releases/tag/en_core_web_md-3.5.0. We plot the top 20 most frequently occurring root verbs alongside their top 4 direct noun objects. Our findings reveal that the sentence structure of responses exhibits greater diversity compared to that of instructions. Moreover, we demonstrate diversity in terms of the length of instructions/responses, the number of images per instruction, and the number of in-context examples per instruction, as depicted in Fig. 4 (c).

Empricial Evaluation

In this section, we showcase the diverse applications of the MIMIC-IT dataset and the potential capabilities of a vision-language model (VLM) trained on it. Firstly, in Sec. 4.1, we introduce Otter, an in-context instruction-tuned model developed using the MIMIC-IT dataset. Next, in Sec. 4.2, we explore various methods for training Otter on the MIMIC-IT dataset and discuss numerous scenarios in which Otter can be effectively employed. Finally, in Sec. 4.3 to Sec. 4.5, we present a comparative analysis of Otter’s performance against other VLMs across an array of benchmarks.

Otter is designed to support multi-modal in-context instruction tuning based on the OpenFlamingo model, which involves conditioning the language model on the corresponding media, such as an image that corresponds to a caption or an instruction-response pair.

2 Usage Examples and Demonstrations

Scene Understanding and Reasoning. The MIMIC-IT dataset comprises approximately 2.8 million in-context instruction-response pairs, which are structured into a cohesive template to facilitate various tasks. The following template encompasses images, user instructions, and model-generated responses, utilizing the Human and Assistant role labels to enable seamless user-assistant interactions.

Training the Otter model on the MIMIC-IT dataset allows it to acquire different capacities, as demonstrated by the LA and SD tasks. Trained on the LA task, the model exhibits exceptional scene comprehension, reasoning abilities, and multi-round conversation capabilities. Meanwhile, on the SD task, the model can acquire the ability to adeptly spot general differences or subtle distinctions within daily scenes.

We showcase response examples from the Otter after training on the MIMIC-IT dataset in Fig. 5, highlighting its ability to understand situations and reasoning in a multi-round conversation style.

Learning with In-context Examples. As mentioned in Sec. 3.1, regarding the concept of organizing visual-language in-context examples, we demonstrate here the acquired ability of the Otter model to follow inter-contextual instructions after training on the LA-T2T task (refer to Appx. for other tasks). The organized input data format is as follows:

The Otter model’s demonstration of regulating its expressions by referencing in-context examples is illustrated in Fig. 5.

Egocentric Visual Assistant. A distinctive feature of the MIMIC-IT dataset is its inclusion of a comprehensive collection of videos and sequential images in an egocentric view, derived from the IEP, E4D scenarios. In the IEP scenario, the content emphasizes understanding and planning within indoor environments, incorporating instructions and responses designed to guide the model in event planning based on interior layouts.

The E4D scenario, on the other hand, tailors instructions and responses specifically for first-person augmented reality (AR) headset assistant applications. These two datasets collectively serve to bolster the model’s proficiency in perceiving scenes from a first-person viewpoint, strategizing for impending tasks, and providing valuable insights and suggestions to AR headset users. Tailored this part of data, we train an egocentric visual assistant, termed Otter-E, which is specifically designed for AR headset applications. MIMIC-IT bolsters the model’s proficiency in perceiving scenes from a first-person viewpoint, strategizing for impending tasks, and providing valuable insights and suggestions to AR headset users. As a result, the Otter-E model emerges as an exceptional and visionary Visual Language Model for AR headsets, paving the way for a groundbreaking and immersive experience.

In the bottom image of Fig. 5, Otter-E demonstrates its ability to perceive the first-person view and respond to users’ questions, such as guiding users to land a small aircraft (In real-life scenarios, you are not encouraged to consult visual assistants for such hazardous actions).

3 ChatGPT Evaluation

In Tab. 2, we utilize the MMAGIBench framework to provide an extensive evaluation of the perception and reasoning capabilities of vision-language models. The perception benchmark consists of data derived from COCO images and social network images (e.g., , Twitter), covering tasks such as coarse scene and object recognition, fine-grained OCR, celebrity identification, and recognition of well-known locations. The reasoning benchmark, on the other hand, is performed across three dimensions: attribute reasoning, relation reasoning, and future prediction.

Current evaluation metrics for vision-language models, like VQAv2 , exhibit shortcomings in terms of robustness. For instance, VQAv2 primarily assesses single-word or phrase responses, while many modern models generate sentence outputs. To bridge this gap, we evaluate the models by asking ChatGPT to compare their label predictions with the ground truth labels for each input. A test sample is considered correct if ChatGPT’s response indicates that the prediction aligns with the corresponding label. For a more in-depth understanding of MMAGIBench, we recommend referring to the original source . Fig. 6 (a) demonstrates that Otter outperforms VideoChatGPT by 6.8% accuracy and 1.8% on MSVD 0-shot question answering and captioning benchmarks respectively. Similar substantial margins are also observed on the MSRVTT dataset.

4 Human Evaluation

Multi-Modality Arena uses an Elo rating system to evaluate the usefulness and alignment of VLM responses. The Elo rating system calculates the relative skill levels of players, as commonly used in chess and other competitive games. The difference in Elo ratings between the two models predicts the outcome if they were matched against each other. This system works well for evaluating conversational AI models, because multiple models can have pairwise "battles" responding to the same inputs in a user-blind evaluation. Fig. 6(b) shows that Otter demonstrates superior usefulness and alignment, achieving the highest Elo rating among recent VLMs.

5 Few-shot In-context Learning Metric Evaluation

Otter is finetuned based on OpenFlamingo, an architecture designed for multi-modal in-context learning. Finetuned with the MIMIC-IT dataset, Otter outperforms OpenFlamingo by a substantial margin on COCO caption (CIDEr) few-shot evaluation (see Fig. 6(c)). As expected, the finetuning also brings marginal performance gain on zero-shot evaluation.

Discussion

Limitations. Though we have iteratively refined the system message and instruction-response examples, ChatGPT is prone to language hallucinations therefore it might generate incorrect responses. Generally, more trustworthy language models are desired for self-instruct data generation.

Future Works. In the future, we plan to support more embodied AI datasets such as Language-Table and SayCan . We also consider improving the instruction collection with more trustworthy language models or generation techniques.

Conclusion. In this work, we propose MIMIC-IT, a large-scale multi-modal in-context instruction tuning dataset. We leverage an automatic pipeline, Syphus, to enable this dataset to cover a diverse set of visual scenes and creative instructions in eight languages. MIMIC-IT empowers our model, Otter, to achieve state-of-the-art performances in perception and reasoning benchmarks as well as human evaluations.

Acknowledgments and Disclosure of Funding

This study is supported by the Ministry of Education, Singapore, under its MOE AcRF Tier 2 (MOE-T2EP20221- 0012), NTU NAP, and under the RIE2020 Industry Alignment Fund – Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s). We thank Peiyu Fu, Xuli Chen, and Mehdi Cherti for their professional advice on the in-context example of the translation query of Japanese, French, German, Spanish, Korean, and Arabic.

References

Appendix A Total Cost and ChatGPT Version

We construct MIMIC-IT using the ChatGPT-0301 version. Overall, we query 1,006,746,240 tokens (859,677,150 and 147,069,090 for input and output tokens respectively). The estimated total cost is $20134.9248. https://openai.com/pricing

Appendix B Content Copyright and License

The license of the datasets we used in this work is illustrated below.

Appendix C Sythus: Automatic Instruction Generation Pipeline

Since we use GPT to generate instructions and responses, we generally follow the GPT content policy for safe and ethical use. This policy eliminates output that is suspicious for unfair opportunities, stereotyping, overrepresentation/underrepresentation, explicit content, disinformation, or unreliable information.

We enrich the datasets by translating the English instruction-response pairs by GPT into 7 additional languages: Chinese, Japanese, Spanish, German, French, Korean, and Arabic. See the prompt for multi-lingual translation query in Fig. 7.

Appendix D Annotation Prompt

In this section, we will present prompts for querying ChatGPT of all datasets in detail. Each prompt contains system message, in-context emample.