LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, Chunyuan Li

Introduction

Recent advancements in Large Multimodal Models (LMMs) have showcased impressive capabilities in diverse multimodal contexts, advancing the pursuit of artificial general intelligence. With extensive vision-language data , they empower Large Language Models (LLMs) with visual modality by aligning vision encoders . This integration has propelled forward the field of AI, enabling complex image and language understanding tasks to be performed with unprecedented accuracy.

However, most open-source LMMs have primarily focused on pushing the performance limit of the single-image scenario, the more complex multi-image scenarios remain largely less explored. This oversight is significant given that many real-world applications demand multi-image capabilities, such as comprehensive multi-image analyses. Traditionally, researchers have approached these challenges by training separate task-specific models for each application scenario, e.g., multi-image , video , and 3D . This is both labor-intensive and time-consuming, resulting in fragmented methodologies that are inefficient and often unscalable. Considering the diverse range of computer vision settings and data formats, there is a pressing need to develop a general framework for LMMs that can operate effectively across these varied contexts.

In this paper, we observe that the image-text interleaved format can naturally serve as a general data template to unify different scenarios, e.g., single-image or multi-image as special cases, video as multi-frames, and 3D as multi-views, as illustrated in Figure 1. Therefore, we present LLaVA-NeXT-Interleave, an all-around LMM that extends the model capabilities to various real-world settings such as Multi-image, Multi-frame (videos), Multi-view (3D) while maintains the performance of the Multi-patch (single-image) performance. We denote the four settings as M4.

The core innovation of our approach lies in the perspective to leverage an image-text interleaved format as a universal data template capable of accommodating different scenarios, and construct the related visual instruction-following data. This perspective not only simplifies the training process across various domains, but also allow the model to emerge new capabilities due to cross-domain task composition.

Our contributions are summarized as below:

Interleave data format unifies different tasks. We convert multi-image, video, 3D, and single-image data all into an interleaved training format, which unifies different tasks in a single LMM.

New dataset and benchmark. We compile a high-quality training dataset, M4-Instruct, with 1177.6 samples to empower LMMs with the M4 capabilities, which spans 4 primary domains (multi-image, video, 3D, and single-image) with 14 tasks and 41 datasets. We also curate LLaVA-Interleave Bench, a diverse set of benchmarks to evaluate the multi-image performance, including 7 newly collected and 13 existing in/out-domain benchmarks.

SoTA performance. With a single model, LLaVA-NeXT-Interleave can achieve leading results across different multi-image tasks compared to the previous SoTA, while maintaining the single-image performance, as exemplified in Figure LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Emerging capabilities with cross-task transfer. By jointly training on a diverse set of tasks, our model showcases emerging capabilities to transfer tasks across different settings and modalities. e.g., from spotting differences between images to videos.

Related Work

As a more general format, interleaved image-text data can enable LMMs with two distinctive capabilities: multimodal in-context learning (ICL) capability and instruction-following capability in real-world multi-image application scenarios. The former in-context scenarios interleave several image-text examples within the prompt as task demonstrations, adapting LMMs to new tasks in the inference stage in a few-shot manner. Flamingo is first model to demonstrate this capability, and thus is considered as GPT-3 moment for multimodal community. Typically, the multimodal ICL ability is emerged after pre-training on web-scale raw interleaved image-text sequences. In the open-source community, MMC4 introduces a public 101.2M interleaved dataset spanning everyday topics, OBELICS also presents a filtered dataset comprising 141M interleaved web pages. Kosmos-1 curates a 71M multimodal corpora, including arbitrarily interleaved documents. To explicitly enable the ICL capability, MIMIC-IT proposes an automatic pipeline to create 2.8M multimodal samples in the instruction-tuning stage. On the other hand, the latter multi-image scenarios aim to tackle diverse real-world applications scenarios that involve multi-images. The training data of VPG-C collected 4 new datasets with ChatGPT. Mantis-Instruct compiles existing 11 interleaved datasets and creates 4 new datasets. The proposed M4-Instruct compiles existing 41 interleaved datasets and creates 6 new datasets, covering a much higher scenarios diversity than Mantis-Instruct.

Interleaved LMMs.

As representative closed-source LMMs, both GPT-4V and Gemini support real-world multi-image application scenarios with leading performance. With various public datasets aforementioned, the community has developed open-source LMMs equipped with remarkable multi-image proficiency. The ICL performance is typically considered to evaluate multimodal pre-training, which has been adopted in several known LMMs, such as OpenFlamingo , IDEFICS series , VILA and MM1 , Emu2 . Otter is initialized from OpenFlamingo, and is fine-tuned on the MIMIC-IT dataset to further improve ICL ability with instruction-tuning. In contrast, the use of instruction-tuning in LMMs for various real-world multi-image applications has been less explored, despite of Mantis . The proposed LLaVA-NeXT-Interleave not only broadens the multi-image scenario itself as demonstrated by the improved experimental results, but also generalize the settings to diverse scenarios with one model, e.g., video, 3D, and single-image. The cross-scenario training leads to emerging capabilities, achieving zero-shot task composition in new multi-image contexts.

Interleaved Benchmarks.

To assess the interleaved multi-image capabilities of LMMs, there have been several high-quality benchmarks in various scenarios. The ICL benchmarks for LMMs comprehensively evaluate their interleaved skills from few-shot to many-shot settings. For the more challenging multi-image scenarios, previous works mainly focus on a specific domain for evaluation, including NLVR2 for daily-life VQA, MMMU for colleague-level problem-solving, MathVerse-mv and SciVerse-mv for mathematical and scientific reasoning, BLINK to challenge LMMs, and Mantis-Eval for multi-image understanding. To further evaluate LMMs on a collection of multi-image scenarios, DEMON is the first benchmark that compiles dozens of datasets with 477K samples. With the large amount of data and high diversity, DEMON lays a good foundation for multi-image research. Unfortunately, it also inherits a significant amount of low-quality data samples from existing datasets. To facilitate evaluation, the proposed LLaVA-Interleave Bench curate high-quality samples, comprising both specific (synthetic, mathematical, low-level) and general (daily, real-world, text-rich) multi-image scenarios. With 9 newly curated and 13 existing datasets, we categorize them into in-domain (12.9K) and out-domain (4.1K) schemes. Con-current multi-image evaluation benchmarks include MuirBench and ReMI .

Interleaved Multi-image Tasks & Data

We observe different computer vision scenarios can be generally represented by the interleaved multi-image format, such as video, 3D, and single-image data. Therefore, to endow LLaVA-Interleave with diverse capabilities, as shown in Figure 1, we adopt the interleaved multi-image format to unify the data input of the following four tasks:

Multi-image scenarios include visual instructions incorporating interleaved vision-language input with multiple images. This setting covers 12 challenging real-world tasks included in our training data, such as spotting the difference, visual story telling, image editing instruction generation, interleaved multi-image dialogue, multi-image puzzle, low-level multi-image assessment, etc.

Multi-frame scenarios refer to taking video as input data by sampling it into multiple frames, preserving temporal visual cues across the multi-image sequence. We mainly focus on 2 tasks: video detailed captioning and video VQA.

Multi-view scenarios depict 3D environments by multi-view images from different perspectives, where the visual correspondence and disparity can indicate spatial information in the 3D world. For 3D perception, we include 2 tasks: embodied VQA (dialogue and planning), and 3D scene VQA (captioning and grounding).

Multi-patch scenarios represent the conventional single-image tasks. With the design of ‘any resolution’ in LLaVA-NeXT , we divide a high-resolution image into multiple low-resolution patches for efficient visual encoding, compatible with our interleaved multi-image format.

2 M4-Instruct

To empower all-round multi-image capabilities, we meticulously curate a comprehensive training dataset including 1177.6K instances, termed M4-Instruct, widely spanning multi-image, multi-frame, and multi-view scenarios with 14 tasks and 41 datasets, along with multi-patch data to preserve basic single-image performance. We showcase task examples of the first three scenarios in Figure 2.

We exhibit a data overview of M4-Instruct in Figure 3, and the detailed data statistics in Table 15. For the multi-image data, most of the datasets are collected from previous public efforts and rigorously converted into our unified format with task-specific instructions, some inspired by DEMON and Mantis . On top of that, we also utilize GPT-4V to annotate 3 new tasks to enable more diverse capabilities, i.e., Real-world Difference, Synthetic Difference, and Twitter Post. For the video data, we collect a 255K subset from LLaVA-Hound , including 240K video VQA and 15K video detailed captioning. We also include NExT-QA and STAR to expand our video training data. For the 3D data, we widely gather the training set from nuScenes QA , ALFRED , ScanQA , and 3D-LLM , covering both outdoor and indoor scenarios. For the single-image data, we randomly sample 40% of the stage-2 fine-tuning data from LLaVA-NeXT , which aims to preserve the single-image capacity.

To comprehensively evaluate the interleaved multi-image performance, we introduce the LLaVA-Interleave Bench for LMMs, consisting of 13 challenging tasks with 17K instances. We present a data overview of the benchmark in Figure 2, and the detailed data statistics in Table 16. In detail, we categorize multi-image tasks into two classes:

In-domain Evaluation includes tasks that have been ‘seen’ during our training, designed to verify the model performance within familiar scenarios. We adopt 5 newly curated multi-image tasks corresponding to training datasets, and 2 existing benchmarks, Q-Bench and NLVR2 , with 12.9K in total.

Out-domain Evaluation involves tasks that don’t overlap with training scenarios, aiming to reveal the generalization capacity of LMMs. We construct 2 new tasks for multi-image mathematical (MathVerse ) and scientific (SciVerse ) comprehension, and utilize 3 existing benchmarks, Mantis-Eval , BLINK , and MMMU , with 4.1K in total.

Interleaved Visual Instruction Tuning

In this section, we introduce several key techniques during the interleaved visual instruction tuning of LLaVA-NeXT-Interleave. For architecture designs, we follow LLaVA-NeXT to adopt the most general framework, i.e., a vision encoder , an intermediate projector, and a powerful LLM . Then, we consider the following three techniques to achieve improved multi-image performance.

The interleaved multi-image tasks can be regarded as an extension of single-image scenarios, more flexible in formats and challenging in reasoning. Therefore, to better leverage the pre-trained single-image proficiency, we adopt an off-the-shelf LLaVA-NeXT-Image as the base model, which has gone through a stage-1 image-caption pre-training and a stage-2 single-image fine-tuning. On top of this model, we perform the interleaved multi-image instruction tuning with our M4-Instruct dataset.

Technique 2: Mixed Interleaved data formats during training.

We adopt two format choices for the positions of image tokens during the interleaved multi-image training. The first is to place all the image tokens in front of the prompt, while maintaining the placeholders ⟨\langleimage⟩\rangle within the text, denoted as ‘In-the-front format’. The second preserves the interleaved format to put image tokens in the place they are originally in, i.e., the positions of ⟨\langleimage⟩\rangle, denoted as ‘interleaved format’. In this way, LLaVA-NeXT-Interleave supports more flexible inference modes, exhibiting robustness to different input formats.

Technique 3: Combining different data scenarios improves individual task performance.

Most existing works conduct supervised fine-tuning with only one type of data source, e.g., multi-image tuning of Mantis and multi-frame tuning of LLaMA-VID . Instead, we utilize the M4-Instruct to simultaneously conduct instruction tuning with four different tasks (multi-image/frame/view/patch). With a unified interleaved format, distinct data scenarios have the potential to provide complementary semantics and instruction-following capabilities.

Experiments

In Section 5.1, we first introduce our evaluation schemes and implementation details. Then, in Section 5.2, we report and analyze the quantitative results in four interleaved multi-image scenarios.

We evaluate our LLaVA-NeXT-Interleave model on four real-world interleaved scenarios, i.e., multi-image, multi-frame (video), multi-view (3D), and multi-patch (single-image).

For multi-image evaluation, we adopt the proposed LLaVA-Interleave Bench covering comprehensive in-domain and out-domain tasks.

For video evaluation, we utilize the existing NExT-QA , MVBench , Video Detailed Description (VDD) , and ActivityNet-QA (Act) . For ActivityNet-QA, we present both the accuracy and GPT score (Acc/Score). We also evaluate on VideoChat-GPT (VCG) with five metrics: CI (Correctness of Information), DO (Detail Orientation), CU (Context Understanding), TU (Temporal Understanding), and CO (Consistency).

For 3D evaluation, we select ScanQA , two tasks from 3D-LLM , i.e., 3D-assisted Dialogue and Task Decomposition, and also curate two new test set from nuScenes VQA and ALFRED .

Implementation Details.

Following the same architecture in LLaVA-NeXT , our LLaVA-NeXT-Interleave adopts Qwen 1.5 as the base LLM with 0.5B, 7B and 14B parameters, SigLIP-400M with 384×\times384 resolutions as the vision encoder, and a two-layer MLP as the projection layer.

2 Main Results

As reported in Table 1, the average multi-image performance of LLaVA-NeXT-Interleave surpasses previous open-source models in both in- and out-domain benchmarks. For in-domain evaluation, our model demonstrates significant advantages across various tasks as expected, due to the multi-image instruction tuning with M4-Instruct. For out-domain evaluation, LLaVA-NeXT-Interleave also showcases superior generalization capacity within novel scenarios, e.g., comparable to GPT-4V on Mantis-Eval and BLINK.

Multi-frame (Video) Results.

Compared with previous video-based LMMs under similar model sizes, LLaVA-NeXT-Interleave achieves superior results on many benchmarks in Table 2, though not specifically designed for video tasks. We also follow LLaVA-Hound to add DPO training after our M4-Instruct tuning. After adding DPO, our 7B model attains SoTA performance on VDD and VideoChat-GPT benchmarks, surpassing the previous LLaVA-NeXT-Video (34B). This demonstrates the effective temporal understanding and reasoning capabilities of our model across sequential frames. Note that we calculate the average scores by multiplying a weight of 10 times by the score of Video Detailed Description and VideoChat-GPT.

Multi-view (3D) Results.

For 3D perception in Table 3, our model also obtains leading results for both indoor and outdoor scenarios on five in-domain benchmarks. Compared to 3D-LLM and Point-LLM with additional point clouds as input, LLaVA-NeXT-Interleave only accepts multi-view images to interpret the 3D world, attaining significantly higher scores in challenging 3D scenarios.

Multi-patch (single-image) Results.

We also add 307k (40%) of original LLaVA-NeXT single-image data, which makes our model capable of doing single-image tasks. We use the anyres training for single-image data, which divides an image into multiple patches, forming another multi-image setting. As shown in Table 4, we maintain the single-image performance of LLaVA-NeXT-Image. As single-image data is of high quality and diversity, adding single-image data also improves the instruction-following ability and enables task transfer from single-image to multi-image, which is demonstrated in Section 6.

3 Ablations of Proposed Techniques

We study the effectiveness of the three proposed training techniques in Section 4 as below.

In Table 5, we compare training strategies. It is seen that initialization from a good single-image model checkpoint (from Stage-2) can consistently enhance the interleaved multi-image performance, than directly from a Stage-1 model checkpoint.

In Table 5.1, our mixed-format training can benefit the results of both two input formats.

In Table 5.1, we progressively incorporate single-image and multi-image data upon the video data. The integration of more sources contributes to enhanced performance, compared with models from individual visual scenarios.

Emerging Capabilities

In this section, we show some example to demonstrate the emerging capabilities of our model. Emerging capabilities means the capabilities do not trained during training but demonstrated when inference. We mainly showcase the emerging capabilities from three aspects:

Task Transfer from Single-image to Multi-image: The capability to reason over one image and tell the funny part is initially observed in single-image models , and not included in our multi-image training. As shown in Table 8, our model is capable of analyzing the fun part within multiple images. This new task is probably emerged by the composition of the single-image capability and multi-image VQA training.

Task Transfer from Image to Video: We only include the multi-image Twitter post task in the M4-Instruct training, while our model can directly perform the witter post from a video, as shown in Table 9. This new task is probably composed by the training data of multi-image Twitter post and video VQA tasks.

Real-world Applications: In Tables 10, 11, and 12, we showcase three real-world scenarios that are not explicitly contained in our interleaved training data, which are multi-image painting style recognition, PPT summary & QA, and multi-doc VQA. This demonstrates our generalization potentials to a broader spectrum of applications.

More interesting demos can be found in our project pagehttps://llava-vl.github.io/blog/2024-06-16-llava-next-interleave/.

Conclusion

In conclusion, our research highlights the transformative potential of LLaVA-NeXT-Interleave in unifying and advancing the capabilities of Large Multimodal Models (LMMs) across diverse visual tasks. By leveraging the interleaved data format, we effectively integrate multi-image, video, 3D, and single-image scenarios, offering a cohesive approach to handling thwoese varied challenges. The introduction of the comprehensive M4-Instruct dataset and the LLaVA-Interleave Bench provides a solid foundation for training and evaluating LMMs, ensuring high-quality performance across multiple domains. Our extensive experiments validate that LLaVA-NeXT-Interleave not only sets new state-of-the-art benchmarks in multi-image tasks but also maintains exceptional performance in single-image tasks. Furthermore, the model exhibits promising emerging capabilities, such as cross-task transfer, showcasing its versatility and potential for broader applications. This work sets a new precedent in the field, paving the way for future advancements in multimodal AI and complex visual understanding tasks.

References

Appendix A Data Statistics

The detailed data statistics of M4-Instruct is shown in Table 15.

The detailed data statistics of LLaVA-Interleave Bench is shown in Table 16.

Appendix B Ablation Study

Similar to LLaVA-NEXT-Video, we adopt a ”Pooling to 1/4” strategy for which we pool the width and heighs of feature maps to 1/2 therefore reducing the number to totals to 1/4. We study the impact of image token pooling. We train and infer our model under two settings: pooling to 1/4 and not pooling with ShareGPTVideo-Caption+QA(255K) data. Pooling to a 1/4 setting is similar to LLaVA-NEXT-Video, which uses the pooling technique to trade-off between the number of image tokens and the number of frames. In our experiment, we find that not pooling yields better performance under similar #image tokens. During training, we sample 10 frames for videos. In this table, we also observe that adding more frames (from 10 to 16) during inference improves performance.

B.2 Impact of video DPO training on other tasks.

In Table 14, we compare the results of doing video DPO on other tasks. Though DPO significantly improves the video performance as shown in Table 2, it slightly impacts the performance of other tasks.