MLVU: Benchmarking Multi-task Long Video Understanding

Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Zhengyang Liang, Shitao Xiao, Minghao Qin, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, Zheng Liu

Introduction

Large language models (LLMs) are growing into a general solution for numerous AI tasks [6; 40]. In recent years, it becomes increasingly emphasized to extend LLMs with multi-modal capabilities and thus bring the Multi-modal LLM, namely, MLLM. Remarkably, it has been made possible for today’s MLLMs to perceive information in texts, images, videos, etc., and solve complicated problems in physical environments [1; 39]. Along with the development of MLLMs, new benchmarks are continuously created to facilitate comprehensive and in-depth analysis of MLLMs [51; 27; 9; 21].

However, it remains a great challenge to evaluate the MLLMs’ long-video understanding (LVU) performances given the following limitations. Firstly, the majority of existing video understanding benchmarks are made up of short videos [48; 21; 19; 32; 15], whose lengths can be merely a few seconds. As a result, they are insufficient to reflect the MLLMs’ long-video understanding capabilities. Secondly, there is a severe lack of diversity for the existing LVU benchmarks in terms of video genres and evaluation tasks. On one hand, the existing benchmarks are usually created based on one type of videos, e.g., movies [36; 22]. On the other hand, it may only contain one specific type of task in each individual benchmark, for example, the captioning task in , the temporal perception task in , and the action understanding task in [42; 11]. Due to the above limitations, existing benchmarks have difficulty in performing a comprehensive evaluation of LVU performances. Last but not least, many previous evaluation tasks are not properly designed for LVU, as they can be solved without using the complex information from long videos. For example, many questions are simply about one single frame in the long videos . Besides, numerous others are about popular movies and celebrities , which can be answered directly by MLLMs based on the textual prompts.

Conceptually, the MLLMs are expected to handle any type of long video and accomplish any type of related tasks. Therefore, the evaluation of LVU should emphasize two important properties: length and diversity. Based on such a principle, we propose a novel benchmark called MLVU (Mult-task Long Video Understanding Benchmark), which presents the following critical advantages.

It makes a substantial extension for the video length. MLVU is created based on long videos of diversified lengths, ranging from 3 minutes to 2 hours. The average video length is about 12 minutes, which makes it much longer than most of the existing benchmarks. Additionally, each video is further segmented so that evaluation tasks can be created w.r.t. different video clips (e.g., summarization for the first 3 minutes, the first 6 minutes, and the entire duration of the video). Therefore, it is able to flexibly evaluate the MLLMs’ performance across different video lengths.

It covers a wide variety of video genres. On one hand, it contains different types of real-world videos, such as movies, documentaries, surveillance videos, and ego-centric videos. On the other hand, it also includes typical simulated videos, like games and cartoons. As a result, it is able to comprehensively reflect the MLLMs’ performance across different application scenarios.

It introduces diversified evaluation tasks tailored for LVU. There are 9 different tasks in MLVU, which jointly examine a wide range of MLLMs’ key abilities, such as reasoning, captioning, recognition, summarization, etc. Both multi-choice and free-form generation tasks are included in MLVU, which reflect the MLLMs’ performances in handling different forms of tasks. Finally, some of the tasks are designed to leverage global information from entire videos, while others need to utilize proper local information from specific clips. Therefore, it can provide a unified perspective for the degree of completeness and nuance in understanding long videos.

We extensively investigate 20 popular MLLMs with MLVU, which brings in several critical insights. Firstly, long-video understanding remains a technically challenging problem for the existing MLLMs. While GPT-4ohttps://openai.com/index/hello-gpt-4o/ achieves the leading performance in the experiment, it only attains an average score of 64.6% in multi-choice tasks. All methods struggle with tasks requiring fine-grained information from entire videos, such as action counting, ordering, and summarization. Additionally, all methods experience significant performance drops as video lengths increase. Secondly, a significant gap persists between open-source and proprietary models. Although some of the latest open-source MLLMs can handle very long videos and excel at single-image tasks, they notably lag behind GPT-4o in understanding long videos. Finally, the empirical results underscore the influential factors in LVU, such as the extension of context length, the improvement of image understanding ability, and the utilization of strong LLM-backbones. In addition to the benchmark’s overall conclusion, the individual tasks enable the fine-grained analysis of MLLMs’ performances in each of the specialized aspects. Therefore, we anticipate the benchmark to assist in improving MLLMs’ long-video understanding capabilities by providing insights into their current strengths and weaknesses.

Related Work

Multimodal large language models (MLLMs) have gained wide interests from both academic and industrial sectors. Recent studies have made significant progresses in this field by integrating LLM backbones with visual encoders and adapters, and fine-tuning the entire architecture through visual instruction [25; 56; 7]. Based on the same philosophy, MLLMs have been further developed for video processing using video instruction datasets and specialized video adapters [53; 29; 20; 49; 23; 21]. However, most existing models are optimized for short videos, typically under one minute, due to the difficulty in establishing sufficient context for longer videos. To address this challenge, researchers have explored compact representations of videos. For example, LLaMa-Vid compresses each video frame into two tokens, allowing the model to handle videos several hours long. Methods like MovieChat and MA-LMM introduce specialized memory components for recursive video processing. Additionally, it is also explored to make selective usage of frames or clips from long videos based on retrievers or agents [47; 34; 41]. Despite these progresses, it remains an open problem for MLLMs to effectively handle long videos.

With the unprecedented interest in MLLMs, the creation of benchmarks for these models has become increasingly emphasized (as advanced by MMMU , MME , and many other pioneering works). In video understanding, the research community has made significant efforts as well, particularly for short videos. There are specialized benchmarks for temporal perception [50; 43; 28], action understanding [43; 42], video classification , video reasoning [46; 45], and video captioning [48; 31]. Recently, MVBench provides a comprehensive short-video benchmark to evaluate general capabilities via question-answering. For long video understanding, people seek to leverage long-form videos, like movies, to create benchmarks. For example, the LVU dataset presents movie understanding tasks, such as predicting release years and identifying character relationships. Similarly, LLaMA-Vid created a movie question-answering dataset based on MovieNet . Despite using long videos, many questions are about common-sense knowledge of the movies. As a result, they could be directly answered without using the video’s information. In contrast, MovieChat makes use of diversified videos and avoids specific character names or plot details in its questions. Considering that many of the questions target on exact time segments, it potentially degrades the tasks to short-video or image understanding problems. Beyond movies, there are other task-specific benchmarks like EgoSchema , which presents video reasoning tasks using first-person footage from Ego4D . However, these specialized benchmarks only focus on one single aspect of MLLMs, rather than offering a comprehensive analysis of long video understanding. Therefore, it remains essential to develop a comprehensive benchmark with carefully designed tasks to effectively evaluate MLLMs’ capabilities in understanding long videos.

MLVU: Multi-task Long Video Understanding Benchmark

In this section, we start with an overview of MLVU, which highlights its constitution and explains its values over the previous works. Then, we discuss how each evaluation task is constructed in MLVU.

MLVU is a comprehensive benchmark made up of 2593 evaluation tasksOnly consider the tasks for the entire videos. Others for fine-grained video segments are not accounted. (shown as the right side of Figure 1), which belong to 9 task categories tailored for long-video understanding. These categories include 1) Topic Reasoning, 2) Anomaly Recognition, 3) Video Summarization, 4) Needle Question Answering, 5) Ego Reasoning, 6) Plot Question Answering, 7) Sub-Scene Captioning, 8) Action Count, and 9) Action Order. The benchmark is distinguished by the following features.

MLVU is made up of videos of diversified lengths, spanning from 3 min to more than 2 hours (Figure 1 Left). The average length is about 12 min, which is much longer than the existing LVU benchmarks. Besides, each video is further partitioned as incremental segments, e.g., the first 3 min, the first 6 min, and the entire video, where tasks are created for each individual segment. Thus, the MLLMs can be flexibly evaluated across different video lengths.

MLVU offers a comprehensive collection of videos of different categories (Figure 1 Middle). There are typical real-world videos, which includes movies [52; 36], documentaries , TV series, egocentric videos , and surveillance footage . Meanwhile, there are also important simulated videos from animated series and game videos .

MLVU also provides a diversified collection of evaluation tasks, which are closely related to the common visual capabilities of MLLMs, such as reasoning, captioning, recognition, perception, and summarization (Figure 1 Right). All the tasks are tailored for LVU. That is to say, the tasks need to be solved based on the in-depth understanding of video. Some of tasks are to examine whether the global information from the entire video can be effectively utilized (holistic LVU); while others focus on whether the MLLMs can make precise usage of proper local information within the long video (detail LVU). Additionally, both multi-choice and free-form generation tasks are included in MLVU, which help to examine MLLMs’ capabilities in handling different task formats.

2 Construction of MLVU

The evaluation tasks of MLVU which can be categorized into three types: 1) holistic LVU, which needs to be solved by making use of the global information from the entire video; 2) single-detail LVU, which needs to leverage one critical plot within the long video; 3) multi-detail LVU, which calls for the joint utilization of multiple plots within the long video. The construction process of MLVU is discussed w.r.t. the above three categories. To facilitate the discussion, we define ULVC (Universal Long Video Collection) as the universal collection of miscellaneous long videos, which includes movies, documentaries, TV-series, ego-centric videos, surveillance footages, animated series, game videos, et al. (more details about ULVC is presented in the Appendix B).

Topic Reasoning (TR). The topic reasoning task requires MLLMs to respond to questions about the principal subject of a long video, as shown with Figure 2 (a). This includes elements such as the video’s genre, pivotal events, or primary settings. Videos for TR tasks are sourced from movies, documentaries, animated series, egocentric videos, and game videos from ULVC. All questions and answers undergo manual annotationDetailed information and annotation guidelines for annotators are presented in Appendix E, resulting in a total of 264 questions. TR tasks are formatted as multiple-choice questions, with the model’s performance assessed based on accuracy.

Anomaly Recognition (AR). The anomaly recognition task involves identifying the anomalous behavior within a surveillance footage (Figure 2 b). We leverage the surveillance video clips from UCF Crime dataset for this task. The selected video clips are longer than three minutes. We create 200 questions based on the original annotations provided by the dataset. The AR task is also conducted in the multiple-choice format, whose performance is measured by accuracy.

Video Summarization (VS). This task requires MLLMs to summarize the key events in a long video (Figure 2 c). We select the narrative-rich videos from ULVC for this task, including movies, TV series, documentaries, and animated series. There are 217 selected videos in total, whose summaries are manually annotated. During evaluation, the MLLMs are prompted with "Please summarize the main content of this video". We employ GPT-4 to assess the generated summaries by comparing with the annotation results. Details about annotation and evaluation are presented in Appendix E.3 and F.3.

2.2 Single-Detail LVU

Needle Question-Answering (NQA). Needle-In-the-Haystack-Search (NIHS) is a popular evaluation task for long-context LLM . Taking the inspiration from NIHS, we create Needle Question-Answering (NQA), shown as Figure 2 (d). In this task, the MLLM is required to answer a question related to a specific segment (referred as needle) within a long video (referred as background video). The needles are short video clips sampled from WebVid , while the background videos are sampled from our ULVC. The needle is randomly inserted into the background video, where a question-answer pair is annotated. By incorporating necessary details, the question can always correspond to the needle without ambiguity. During evaluation, the MLLM needs to infer the location of the needle based on the details provided in the question, and solve the problem on top of the needle’s information. The NQA task is structured as multiple-choice, whose performance is measured by accuracy.

Ego Reasoning (ER). Ego-centric videos capture a series of consecutive actions from a first-person perspective. The MLLM needs to reason for a question about a specific behavior in the video, e.g., predicting for the event which is correlated or satisfies a certain causal relationship with the behavior (Figure 2 e). Both videos and QA annotations are collected from the NLQ task of Ego4D . The ER task is structured as multiple-choice, with a total of 352 questions created for this task.

Plot Question-Answering (PQA). In this task, the MLLM needs to reason for questions about a plot in a narrative video, shown as Figure 2 (f). The video is sampled from the movies, TV series, and animated series in our ULVC. There are 539 question-answer pairs created by manual annotation. During annotation, the human annotators are asked to only provide necessary details about the plot but not to suggest any objective hints, e.g., the two characters in the example video are referred as cat and mouse, rather than Tom and Jerry. Therefore, it can prevent the question from being short-cut by the MLLM’s common-sense knowledge (more details about PQA can be found in the Appendix E.6).

Sub-Scene Captioning (SSC). In this task, the MLLM needs to generate the caption for a sub-scene in a long video. The long videos in SSC are sampled from the Movie101 dataset , while the questions and answers are manually annotated. During annotation, the human annotator is asked to provide a detailed description for the sub-scene as the ground-truth answer. Besides, they need to offer necessary clues in their questions such that the referred sub-scenes can be identified without ambiguity. During evaluation, we employ GPT-4 to measure the quality of caption in comparison with the ground-truth. Details about annotation and evaluation are presented in Appendix E.7 and F.3.

2.3 Multi-Detail LVU

Action Order (AO). In this task, the MLLM needs to predict the right order for a sequence of actions (Figure 2 h). The actions are presented by short video clips, called probes. The probes are formulated in two different ways. One is made up of clips from the Kinetics dataset , where each clip represents a distinct action. The other one is from the consecutive clips of an action in the ActivityNet-Caption dataset . The probes are inserted into a long background video, which is sampled from ULVC. There are 259 AO questions in total. The task is structured as a multiple-choice prblem, where the right order is selected from the misleading options provided by the annotator.

Action Count (AC). This task requires the MLLM to count the occurrences of an action within a long video (Figure 2 i). Each action corresponds to multiple short probe clips sampled from the Kinetics dataset . The probes of an action are inserted into a long background video sampled from ULVC. We also perform manual examination to ensure that the inserted action does not exist in the original background video. A total of 206 evaluation instances have been created. The AC task is structured as a multiple-choice problem, with performance measured by accuracy.

Experiments and Analysis

We perform a comprehensive investigation of 20 MLLMs based on MLVU, both open-source and proprietary. The experimental MLLMs can be partitioned into three categories. 1. Image MLLMs, which are primarily fine-tuned by image-related instructions. We consider the following models for this category: Otter-I , LLaVA-1.6 , InternVL-1.5 , Claude3-Opus , Qwen-VL-Max , and GPT-4 Turbo ). 2. Short Video MLLMs, which are fine-tuned by short-video related instructions. This category includes: Otter-V , mPlug-Owl-V , Video LLaMA-2 , Video ChatGPT , VideoChat , VideoChat2 , and Video-LLaVA . 3. Long Video MLLMs, which are optimized for their long-video understanding capability. This category includes: MiniGPT4-Video , LLaMA-VID , Movie-LLM , MA-LMM , MovieChat , TimeChat , and GPT-4o . For Image MLLMs, we leverage their multi-image inference capabilities to process segmented frames from original videos. In the case of Video MLLMs, we employ either a uniform sampling strategy or a frame rate sampling strategy for video processing. All models are evaluated based on their official implementations or available APIs, where the evaluation is conducted in a zero-shot manner. More details about the evaluation are provided in the Appendix F.

2 Overall Performance

The overall evaluation results for all investigated MLLMs are shown in Table 1. The individual performance is reported for each individual task; meanwhile, the average performances are reported for the multiple-choice (M-Avg) and generation tasks (G-Avg), respectively. We can obtain two primary observations from the demonstrated results.

On one hand, the recently released GPT-4o achieves an overwhelming advantage in our benchmark. It outperforms the rest of methods by a big margin in terms of the average performances, with an M-Avg of 64.6% (within 0-100%) and a G-Avg of 5.80 (within 0.0-10.0). Besides, it also maintains the leading position in every individual task. As for the open-source models, InternVL-1.5 achieves the highest average performance in multiple-choice tasks, whose M-Avg scores reaches 50.4%. Such a performance is impressive knowing that it even slightly goes beyond GPT-4 Turbo . Meanwhile, LLaMA-VID presents the top performance for the generation tasks, whose G-Avg score reaches 4.22. However, such a score is significantly lagging behind GPT-4 Turbo and GPT-4o.

On the other hand, although GPT-4o achieves a huge advantage in comparison with other methods, it actually struggles to handle most of the tasks in the benchmark. For example, it only ends up with 64.8% in dealing with the needle question-answering (NQA) task. Whereas its analogue tasks in the text domain, e.g., NIHS (Needle-In-the-HayStack-Search) and Passkey Retrieval, can be effectively conquered by many of the existing long LLMs [10; 54]. At the same time, it exhibits even less reliability given tasks like ego-reasoning (ER), action ordering (AO), and action count (AC), while other baseline methods produce even worse performances in these scenarios. The above observations indicate that long-video understanding remains a tough challenge for today’s MLLMs.

In addition to the primary conclusions from the overall performances, we can also make the following interesting observations about the individual tasks. First of all, the multiple-choice holistic tasks, i.e., topic retrieval (TR) and anomaly recognition (AR), present much higher differentiation than other tasks. Proprietary MLLMs, like GPT-4o, GPT-4-turbo, and superior open-source models, like InternVL-1.5, can accurately solve such problems; meanwhile, many other popular MLLMs still fail to generate meaningful performances. Knowing that the two tasks only require an overall understanding of the long videos, they can serve as a preliminary indicator of MLLMs’ LVU ability.

Besides, it’s hard to deal with tasks which need nuanced understanding of multiple details. Although several MLLMs can handle single-detail LVU tasks to some extent, their performances suffer from catastrophic degradation when addressing multi-detail LVU tasks. Most methods, except for GPT-4o, fail entirely in action order (AO) and action count (AC) tasks. Additionally, most approaches struggle with the summarization task, which require recalling multiple nuanced details from long videos.

Finally, the existing long-videos MLLMs (except for GPT-4o) are no better than those primarily trained from short-video or image related instructions. Although such models can intake much longer videos (i.e. much more frames than other baselines), it turns out that the extended input information is not effectively utilized. On the contrary, Video-LLaVA can achieve a relatively competitive performance, especially in multiple-choice, despite that it can only intake 8 frames.

As a brief conclusion, although today’s MLLMs can deal with some preliminary LVU tasks, it remains a tough challenge to achieve an in-depth understanding of nuanced information within long videos.

3 Detailed Analysis

We analyze the impact from video length and three factors: context length, image understanding (IU) ability, LLM-backbone. All the factors are empirically critical to MLLMs’ LVU performances.

In the first place, we evaluate MLLMs’ performances across various video lengths. For this purpose, we introduce a derivative dataset alongside MLVU, called MLVU Time-ladder. In this dataset, the same kinds of evaluation tasks are created for videos of variant lengths, including 180s, 360s, and 600s (more details presented in Appendix D). As shown in Figure 3, the performances of all models tend to decline as the video length grows, which indicates that the existing MLLMs’ LVU abilities are severely constrained by the video length. By comparison, image models, like GPT-4-turbo , and short-video models, like VideoChat2 , VideoLLaMA , are more vulnerable to the growth of video length, while long video models, like MiniGPT4-Video, can be relatively more resilient.

We further examine MLLMs’ performances (measured by M-Avg) under varying context lengths. Specifically, we increase MiniGPT4-Video’s input from 16 to 90 frames and GPT-4o’s input from 16 to 256 frames (the left side of Table 2). As the input length extends, both models consistently show improved performance. To investigate the impact of MLLMs’ image understanding (IU) ability, we reference the experiment results from MMMU (the middle of Table 2). It’s evident that the MLLMs’ LVU performances basically align with their image understanding performances in MMMU. Finally, we compare MLLMs with different backbones (the right side of Table 2). The results show that LVU performances improve with larger (Vicuna-13B vs. Vicuna-7B) and more capable (Mistral-7B vs. Llama-2-7B) backbone encoders. These observations indicate that LVU is the result of multiple complex factors, with the ability to perceive longer videos and effectively utilize the perceived information being crucial for the improvement of LVU.

Conclusion & Discussion

This paper presents MLVU, a novel benchmark for the assessment of long video understanding. With several critical innovations: the substantial extension of video lengths, the inclusion of various video genres, and the development of diversified LVU-oriented evaluation tasks, the new benchmark is able provide a comprehensive and in-depth analysis for MLLMs’ long-video understanding performance. The empirical study on MLVU reveals LVU remains a technically challenging problem for today’s state-of-the-art MLLMs. Future advancements may call for the joint optimization of complex factors, such as context length, image understanding ability, and even LLM backbones. We anticipate this benchmark will facilitate future research in long-video understanding of MLLMs.

Limitations. While MLVU basically covers the major dimensions of long-video understanding, it can still be improved for even better comprehensiveness. For instance, there can be tasks related to high-resolution videos, or more specific tasks such as tracking and low-level processing. Our work will be a persistent effort, where new tasks and video types will be continually introduced.

References

Appendix A Overview of Appendix

B: Collecting Details of our Universal Long Video Collection (ULVC).

C: Distribution of Video Durations for Each Task.

F: Details of Baselines and the Evaluation Process.

G: Explorations of Video Retrieval Augmented Generation.

J: Licensing, Hosting and Maintenance Plan.

Appendix B Collecting Details of our Universal Long Video Collection (ULVC)

In the initial stage of our Multi-task Long Video Understanding (MLVU) benchmark creation, we first collected long-form videos from a variety of sources to form our Universal Long Video Collection (ULVC). The entirety of the long videos incorporated into our MLVU benchmark were selected, edited, or synthesized from ULVC.

Specifically, our ULVC consists of a diverse set of 757 long videos. This collection includes 168 movies sourced from the Movie101 dataset and the MovieChat dataset , as well as 60 documentaries from the MovieChat dataset . Additionally, it comprises 65 game videos from the MineDojo dataset , 200 surveillance videos from the UCF-Crime dataset , and 100 ego-centric videos from the Ego4D dataset . Furthermore, our team independently collected 72 cartoons and 92 TV series to enrich the dataset.

It’s important to clarify that the quantity of videos in the ULVC does not directly correspond to the number of videos and questions in our MLVU benchmark, which are 1334 and 2593 respectively. For example, a two-hour movie from the ULVC might be utilized in its entirety for the Sub-Scene Captioning task, or it could be segmented into several approximately 10-minute clips for the Video Summary task, or even used as a background video for synthetic video generation. Moreover, a single video could be annotated with multiple questions simultaneously.

Appendix C Distribution of Video Durations for Each Task

Figure 4 illustrates the detailed video duration distribution for each task in our MLVU, featuring videos of varying lengths from three minutes to over 120 minutes.

Appendix D Details of the MLVU Time-Ladder

As discussed in Section 3.1, most tasks in our MLVU are subject to segment-level annotation. This approach provides us with the flexibility to adjust the length of the video without requiring additional human annotators. Building on this strategy, as mentioned in Section 4.3, we have generated a derivative dataset, MLVU Time-Ladder, which includes videos of varying durations - specifically 3, 6, and 10 minutes. This dataset allows us to investigate how video duration impacts LVU task difficulty.

Specifically, during the annotation process of the VS task, we guided annotators to delineate the summarization in accordance with the initial 3 and 6-minute segments. For the PQA and SSC tasks, we requested annotators to identify the segments within the extended video where the pertinent answers are located. In the case of the ego reasoning task, the Ego4D dataset already comprises the intervals where the answers reside. Lastly, for the synthetic tasks of NQA, AO, and AC, we possess the capability to directly generate the necessary video lengths.

Appendix E Annotation Details of MLVU

The questions and corresponding answers for the TR task were meticulously annotated by human annotators, following the specific guidelines illustrated in Figure 5. We required the annotators to design questions related to the reasoning of the video topic, rather than focusing on the creation of questions about minor details. More visualized examples of TR task can be found in Figure 14.

E.2 Anomaly Recognition (AR).

The anomaly recognition task did not involve manual annotation. We utilized videos exceeding three minutes in duration, extracted from the UCF-Crime dataset . We also modified the original labels to fit a multiple-choice format.

E.3 Video Summarization (VS).

The ground truth data for the VS task were derived from manual annotations. We instructed the annotators to use pronouns instead of specific character names in all annotations. This guideline stemmed from the inherent constraints of most existing MLLMs, which generally lacked the capacity to process audio or subtitles. This made it difficult for these models to identify specific characters. The annotation instructions and examples provided to the annotators are elaborated in Figure 6. More visualized examples of VS task can be found in Figure 14.

E.4 Needle Question-Answering (NQA).

We leveraged the GPT-4 and the detailed video caption data from the WebVid dataset to facilitate a semi-automated generation of annotated questions and answers for the NQA task. Initially, we selected video clips from WebVid, which we refered to as needle clips. The corresponding captions of these needle clips were then fed into GPT-4, which generated question-answer pairs based on the information encapsulated in the captions. The specific prompt provided to GPT-4 is depicted in Figure 7. The generated questions were carefully crafted to focus on a particular detail within the needle clip. These questions were structured to incorporate the maximum number of hints to effectively guide MLLMs in grounding the content of the needle within the context of the longer video. Following this, we randomly selected longer background videos from our ULVC and manually ensured that the scene indicated by the needle’s question did not feature in these background videos. The final step involves integrating the needle into the longer video, thereby producing the final needle question video. More visualized examples of NQA task can be found in Figure 15.

E.5 Ego Reasoning (ER).

The video resources, questions, and correct responses used in the ER task were derived from the Natural Language Queries (NLQ) task within the Ego4D dataset . This data was restructured to fit a multiple-choice question format.

E.6 Plot Question-Answering (PQA).

The PQA task’s questions and answers were annotated by human annotators, following specific guidelines illustrated in Figure 8. We instructed the annotators to craft questions that probe into the intricate plot details encapsulated within the videos. These questions were designed to encompass both perception and reasoning aspects. We stipulated that both questions and their corresponding answers should avoid the use of specific character names or any objective hints, and should instead utilize pronouns. This approach was strategized to prevent potential information leakage, given that MLLMs often demonstrate a familiarity with the storylines of well-known movies and TV series. Such common-sense knowledge could potentially allow the MLLMs to answer questions correctly without the essential requirement of analyzing the input video.

Nonetheless, the complexity of character interactions and actions in longer videos poses a challenge to conveying plot details using only pronouns and feature descriptions. Previous datasets for plot question answering that avoided the use of character names often resulted in compromised question diversity and tended towards generalized queries. We illustrate this through a comparative analysis of TVQA , Moviechat , and our PQA dataset’s question word clouds in Figure 9. While TVQA provides a diverse range of questions, it does so by employing specific character names. In contrast, Moviechat avoids character names, but its questions are frequently overly broad, lack specific plot details, and exhibit diminished diversity. Our PQA dataset successfully navigates these challenges, offering a diverse range of questions without resorting to the use of character names. More visualized examples of PQA task can be found in Figure 15.

E.7 Sub-Scene Captioning (SSC).

In the development process of the SSC task, we employed human annotators to generate both prompts and standard caption data. The specific guidelines provided to annotators are illustrated in Figure 10. Initially, the annotators identified a specific, easily referable sub-scene within a lengthy movie. Subsequently, they crafted a prompt replete with adequate clues to reference this scene, ensuring the uniqueness of these clues throughout the entire film. To prevent any leakage of information, the prompt was designed to exclude any character-specific names or objective hints, instead incorporating rich descriptive details to allude to the plot. Following this, the annotators produced a detailed caption for this sub-scene, and deconstructed the caption into multiple, non-redundant "scoring points" to facilitate quantitative assessment (the details of the evaluation metric can be found in Section F.3). More visualized examples of PQA task can be found in Figure 16.

E.8 Action Order (AO).

The videos, questions, and answers for the action order task were all synthetically generated. In order to maintain the high quality of our evaluation data, we adopted a dual-strategy approach. Firstly, we selected actions for the probe videos that were not commonly seen in most films, such as making jewelry and water skiing. Secondly, in the selection of background videos, we conducted a cursory review of the video content to further ensure that the actions referenced in the questions were not present in the video. This rigorous methodology ensured the reliability of our data.

E.9 Action Count (AC).

The process of data acquisition and annotation for the action count task closely mirrored that of the action order task. All videos, questions, and answers were synthetically generated. We employed a strategy consistent with the action order task to ensure the validity and reliability of our evaluation data.

Appendix F Details of Baselines and the Evaluation Process

In this section, we detail the primary baselines evaluated on our MLVU. For image-based MLLMs, most available models lack multi-image inference capabilities. Consequently, we select Otter-I, LLaVA-1.6, and InternVL, which offer official multi-image implementations. Additionally, we include three proprietary models—Claude-3-Opus, Qwen-VL-Max, and GPT-4 Turbo—that provide APIs for multi-image inference. For available models, we estimate the maximum input frames based on their maximum LLM context length. For Claude and Qwen, the supported maximum image numbers are approximately 20; therefore, we select 16 frames to ensure fair comparisons. Similarly, we choose 16 frames for GPT-4 Turbo to maintain consistency. In terms of video MLLMs, we adhere to the default settings for frame sampling strategies. For instance, VideoChat2 uniformly samples 16 frames, whereas LLaMA-Vid samples 1 frame per second. Specifically, the GPT-4o can support a maximum of approximately 500 images when each image’s resolution is set to 512×\times512 pixels. Consequently, we opt for a sampling rate of 0.5 fps to accommodate the majority of our videos. Table 3 and Table 4 show the LLM-backbone as well as evaluated weight links used in each MLLM.

F.2 Inference Detatils

We have developed two templates specifically for Multiple-Choice and Generation tasks, as illustrated in Figure 11. Distinct system prompts were designed to accommodate the differences between video-based and image-based MLLMs. Considering the variances in task requirements, we incorporated “option prediction guidance” into the Multiple-Choice template to aid in option extraction. Conversely, in Generation tasks, we do not implement any additional interventions but employ fixed-question guidance to enable models to respond to diverse task questions. In our evaluation, the templates are seamlessly integrated into the evaluation code of open-release models or available API of proprietary models.

F.3 Evaluation Metrics

For the evaluation of Multiple Choice tasks, we directly compute absolute accuracy by matching the predicted option with the ground truth. In Generation tasks, we develop multiple criteria for assessment and employ GPT-4 to rank the alignment between generated texts and the provided answers. As illustrated in Figure 12, we use “Accuracy” and “Relevance” to benchmark Sub-scene Captioning, and “Completeness” and “Reliability” to evaluate the capabilities of Video Summary.

Appendix G Explorations of Video Retrieval Augmented Generation

As discussed in Section 4.3, most MLLMs are adversely affected by video length. Drawing inspiration from the use of Retrieval Augmented Generation (RAG) in video understanding, we have developed a zero-shot RAG strategy and seamlessly integrated it into existing MLLMs. Table 5 displays the performance comparison between the baseline models and the models employing our RAG strategy. It is noteworthy that all methods benefit from the RAG strategy in Needle QA, Ego Reasoning, and Plot QA. Conversely, minimal improvement is observed in Action Count, and a decrease is noted in Action Order and Overall Reasoning. This is primarily because RAG facilitates the retrieval of detail-oriented video clips, which makes models more likely to focus on answer-related cues in specific single-detail reasoning tasks. However, RAG exhibits limited capabilities in multi-detail reasoning and holistic understanding tasks, which require global perception and knowledge aggregation.

Appendix H More Experimental Results

Table 6 shows the leaderboard of each task in our MLVU.

H.2 Comprehensive evaluation results on MLVU Time-ladder

Table 7, Table 8 and Table 9 show the experimental results on MLVU Time-ladder within 3 minutes, 6 minutes and 10 minutes. As anomaly recognition and topic reasoning are global tasks, thus, we can not annotate them in MLVU Time-ladder as the way provided in Appendix D.

Appendix I More Visualized Examples of MLVU.

Appendix J Licensing, Hosting and Maintenance Plan

We bear all responsibilities for the licensing, distribution, and maintenance of our dataset.

MLVU can be viewed and downloaded on GitHub at https://github.com/JUNJIE99/MLVU or on Huggingface at https://huggingface.co/datasets/MLVU/MVLU. We assure its long-term preservation for future reference and use. The annotations for questions and answers are provided in the JSON file format, while the raw videos are available in the MP4 format.

For the video files, we do not hold any copyright, and all video rights belong to the video authors. To facilitate user usage, under the premise of a user agreement that the data will be used solely for research purposes and not for commercial use, we provide the download methods for these videos. We have taken measures such as reducing resolution, modifying aspect ratio, and editing to minimize the impact on the original work rights. However, there is still a risk that copyright holders may request the removal of some data. If this happens, we will follow the practice of Movienet to sparsely collect these video frames. This will not significantly impact our data usage, as all our annotations are focused on the visual information in the videos, unrelated to the audio. Moreover, most of the models we evaluate are essentially still frame-based video processing and do not involve audio processing. If retaining video frames is also not allowed, we will still preserve the annotation data and provide metadata for the corresponding videos.

Metadata can be found at https://huggingface.co/datasets/MLVU/MVLU/tree/main/MLVU/json.

Appendix K Datasheet

For what purpose was the dataset created?

Answer: The creation of MLVU is to facilitate the evaluation and development of MLLM’s long video understanding capabilities. Compared to previous video understanding benchmarks, MLVU possesses a substantial extension of video length, diversified video categories, and diversified evaluation tasks. MLVU is the first comprehensive benchmark for long video understanding.

Who created the dataset (e.g., which team, research group) and on behalf of which entity (e.g., company, institution, organization)?

Answer: The MLVU is created by Junjie Zhou (BUPT, BAAI), Yan Shu (BAAI), Bo Zhao (BAAI), Boya Wu (BAAI), Shitao Xiao (BAAI), Xi Yang (BAAI), Yongping Xiong (BUPT), Bo Zhang (ZJU), Tiejun Huang (PKU, BAAI), and Zheng Liu (BAAI)

Answer: Beijing Academy of Artificial Intelligence, Beijing University of Posts and Telecommunications, Peking University, and Zhejiang University.

K.2 Composition

What do the instances that comprise the dataset represent? (e.g., documents, photos, people, countries)

Answer: Each instance in our dataset represents a long video ranging from 3 minutes to 2 hours in duration, a question, a standard answer, and for the multiple-choice task, there are also 4 options. Videos are stored in MP4 file format, while the questions, standard answers, and options are all stored in JSON format files.

How many instances are there in total (of each type, if appropriate)?

Answer: In total, we collect 2593 instances, including 2175 multiple-choice questions and 418 free-form generation questions. The number of distinct videos is 1334, as one video may correspond to multiple questions. The specific distribution of question types and video duration can be found in Figure 1 in the main paper.

Does the dataset contain all possible instances or is it a sample (not necessarily random) of instances from a larger set?

Answer: Only a part of the questions was sampled from existing datasets (the ego reasoning task and the anomaly recognition task). The rest of the questions were newly annotated by us (for example, plot question answering, sub-scene captioning, video summarization, topic reasoning), or they were substantially modified based on existing annotations (needle question answering, action count, action order). During the data collection and annotation process, we ensured that the instances within each task were diverse in terms of videos and questions.

Is there a label or target associated with each instance?

Answer: Yes, for the multiple-choice question type, each instance provides the correct option. For the free-form generation task, each instance comes with a standard caption or summary answer.

Is any information missing from individual instances?

Are relationships between individual instances made explicit (e.g., users’ movie ratings, social network links)?

Answer: Some instances may have the same video but different questions and answers. Each instance clearly indicates the corresponding video, and each video has a unique identifier.

Are there recommended data splits (e.g., training, development/validation, testing)?

Answer: No. Our MLVU is specifically designed for evaluation, with its core objective being to assess the capability of MLLMs to understand long-term video.

Are there any errors, sources of noise, or redundancies in the dataset?

Is the dataset self-contained, or does it link to or otherwise rely on external resources (e.g., websites, tweets, other datasets)?

Answer: All data will be publicly accessible in the dataset repository. Our annotations will be stored in JSON format. As our collected videos include parts of movies, TV series, documentaries, and cartoons, we have taken measures such as reducing resolution, modifying aspect ratio, and editing to minimize the impact on the original work rights. However, there is still a risk that copyright holders may request removal of some data. If this happens, we will follow the practice of Movienet to sparsely collect these video frames. This will not have much impact on our data usage, as all our annotations are focused on the visual information in the videos, unrelated to the audio. If retaining video frames is also not allowed, we will still keep the annotation data and provide metadata for the corresponding videos.

Does the dataset contain data that might be considered confidential?

Does the dataset contain data that, if viewed directly, might be offensive, insulting, threatening, or might otherwise cause anxiety?

K.3 Collection Process

The data collection process is described in Section 3.2 of the main paper and Appendix B, D, and E.

K.4 Uses

Has the dataset been used for any tasks already?

Answer: Yes, MLVU has been used to evaluate the long-video understanding capabilities of as many as 20 different MLLMs. For specific details, please refer to Section 4.

What (other) tasks could the dataset be used for?

Answer: MLVU is primarily used to comprehensively evaluate the long-video understanding capabilities of MLLMs. Since MLVU includes multiple different tasks, it can also be used to individually assess the ability of MLLMs or other video-specific models on a particular task, such as video summarization.

Is there a repository that links to any or all papers or systems that use the dataset?

Is there anything about the composition of the dataset or the way it was collected and preprocessed/cleaned/labeled that might impact future uses?

Answer: As our collected videos include parts of movies, TV series, documentaries, and cartoons, we have taken measures such as reducing resolution, modifying aspect ratio, and editing to minimize the impact on the original work rights. However, there is still a risk that copyright holders may request removal of some data. If this happens, we will follow the practice of Movienet to sparsely collect these video frames. This will not have much impact on our data usage, as all our annotations are focused on the visual information in the videos, unrelated to the audio. If retaining video frames is also not allowed, we will still keep the annotation data and provide metadata for the corresponding videos.

Are there tasks for which the dataset should not be used?

Answer: The MLVU cannot be used to evaluate the image or video generation capabilities of the MLLM.

K.5 Distribution

Will the dataset be distributed to third parties outside of the entity (e.g., company, institution, organization) on behalf of which the dataset was created?

Answer: Yes. The benchmark is publicly available on the Internet.

How will the dataset will be distributed (e.g., tarball on website, API, GitHub)?

Answer: The benchmark is available on GitHub at https://github.com/JUNJIE99/MLVU or on Huggingface at https://huggingface.co/datasets/MLVU/MVLU.

Will the dataset be distributed under a copyright or other intellectual property (IP) license, and/or under applicable terms of use (ToU)?

Have any third parties imposed IP-based or other restrictions on the data associated with the instances?

Do any export controls or other regulatory restrictions apply to the dataset or to individual instances?

K.6 Maintenance

Who will be supporting/hosting/maintaining the dataset?

Answer: The authors will be supporting, hosting, and maintaining the dataset.

How can the owner/curator/manager of the dataset be contacted (e.g., email address)?

Answer: Please contact the official email of our project (mlvubenchmark@gmail.com), or contact the one of the authors (zhoujunjie@bupt.edu.cn; shuyan9812@gmail.com)

Answer: No. We will make announcements if there are any.

Will the dataset be updated (e.g., to correct labeling errors, add new instances, delete instances)?

Answer: Yes. We will post new update in https://github.com/JUNJIE99/MLVU and https://huggingface.co/datasets/MLVU/MVLU if there is any.

If the dataset relates to people, are there applicable limits on the retention of the data associated with the instances (e.g., were individuals in question told that their data would be retained for a fixed period of time and then deleted)?

Answer: People may appear in the reference videos. People may contact us to exclude specific data instances if they appear in the reference videos.

Will older versions of the dataset continue to be supported/hosted/maintained?

Answer: Yes. Old versions will also be hosted in https://github.com/JUNJIE99/MLVU and https://huggingface.co/datasets/MLVU/MVLU.

If others want to extend/augment/build on/contribute to the dataset, is there a mechanism for them to do so?

Answer: If others wish to add data, they can apply to do so provided the data is compliant and reasonable. However, making other modifications based on our dataset is currently not allowed.