MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, Kai Chen

Introduction

As a ubiquitous format for multimedia, video holds a pivotal role in people’s lives, serving purposes such as knowledge dissemination, sharing life experiences, and entertainment. The rapid proliferation of video content has reshaped communication, learning, and connection in the digital age. The vast amount of online video content underscores the importance of algorithmic video understanding. Current video understanding paradigms, which often focus on specific tasks , typically excel only on in-domain data. An ideal video understanding system should demonstrate robust zero-shot capabilities, accurately discern contextual, emotional, and linguistic details within a video, and engage in free-form dialogues with humans .

With the rapid development of Large Language Models (LLMs) , Large Vision Language Models (LVLMs) have also seen significant advancements. Typical video-language models developed by researchers utilize frame-level or clip-level visual features extracted by vision encoders , align these features with language embeddings via a projector, and process these embeddings with a fine-tuned large language encoder . The models are fine-tuned with video instruction data and quantitatively assessed on free-form VideoQA benchmarks . The current evaluation of Video-LLMs is characterized by the following limitations:

Short Videos: Existing VideoQA datasets primarily consist of short videos, typically lasting less than a minute. Meanwhile, most web video content spans several minutes or longer, creating a discrepancy between the evaluation benchmark and real-world application scenarios.

Limited Capabilities: Current VideoQA benchmarks are limited to several basic video tasks , including concept existence, object relationship recognition, and activity recognition. There are more fine-grained perception and reasoning capabilities not encompassed by existing benchmarks.

Biased Evaluation: Existing evaluation paradigms employ GPT-3.5 to score open-ended answers generated by video-language models. Our preliminary study indicates that GPT-3.5-based evaluation is less accurate and exhibits significant discrepancy relative to human preferences, diminishing the credibility of the evaluation results.

To address these problems, we develop a new VideoQA benchmark, MMBench-Video, to evaluate the effectiveness of LVLMs in video understanding. It incorporates approximately 600 web videos with rich context from YouTube, spanning 16 major categories, including News, Sports, etc., covering most video topics people watch in their daily lives. Each video ranges in duration from 30 secs to 6 mins, to accommodate the evaluation of video understanding capabilities on longer videos. The benchmark includes roughly 2,000 original question-answer (QA) pairs, contributed by volunteers, covering a total of 26 fine-grained capabilities. During dataset collection, we implement quality control strategies to explicitly increase the proportion of temporal indispensable questionsA visual question is temporal indispensable if it can not be correctly solved by viewing any random frame. . Quantitative statistics show that MMBench-Video significantly differs from existing benchmarks in terms of temporal duration, context richness, and temporal indispensability.

During evaluation, an LVLM produces free-form responses to visual questions. Given the variability in the lengths and styles of ground-truth answers, accurately assessing these responses presents a significant challenge. In light of the limitations observed in previous evaluations powered by GPT-3.5, we propose the use of the more powerful GPT-4 for automated scoring. This approach prioritizes semantic similarity while overlooking minor discrepancies in language organization. Employing a carefully crafted evaluation prompt, our GPT-4-based evaluation exhibits improved quality in terms of accuracy, consistency, and alignment with human judgment.

Based on MMBench-Video, we perform a thorough evaluation of mainstream LVLMs, including open-source video-language models (Video-LLMs), as well as both open-source and proprietary LVLMs for image understanding. We report their performance across diverse capabilities, as depicted in Fig. 1. The performance rankings enable direct comparisons between models, revealing critical insights into their limitations. Surprisingly, existing Video-LLMs exhibit subpar performance on MMBench-Video, significantly underperforming proprietary LVLMs and even lagging behind open-source LVLMs, such as Idefics2 and InternVL-Chat-v1.5 . To further investigate these models’ capabilities, we employ image VQA benchmarks to assess their image understanding skills, again observing a substantial gap between Video-LLMs and the state-of-the-art LVLMs. The comprehensive assessment underscores the significant performance disparities between Video-LLMs and leading LVLMs in both spatial and temporal understanding, highlighting areas requiring future improvement.

In summary, the contributions of this work are as follows:

Innovative VideoQA Benchmark: MMBench-Video features long-form, diverse videos sourced from the web, encompassing a broad spectrum of topics. It includes original, high-quality visual questions crafted by volunteers, spanning dozens of fine-grained capabilities.

Enhanced Scoring Methodology: We assess the limitations of using low-quality LLMs, such as GPT-3.5, for scoring model responses. To address this, we implement a GPT-4-based evaluation paradigm, which offers superior accuracy, consistency, and a closer alignment with human judgments.

In-depth Evaluation: Our comprehensive assessment of various LVLMs on MMBench-Video reveals detailed insights into their performance across multiple fine-grained capabilities. The results underscore the current limitations of Video-LLMs in spatial and temporal understanding, guiding future research and development.

Related Work

The success of Large Language Models (LLMs) such as GPTs and LLaMA has spurred significant advancements in Large Vision-Language Models (LVLMs). Flamingo has demonstrated impressive few-shot capabilities by integrating gated cross-attention blocks to connect pre-trained vision and language models. BLIP employs a Querying Transformer to bridge the modality gap between a frozen image encoder and a language encoder. LLaVA leverages GPT-4 to create instruction-following data for vision-language tuning, with its learning paradigm and instruction tuning corpus being widely adopted by subsequent works . In the realm of video-language models, Video-ChatGPT aligns frame-level vision features with language embeddings via a linear projector, whereas VideoChat utilizes a learnable Q-former, inspired by BLIP-2. Subsequent works like Video-LLaMA integrate audio features, and Video-LLaVA learns from a mixed dataset of images and videos. Additionally, proprietary APIs such as GPT-[4v/4o] , Gemini , and Reka have been made publicly available, supporting various input formats including single or multiple images. We present a comprehensive evaluation of existing LVLMs, encompassing Video-LLMs as well as open-source and proprietary LVLMs for images, using the proposed MMBench-Video to provide a detailed landscape of their capabilities.

2 Video Question Answering

Video Question Answering (VideoQA) is a critical method for assessing the depth of understanding that models possess regarding video content. The research community has progressively developed a wide array of VideoQA benchmarks, spanning various visual domains such as movies , TV shows , video games , synthetic scenarios , and egocentric videos . These benchmarks typically assess models trained on their respective training sets, demanding concise answers for evaluation. However, Large Vision and Language Models (LVLMs), which are often not trained on domain-specific data, face challenges in adapting to these benchmarks due to their diverse answer styles. To mitigate this, Video-ChatGPT employs GPT-3.5 as a scoring mechanism for free-form responses from VLMs. The method was applied to evaluate several popular benchmarks , covering topics including concept existence, objection relationship, and activity recognition. Despite its broad adoption , this approach is limited by suboptimal accuracy and stability, as well as poor alignment with human preferences. Additionally, those benchmarks primarily consist of short videos, which contrasts with the typical length of web videos. In response, we present MMBench-Video, a novel dataset tailored for longer videos, challenging models to generate detailed, free-form responses to complex questions. We adopt GPT-4-based evaluation, which improves correctness and robustness, offering a more stable evaluation strategy compared to previous methods.

MMBench-Video

In this section, we delve into the meticulous construction of MMBench-Video, outlining our strategic approach to video question selection, the conceptualization and design of a comprehensive capability taxonomy, and the innovative methods employed to enhance the temporal relevance and quality of questions. Additionally, we present detailed statistics of MMBench-Video and contrast it with existing VideoQA benchmarks, thereby illustrating its unique features and contributions to the field.

Video Collection. To create a VideoQA benchmark, a prevalent approach involves generating question-answer pairs for videos sourced from existing datasets. For example, MSRVTT-QA and MSVD-QA are derived from video retrieval datasets , while ActivityNet-QA is constructed using an action recognition dataset . Most existing VideoQA datasets are limited to short videos with a constrained number of shots, exhibiting limited diversity in content. To develop a benchmark that more closely mirrors the web video content commonly consumed by viewers, we propose the creation of a long-form, multi-shot VideoQA benchmark. The benchmark draws its content directly from YouTube, offering several distinct advantages. Firstly, YouTube’s extensive metadata, including video titles, click metrics, and subtitles, provides valuable context for video understanding. Secondly, as a leading global streaming platform, YouTube’s vast user base ensures the dataset’s diversity.

Drawing inspiration from the YouTube-8M labels, our categorization scheme encompasses 16 major categories (Fig. 5), spanning from engaging topics like ‘Entertainment and Sports’ to enlightening subjects such as ‘Science and Knowledge’. Volunteers are instructed to navigate through YouTube and collect videos that align with these designated categories. In line with our objective to amass long-form content, volunteers are directed to disregard videos with durations of less than 30 seconds. Although we impose no upper limit on the length of the web videos collected, all question-answer pairs composed for a video will be derived from a clip no longer than 6 minutes. This helps maintain a practical balance between video duration and the task complexity.

Capability Taxonomy. Inspired by MMBench , we have developped a 3-level (L-1 to L-3) hierarchical capability taxonomy (Fig. 2). The top level encompasses two broad capabilities: Perception and Reasoning. Besides the six L-2 capabilities inherited from MMBench, we further introduce three additional L-2 capabilities specific to MMBench-Video: Hallucination, Commonsense Reasoning, and Temporal Reasoning. Hallucination assesses whether a model is prone to generating content that includes misleading or inaccurate information. Commonsense Reasoning evaluates a model’s ability to integrate necessary commonsense knowledge into its reasoning processes. Temporal Reasoning examines a model’s proficiency in understanding the relationships between events unfolding at different video points. This taxonomy comprises a total of 26 leaf capabilities, which collectively address a comprehensive spectrum of cognitive processes involved in video comprehension.

Composing Questions and Answers. A well-known issue in existing VideoQA benchmarks is the prevalence of non-temporal questions, which are those that can be accurately answered based on nearly any frame within a video, rendering them effectively ‘static’. These questions fail to adequately assess a model’s ability to temporal understanding. In the curation of MMBench-Video, we prioritize questions that necessitate temporal reasoning and strive to minimize the occurrence of static questions. Recognizing the necessity of evaluating certain coarse perception capabilities, such as Video Style and Video Topic, it is impractical to entirely eliminate static questions. Instead, we focus on significantly reducing their proportion within the benchmark.

In MMBench-Video, each video is accompanied by multiple independent questions designed to assess one or more specific leaf capabilities. For instance, a question that requires identifying and counting a particular type of object would evaluate both Object Recognition and Counting capabilities. To ensure the quality and relevance of the questions and their corresponding answers, volunteers involved in the collection process are provided with the following five guidelines to adhere to:

Each question should evaluate one or multiple leaf capabilities within the established taxonomy.

You are encouraged to formulate temporal indispensable questions, as long as it’s feasible for the corresponding video content and capability category.

Avoid including specific timestamps in the questions, such as “at 03:20 in the video”. Please use relative expressions like “at the end of the video" or “before/after a specific event" instead.

The questions should be free-form and exhibit linguistic diversified. Besides standard formats like What/Who/How, questions can also adopt a conversational styleFor example, “What is the score of the football game in the video?” (a “what” question) can be expressed as “Tell me the winning team and the final score.” (conversation style)..

Please provide informative and detailed answers for each question.

All generated question-answer pairs in MMBench-Video will be subjected to a meticulous cross-validation process to confirm their accuracy and adherence to the established guidelines. In addition to this, we implement an LVLM-based filtering mechanism to identify and eliminate a portion of static questions, as detailed in the supplementary material. The final MMBench-Video dataset comprises a diverse selection of web videos sourced from YouTube, accompanied by human-composed, original question-answer pairs designed to assess a comprehensive array of fine-grained capabilities.

Evaluation Paradigm. Given the varied length and style of ground-truth answers, automated robust evaluation that aligns with human judgments can be challenging. To address this, we propose a 3-grade marking scheme and utilize GPT-4 as our adjudicator. GPT-4 assigns a score from 0 to 3 based on the content similarity between the model’s output and the ground truth. Our experiments show that this evaluation framework exhibits strong consistency and alignment with human assessments.

2 Dataset Statistics

MMBench-Video comprises 609 video clips across 16 major categories, as depicted in Fig. 5), with durations spanning from 30 seconds to 6 minutes. The dataset has an average video length of 165 seconds, totaling 28 hours in aggregated duration. The duration distribution of the clips within MMBench-Video is illustrated in Fig. 5. The dataset includes 1,998 question-answer (QA) pairs, with each QA assessing one or multiple capabilities of a vision-language model. The distribution of QAs corresponding to each capability is visualized in Fig. 2. To highlight the distinct value of MMBench-Video, we compare its statistics with those of existing VideoQA benchmarks:

Duration & Shot NumbersWe adopt the open-source tools scenedetect to obtain the shot number of a video.. MMBench-Video is specifically designed as a long-form, multi-shot video dataset. As indicated in Tab. 1, our dataset boasts a substantially greater average duration than existing benchmarks. As shown in Fig. 5, videos in our benchmark display a long-tail distribution in shot numbers, with a maximum of 210 shots. This significantly surpasses all other benchmarks in average shot count.

Linguistic Characteristics of QAs. MMBench-Video features free-form video QA with rich linguistic diversity. In benchmarks such as MSVD-QA and MSRVTT-QA, questions are automatically generated and invariably begin with pronouns such as ‘what’, ‘who’, etc. Conversely, a significant proportion of questions in MMBench-Video are framed in a conversational manner, enhancing linguistic diversity (further details are available in the supplementary materials). Regarding answers, previous VideoQA benchmarks often provide responses that are limited to a single word or a brief phrase. In contrast, MMBench-Video strives to offer more comprehensive answers that extend beyond a single word. This is evident in the distribution of answer lengths, as shown in Tab. 1.

Capability Coverage. Existing benchmarks typically cover only a limited set of fine-grained capabilities and often lack an explicit capability taxonomy. For instance, the majority of questions in MSVD-QA and MSRVTT-QA assess the ability to determine the existence of concepts (such as humans or objects) and to recognize relationships between objects. In contrast, ActivityNet-QA and TGIF-QA extend this by including assessments of activity recognition and repetition counting. In MMBench-Video, we have established a comprehensive taxonomy encompassing 26 fine-grained capabilities, with each capability being evaluated using dozens to hundreds of original QAs.

Temporal Indispensability. In contrast to existing VideoQA benchmarks, MMBench-Video is designed to be temporal indispensable. In a preliminary study, we find that a great proportion of QAs in existing datasets can be correctly answered by LVLMs without providing the temporal context. The underlying factors can be categorized into two primary ones: (1) The brevity of source videos, characterized by the limited number of shots, allows for its content to be adequately represented by a single frame. (2) Many of the QAs are too simplistic and can be answered through guesswork rather than comprehension. For instance, MSVD-QA and MSRVTT-QA are replete with ‘who’ questions, which are commonly answered with general terms like ‘someone’, ‘man’, or ‘woman’. In MMBench-Video, we have made significant efforts to mitigate these factors.

To quantitatively measure the temporal indispensability of each VideoQA benchmark, we randomly sample 1000 QAs from MSVD, TGIF, MSRVTT, and ActivityNet, and conduct a study on the subsets as well as MMBench-Video. We evaluate GPT-4o (by far the most powerful LVLM) on these benchmarks under 1-frame and 8-frame settings, and present the results in Tab. 2. Notably, GPT-4o using a 1-frame input achieves a normalized score of approximately 50%, retaining over 80% of its performance compared to an 8-frame input across MSVD, TGIF, MSRVTT, and ActivityNet. In contrast, when assessed on the MMBench-Video, GPT-4o using a 1-frame input preserves only 47.8% of its efficacy compared to its performance with an 8-frame input, yielding a normalized score of just 26.0%. This marked difference underscores the temporal importance of MMBench-Video.

Experiment

Utilizing MMBench-Video, we assess a diverse array of large vision-language models (LVLMs), encompassing Video-LLMs and image-based LVLMs, both open-source and proprietary. For Video-LLMs, we utilize the default hyperparameters specified in their respective open-source implementations for inference. For image-based LVLMs, we conduct evaluations based on VLMEvalKit , employ greedy decoding during inference and cap the maximum number of output tokens at 512.

Open-Source Video-LLMs. We first identify and evaluate representative open-source Video-LLMs using MMBench-Video. Adhering to their default settings, these Video-LLMs process a sequence of video frames, with the number of frames varying from eight to dozens. Interestingly, we observe that all Video-LLMs exhibit comparably subpar performance on MMBench-Video, despite notable performance disparities on other benchmarks. For instance, VideoChat2 surpasses Video-ChatGPT by 18% on the MSVD-QA score (3.9 vs. 3.3), yet the performance gap narrows to just 6% on the MMBench-Video score (0.99 vs. 0.93). All video LLMs attain an average score close to 1 (out of a total of 3), with the top-performing model VideoStreaming reaching a mere 1.12. These findings suggest that the current state of video models’ proficiency in understanding MMBench-Video is nascent, underscoring the challenges and emphasizing the necessity for advancements in video LLMs to enhance their capability and effectiveness in interpreting varied video content.

Open-Source LVLMs for Images. A significant number of LVLMs have been developed to comprehend image content and execute visual reasoning tasks. During our evaluation, we focused on LVLMs that support the multi-image inference interface. We assess four prominent open-source LVLMs: Idefics2-8B , Qwen-VL-Chat , mPLUG-Owl2 , and InternVL-Chat-v1.5 , using MMBench-Video. To ascertain the models’ ability to effectively leverage multiple input frames, we evaluated them under two distinct settings: 1-frame and 8-frame inputs. Results in Tab. 3 indicate that all models, except Qwen-VL-Chat, exhibit a substantial enhancement in performance when utilizing 8 frames compared to a single frame. Notably, InternVL-Chat-v1.5 emerges as the top performer, achieving an impressive average score of 1.26 with 8 frames as inputs, significantly surpassing all other evaluated video LLMs.

Proprietary LVLMs for Images. Unlike their open-source counterparts, most proprietary LVLMs accept arbitrary interleaved images and text as input. However, the context window size still limits the maximum number of images. We evaluate several proprietary LVLMs, including Claude-3v, Gemini-Pro-v[1.0/1.5], GPT-4v, and GPT-4o on MMBench-Video with varying numbers of frames. Claude-3v struggles with 8-frame inputs and is only evaluated under the 4-frame setting. As anticipated, it exhibits the poorest performance among proprietary LVLMs when handling multiple frames. In contrast, other proprietary models demonstrate notably superior performance compared to the state-of-the-art open-source InternVL-Chat. Particularly impressive is GPT-4o, which, when processing 16 frames, achieves an outstanding overall score of 1.86. This result positioned GPT-4o 66% ahead of the best open-source video LLM and 48% ahead of the best open-source image LVLM.

2 Performance of Video-LLMs on Image VQA Benchmarks

Intuitively, a Video-LLM is expected to not only possess all capabilities of an image-based LVLM but also exhibit video-specific competencies, such as future prediction or causal reasoning. In light of the underwhelming performance of Video-LLMs on MMBench-Video, we broaden our evaluation to include image VQA benchmarks to determine if these models have the necessary skills for comprehending static content. We evaluate five Video-LLMs on two extensive image VQA benchmarks: MMBench and MMStar . To accommodate the input format of Video-LLMs, we create pseudo video clips by duplicating static frames, which then serves as the input for the evaluation. In Tab. 4, we list the performance of Video-LLMs alongside several representative image LVLMs for comparative analysis. On both benchmarks, existing Video-LLMs exhibit subpar performance. Notably, top-performing Video-LLMs such as PLLaVA and Video-LLaVA show performance that is either on par with or inferior to LLaVA-v1.5-7B, a rudimentary baseline for image multimodal understanding, and significantly trail behind the state-of-the-art image LVLM, InternVL-v1.5. This evaluation underscores the current limitations in the spatial understanding capabilities of Video-LLMs.

3 Incorporating Speech Further Improves Proprietary LVLMs

Video inherently comprises both visual and audio signals. However, the majority of existing LVLMs for video understanding predominantly focus on visual features, often neglecting the valuable information embedded in audio signals. To explore the potential impact of integrating audio features on video understanding, we conducted experiments using video title tracks (VTT) sourced from YouTube, which are automatically generated through speech recognition techniques. We incorporate these subtitles into the prompt as supplementary context. Experimental results in Tab. 5 reveal that the inclusion of audio/speech information enhances the performance of the state-of-the-art proprietary model, GPT-4o. The subtitles offer a rich source of high-density information, facilitating the LLM’s ability to accurately address the questions, thereby leading to a comprehensive performance improvement. Nonetheless, the increased information richness also heightens the risk of hallucinations, where the model may produce responses about non-existent content. The effectiveness hinges on the information density and the level of redundancy, necessitating a careful balance in applications.

4 The Superior Performance of GPT-4 as a Judge

Due to the discrepancy between the predictions of Video-LLM and the ground truth answers, existing VideoQA benchmarks largely rely on a judge model to assess the model responses. The capability of judge models can significantly influence the final results. To quantitatively study the impact of judge models, we utilize different versions of GPT-3.5 and GPT-4 for evaluation and report the results in Tab. 7 and Fig. 6. We observe that GPT-3.5 tends to assign high scores (typically 2 and 3), potentially leading to inaccuracies. To investigate the alignment with human preferences across different judge models, we conduct study based on a randomly selected subset of 100 questions. Two of the authors manually rate the responses from Video-LLaVA and GPT-4o, and we then report the mean absolute error between different judge models and the averaged human ratings. Tab. 7 shows that GPT-3.5 exhibits a significantly larger discrepancy with human preferences and greater inter-version variance.

Conclusion

This work introduces MMBench-Video, a novel long-form, multi-shot VideoQA benchmark specifically designed to evaluate the capabilities of LVLMs in understanding video content. MMBench-Video encompasses a diverse range of video topics and fine-grained capabilities. Extensive evaluations on MMBench-Video allow us to identify significant performance limitations among existing Video-LLMs in both spatial and temporal understanding.

Appendix A Additional Experiments

In the main paper, we report the quantitative results of speech improvement on the subset of videos with subtitles available from YouTube. In MMBench-Video, approximately half of videos do not include parseable video title tracks. In Tab. 8, we report the impact of incorporating speech information across the entire MMBench-Video dataset. While only half of the VideoQAs are enhanced with speech, the overall performance improvement on the full benchmark remains significant. It is evident that the enhancement in reasoning capabilities surpasses that of perceptual abilities. Speech typically conveys contextual information absent in static visual inputs, facilitating further reasoning by the VLM. Meanwhile, improvements in coarse perception are minimal or remain largely unchanged (for GPT-4o-[16f]). This can be attributable to the fact that perception is predominantly reliant on visual inputs, and speech information does not significantly augment the model’s performance in coarse perception.

A.2 Detailed Analysis of L-2 Capability

Based on Tab. 3, it is evident that hallucination is the most significant limitation in L-2 perceptual capabilities for all Video-LLMs, in contrast to the state-of-the-art proprietary LVLMs. This indicates that existing Video-LLMs are unable to dismiss questions pertaining to videos when uncertain and are inclined to generate answers for questions regarding non-existent visual content.

Regarding LVLMs, the number of frames significantly influences the performance of most L-2 capabilities in MMBench-Video. With an increase in the number of frames, the enhancement in perceptual capabilities becomes more pronounced than that in reasoning capabilities. Owing to a more extensive training corpus and superior safety mechanisms, proprietary LVLMs exhibit superior performance in challenging capabilities such as logical reasoning, commonsense reasoning, and hallucination.

Interestingly, despite Idefics2-8B-[1f] utilizing a single image as input, it still outperforms all Video-LLMs in temporal reasoning tasks. This suggests that Video-LLMs are not effectively leveraging diverse temporal information, underscoring the necessity to enhance the diversity of instruction tuning data for these models.

Appendix B Additional Dataset Analysis

In this section, we present more details about the MMBench-Video dataset: including the technique we adopted to filter temporal dispensable questions and some statistics on the linguistic characteristics of questions in MMBench-Video. In Figs. 8, 9 and 10, we display a selection of samples from MMBench-Video, showcasing videos, images, questions, and reference answers for illustrative purposes.

To ensure that the majority of questions in MMBench-Video are temporally indispensable, we implement an LVLM-based filtering process and subsequently conduct manual verification. Specifically, we employ GPT-4v, one of the most potent LVLMs, to filter out questions exhibiting high temporal irrelevance. Utilizing four distinct random seeds, we sample a single random frame as visual input for each individual VideoQA instance and conduct the inference four times. Subsequently, we utilize GPT-4 to evaluate the responses and compute the average score for each question. Questions with an average score of 2.5 or higher across the four responses were excluded from the benchmark. Based on this approach, we removed a total of 246 temporally dispensable questions from the dataset.

B.2 Question Type Analysis

Given that the majority of existing Video-QA benchmarks are characterized by a limited range of question types, which fail to adequately represent the diverse spectrum of human conversations, the question set within MMBench-Video has been meticulously curated to encompass a wide variety of categories. We visualize the comparison of the question type distribution between MMBench-Video and other popular benchmarks in Fig. 7. In addition to the conventional question archetypes, namely ‘what’, ‘who’, ‘how’, ‘when’, and ‘where’, MMBench-Video extends its corpus to include additional interrogatives such as ‘why’, ‘which’, ‘is / are’, and ‘does / do’. The expansion diversifies the dataset and closely aligns with the style of natural human dialogues. Meanwhile, the question type distribution within MSVD or MSRVTT exhibits a significant skew. The category of ‘what’ predominates, comprising over 60% of the questions, in stark contrast to the significantly underrepresented categories such as ‘when’ and ‘where’, totaling a mere 1%. Nevertheless, the question type distribution within MMBench-Video has been deliberately engineered to achieve a greater equilibrium. While the ‘what’ category maintains its status as the most prevalent, the remaining question types are evenly distributed across various interrogative forms.

Appendix C Prompts Adopted in MMBench-Video

In Sec. 3.1, we outline the LLM-involved evaluation paradigm we employ, utilizing GPT-4 for scoring. The evaluation is conducted through prompts configured with a 3-grade marking (0, 1, 2, 3). In this section, we elaborate on the specific prompts utilized in the evaluation process.

As an AI assistant, your task is to evaluate a candidate answer in comparison to a given correct answer. The question itself, the correct ’groundtruth’ answer, and the candidate answer will be provided to you. Your assessment should range from 0 to 3, based solely on the semantic similarity between the groundtruth and the candidate answer, disregarding any grammatical differences. A rating of 0 suggests no similarity, implying the candidate answer is entirely incorrect. A rating of 1 suggests low similarity, meaning the candidate answer is largely incorrect. A rating of 2 suggests high similarity, meaning the candidate answer is largely correct. Lastly, a rating of 3 indicates complete similarity, which means the candidate answer is entirely correct. Your response should be a single integer from 0, 1, 2, or 3. Question: [QUESTION] Groundtruth answer: [ANNOTATED ANSWER] Candidate answer: [CANDIDATE ANSWER] Your response:

C.2 System Prompt for the Inference of LVLMs with Multi-Frame Inputs

You will be provided with [FRAME NUM] separate frames uniformly sampled from a video, the frames are provided in chronological order of the video. Please analyze these images and provide the answer / answers to the following question / questions about the video content. If multiple questions are provided (with indices I1, I2, I3, ...), you should organize your answers in the following json format: {\{ ‘I1’: ’Answer to Question I1’, ‘I2’: ’Answer to Question I2’, ... }\} Otherwise, please directly reply with your response to the only question. Even if the information in these separate frames is not enough to give an answer, PLEASE TRY YOUR BEST TO GUESS A CLEAR OR VAGUE ANSWER WHICH YOU THINK WOULD BE THE MOST POSSIBLE ONE BASED ON THE QUESTION. Minimize negative responses such as ‘not possible to determine’. STIMULATE YOUR POTENTIAL AND IMAGINATION!

C.3 Input Prompt Template for the Inference of LVLMs with Multi-Frame Inputs

[System Prompt] [Subtitle (Optional): { ‘t0’ - ‘t1’: subtitle 1, ‘t1’ - ‘t2’: subtitle 2, ...... }] [Multi-Frame Inputs] [Question Set: { ‘index 1’: question 1 for this video, ‘index 2’: question 2 for this video, ...... }]

Appendix D Limitations and Broader Impacts

Limitations. In this study, we introduce MMBench-Video, a novel long-form multi-shot VideoQA benchmark, and perform a comprehensive evaluation based on this benchmark. In light of budget constraints, our evaluation is focused on a curated selection of representative open-source and proprietary VLMs, which may not encompass all those most recent high-performing models. GPT-4 is adopted as a more advanced judge model for scoring the responses, while further experiments should be conducted in the future to study the feasibility of using state-of-the-art open-source LLMs as the judge. Taking into account the limited capabilities of existing Video-LLMs, we currently set the upper duration limit of videos to 6 minutes, refraining from scaling to tens of minutes or hours. The evaluation results indicate that, even with relatively modest video durations, MMBench-Video presents a significant challenge to existing Video-LLMs.

Broader Impacts. As an evaluation benchmark, MMBench-Video offers detailed insights into the fine-grained capabilities of diverse vision-language models (VLMs) in the domain of video understanding, providing valuable insights for future model optimization. The new benchmark exhibits enhanced quality and enriched diversity, and employs a more precise scoring strategy, which collectively contribute to comprehensive and reliable evaluation outcomes. Leveraging MMBench-Video and other Image VQA benchmarks, we conduct a comprehensive evaluation of existing video-LLMs, revealing their limited capabilities in both spatial and temporal understanding. Additionally, MMBench-Video, being a small-scale benchmark, may not encompass every video topic and fine-grained capability. There is a risk that MMBench-Video may not adequately reflect the video understanding capabilities of VLMs in specific tasks or scenarios. We encourage users to carefully consider the intended use cases of VLMs when utilizing MMBench-Video for evaluation.

References