CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Guo Chen, Yicheng Liu, Yifei Huang, Yuping He, Baoqi Pei, Jilan Xu, Yali Wang, Tong Lu, Limin Wang

Introduction

Recently, video understanding has made significant progress with the advent of multimodal large language models (MLLMs). To evaluate these models, many recent efforts have been made to create video understanding benchmarks (Li et al., 2023b; Mangalam et al., 2024; Liu et al., 2024e), providing assessments of model comprehension capabilities and clues for future improvement.

Since early benchmarks only focus on short video clips, recent works have started to create benchmarks (Fu et al., 2024a; Wu et al., 2024b; Zhou et al., 2024; Huang et al., 2024) for longer videos (≥\geq 10 minutes). However, these works employ multiple-choice questions (MCQ), where the difficulty level is heavily influenced by the configuration of negative options. In such scenarios, models (Chen et al., 2023; Li et al., 2024; Zhang et al., 2024b; Lin et al., 2024) tend to focus on only general video knowledge and use elimination to avoid selecting the negative options. As a result, the models can achieve correct answers without genuinely engaging with the relevant video content, leading to a lack of trustworthiness. One illustration can be found in question 2 of Figure 1, the option ‘A’ can be easily eliminated based purely on textual information. Recently, the NExT-GQA (Xiao et al., 2024) benchmark tries to address the problem of credible models by incorporating temporal grounding into MCQ. However, NExT-GQA is limited to the NextQA (Xiao et al., 2021) dataset, which lacks diversity and primarily consists of short videos. A comprehensive benchmark for credibly evaluating generalist MLLMs for long video understanding, is still missing in the research community.

To make up this gap, we introduce CG-Bench, illustrated in Figure 1, a novel benchmark designed to evaluate clue-grounded question answering in long videos. In contrast to traditional benchmarks that focus primarily on the accuracy of question answering, CG-Bench goes a step further by evaluating whether the model bases its answers on relevant clues within the video. CG-Bench designs two novel clue-based evaluation methods to provide more reliable model performance assessments. 1) clue-grouded white box evaluation requires the model to directly provide the clue interval corresponding to the question while selecting the correct answer. 2) clue-grouded black box evaluation requires the model to align the accuracy of video-level MCQ and clue-level MCQ. Furthermore, we propose a novel heuristic method, aided by human-annotated clues, for open-ended QA evaluation, to effectively balance the cost and performance.

CG-Bench features 1,219 meticulously curated videos and 12,129 human-annotated question-answer-clue (QAC) triplets, establishing it as the largest and held-out VideoQA and question grounding benchmark for long videos. It employs a highly detailed manual classification system, organizing each video into 14 primary categories, 171 secondary categories, and 638 tertiary categories. The benchmark includes three main question types: perception, reasoning, and hallucination. Perception questions are further divided into 10 subcategories, such as object and attribute recognition, while reasoning questions are categorized into 12 subcategories, including relation reasoning, etc.

We evaluate a range of closed-source and open-source MLLMs using this benchmark. The commercial models, GPT-4o (OpenAI, 2024) and Gemini-1.5 Pro (Anil et al., 2023) achieve scores of 53.9 and 43.4, respectively, with 128 frames for long-video multiple-choice questions. The leading open-source MLLM, Qwen2-VL-72B (Wang et al., 2024a), scores 51.4 under the same conditions, indicating its initial benchmarking against GPT-4o. However, our credibility assessments and open-ended evaluations reveal a significant drop in accuracy for existing MLLMs, with scores decreasing from 53.9 to 21.7. This underscores the considerable room for improvement in current MLLMs for long video understanding. We hope this benchmark can become a vital tool for advancing research and development of more reliable and capable MLLMs.

Related Work

Multimodal Large Language Models (MLLMs) have rapidly gained popularity due to their proficiency in integrating visual and textual information (Liu et al., 2024a; 2023; Chen et al., 2023; Wang et al., 2022; 2024c). Recent advancements, such as LLaVA-Next-Video (Zhang et al., 2024b), LLaVA-OneVision (Li et al., 2024), and InternVL2 (Chen et al., 2024d), focus on enhancing MLLMs by integrating LLM backbones with visual encoders and specialized adapters, or creating higher-quality multimodal instruction data. This results in improved performance across tasks that involve both text and images.

Another area of focus is multimodal video understanding. Most models (Chen et al., 2024d; Li et al., 2023a; Maaz et al., 2023; Pei et al., 2024; Huang et al., 2018) are optimized for short videos, typically a few seconds or at most a few minutes, without exploring their visual understanding with longer context. In response, researchers have explored methods such as compressing video frames into fewer visual tokens to allow for the handling of longer videos, as seen in models like LLaMA-Vid (Li et al., 2023c), LVChat (Wang et al., 2024d), MovieChat (Song et al., 2024), MA-LMM (He et al., 2024) and Oryx (Liu et al., 2024f). In addition, LongVA (Zhang et al., 2024a) and LongViLA (Xue et al., 2024) explore the system-level optimization for long-context MLLMs which can natively support long video understanding. Despite the continuous proposal of various MLLMs, their real-world performance in long video understanding is still under explored.

MLLM Benchmarks. The development of benchmarks is becoming increasingly essential, especially for evaluating the MLLM performance in video understanding tasks. As the field develops, various benchmarks have been established to assess MLLMs across different modalities and video lengths. Previous efforts primarily focused on short videos, with traditional specialized VideoQA datasets like TVQA (Lei et al., 2018), NextQA (Xiao et al., 2021), and benchmarks for MLLM like VideoBench (Ning et al., 2023), MVBench (Li et al., 2023b) and EgoSchema (Mangalam et al., 2024). MVBench provides a comprehensive framework for evaluating general temporal understanding capabilities through question-answering on short clips, while EgoSchema focuses on egocentric video understanding with multi-choice questions. The videos in these benchmarks typically range from a few seconds to several tens of seconds, making them similar to image benchmarks and thus hindering the development of general video LLMs.

Recently, several works such as VideoMME (Fu et al., 2024a), CinePile (Rawal et al., 2024), MLVU (Zhou et al., 2024), LongVideoBench (Wu et al., 2024b), MoVQA Zhang et al. (2023b), and LVBench (Wang et al., 2024b), have introduced long video benchmarks to evaluate MLLMs. VideoMME constructs a diverse video MCQ dataset, incorporating multimodal evaluations with visuals, subtitles, and audio. MLVU designs a range of tasks that focus on granular detail understanding to assess long video comprehension capabilities. However, a common limitation of these benchmarks is their reliance on MCQs, where the difficulty is heavily influenced by the construction of negative options. This allows MLLMs to often eliminate incorrect answers using sparse frames and common sense reasoning, which can inflate performances. With our clue interval annotation, CG-Bench enhances the evaluation quality of MLLMs in long video understanding by introducing new evaluation mechanisms on credibility.

CG-Bench

The dataset construction process of CG-Bench consists of three steps: video collection, question-answering-clue annotation, and quality review iteration. We provide details as follows.

Video Collection. To avoid using videos that have been used for pre-training by existing MLLMs, we manually collect videos from the internet and provide new annotations on them. To facilitate the collection of raw videos from the Internet, we define 14 root domains as listed in Figure 3. During the collection process, we manually assign a brief tag (4-8 words) to categorize the content of each video. This supplementary tagging helps to ensure the diversity of the videos. We define a video to be long if it exceeds 10 minutes in duration. Accordingly, we collected videos longer than 10 minutes while considering the distribution of video duration. Furthermore, we retain the accompanying subtitles and audio to provide multimodal information. We carefully review and filter the videos manually for 7 rounds. More details about the video collection can be found in the supplementary material.

Question-Answer-Clues Annotation. After collecting the raw video data, we annotate it with high-quality question-answer-clue (QAC) triplets. To ensure question diversity, we first establish a taxonomy with three main types: Perception, Reasoning, and Hallucination. As shown in Figure 3, Perception and Reasoning questions are further divided into 10 and 14 subcategories, respectively, while Hallucination questions combine elements of both perception and reasoning. Annotators are instructed to include negative options to create a multiple-choice QA format, facilitating straightforward and cost-effective assessments. To minimize expression loss, annotators use their native language during the annotation process. Each video is annotated 6 to 15 QAC triplets depending on its duration. To ensure consistency in QAC triplets, we standardized the annotation process by first annotating the QA pairs and then identifying the clues. Annotators must watch the entire video, select a question type from the predefined categories, and then annotate a new question and its corresponding answer. Next, they select one or more intervals from the video to form a QAC triplet. Since the actual clue intervals often consist of multiple short moments, annotating each fragment is costly. Therefore, annotators are required to mark intervals that cover these short moments while ensuring the completeness of each event.

Review Iteration. To ensure the difficulty and quality of the dataset, we conduct a repetitive review and iteration process to enhance annotation quality. We reject annotations that do not meet our quality standards and request annotators to revise them. Our quality requirements for annotations and the measures taken to ensure them are as follows: 1) The rationality of the question, options, and answer: we conduct manual reviews; 2) The video dependency of the question, options, and answer: we input questions and options into GPT-4 and filter out QA pairs that can be answered solely based on pure text; 3) The difficulty of negative options in multiple-choice questions: we input the video, questions and options into MLLMs and filter out QA pairs that can be answered using only sparse frames and small models; 4) The positional diversity of clue intervals: We monitor the distribution of clue duration and position and provide timely guidance to annotators.

2 Dataset Statistics & Comparisons

We present the detailed statistics of our dataset to provide a more comprehensive understanding, including meta-information, QAC triplets, qualitative analysis, and comparison to previous works.

Video Meta. Our dataset comprises a total of 1219 videos with multimodal information, including vision, audio, and subtitles. The duration of the videos varies between 10 and 80 minutes, with a distribution illustrated in Figure 5. Notably, videos that last between 20 and 30 minutes are the most prevalent. This selection process is manual, based on content relevance, which mirrors real-world duration distributions and highlights a long-tail effect for longer videos. As illustrated in Figure 3, each video is classified using a three-tiered tagging system that succinctly encapsulates its content and assigns it to fundamental categories. The primary classification is augmented by a secondary layer of 171 tags and a tertiary layer consisting of 638 tags. This multi-level tagging mechanism guarantees a broad diversity of data content. For a more detailed classification of tags, please consult the supplementary materials.

QAC Annotation. CG-Bench includes 12,129 annotations consisting of questions, answers, and clues. Table 1 presents the sentence lengths and totals for the annotated questions and answers, highlighting the linguistic diversity within our dataset. Each QAC triplet is annotated with 4 to 7 negative samples, resulting in an approximately uniform distribution with ratios of options A to H of 12.4%, 14.7%, 12.1%, 14.8%, 15.1%, 16.1%, 11.6%, and 3.1%. There are a total of 14,362 clue intervals across all QAC triplets, with an average duration of 19.24 seconds each. Additionally, we conduct a further analysis of the positions of clue intervals within the video. Figure 5 illustrates the frequency with which each normalized timestamp is represented by intervals. This demonstrates the unbiased nature of our interval annotations and highlights the diversity of our QA content in temporal position.

2.2 Comparison with Previous Benchmarks

CG-Bench is characterized by its diverse features, allowing it to be compared with three distinct types of benchmarks, as depicted in the three sections of Table 2: Question Clue Grounding, Short-Video QA, and Long-Video QA benchmarks. For the question clue grounding benchmarks, NextGQA (Xiao et al., 2024), Ego4D-NLQ (Grauman et al., 2022), MultiHop-EgoQA (Chen et al., 2024c), E.T. Bench (Liu et al., 2024d), and RexTime (Chen et al., 2024a) are primarily centered around action and egocentric domains. Their videos are sampled from academic datasets. In comparison, the question clue grounding part of CG-Bench, CG-Bench-QG, stands out with the highest number of videos and the longest average length, the diversity of which fosters a broad spectrum of question-grounding queries.

Furthermore, we transform QAC triplets to our novel Short-Video QA benchmark, termed CG-Bench-Clue. When contrasted with prior short video benchmarks such as TempCompass (Liu et al., 2024e), MVBench (Li et al., 2023b) and MMBench-Video (Fang et al., 2024), our CG-Bench-Clue emerges as the largest, held-out, open-domain and multimodal Short-Video QA benchmark.

As for the Long-Video QA benchmark, CG-Bench excels in the number of videos, length, quantity of questions, and annotation quality. Owing to our clue interval annotations, CG-Bench further facilitates reliable evaluations for long videos and open-ended evaluations with clue assistance, a feature that sets it apart from existing long video benchmarks like Video-MME (Fu et al., 2024a) and MLVU (Zhou et al., 2024).

3 Evaluation

In this section, we describe the evaluation tasks of our CG-Bench which include traditional MCQ, the unique credibility evaluation, and clue-aided open-ended QA evaluation.

We assess the accuracy of MCQ in two settings: Long-Video MCQ and Clue-based MCQ. In the Long-Video MCQ setting, the model receives the entire video as input and is required to select the correct answer based on the video, the question, and the candidate options. For the Clue-based MCQ setting, the model is given only the video within the annotated clue interval as input. The model has access only to the clue clip, the question, and the candidate options. It does not have access to the original long video. Since a single QA may correspond to multiple clues, we merge these clues and treat the combined clue as a single, cohesive clue segment.

3.2 Credibility Evaluation

The ability of a model to identify relevant clues related to questions is a crucial factor in determining its reliability. Therefore, we define a model’s reliability based on its proficiency in locating accurate clues when addressing problems. To achieve this, we introduce two clue-grounded mechanisms for credibility assessment: white-box evaluation and black-box evaluation.

White-Box Evaluation requires the model to directly output the intervals of clues that can accurately answer the question. This task is similar to video temporal grounding (Lei et al., 2021; Huang et al., 2023). Therefore, we use tIoU (Temporal Intersection over Union) as the evaluation metric. Since each question may correspond to multiple intervals of clues, we allow the model to predict multiple possible intervals. Given a set of prediction P\mathcal{P} and ground truths G\mathcal{G}, the tIoU is defined as:

where aia_{i}, bib_{i} are the start and end timestamps of the ii-th ground truth interval of G\mathcal{G}. cjc_{j}, djd_{j} are the start and end timestamps of the jj-th predicted interval of P\mathcal{P}. We calculate the mean IoU (mIoU) by averaging the tIoU scores obtained by the model across all question queries. To further improve the robustness of question grounding evaluation, we introduce the rec.@IoU metric. This metric measures the probability of successfully recalling clue intervals at various IoU thresholds. We calculate the average recall rate at IoU thresholds of 0.1, 0.2, 0.3, 0.4, and 0.5 to determine the final result.

In addition, we propose a combined metric, acc.@IoU that evaluates both MCQ accuracy and clue-grounding ability. For a question with multiple choice options, the response is considered correct only if the selected answer is accurate and the tIoU between the prediction and the ground truth exceeds (>>) a predefined threshold τ\tau. Since locating short-duration clues in the long videos in CG-Bench is inherently challenging, we set the default τ\tau to be 0 for the more obvious comparison on ablation studies. Setting τ=0\tau=0 ensures that acc.@IoU requires the model to select the correct option and produce a time interval that overlaps at least slightly (tIoU>0\text{tIoU}>0) with the annotated clue interval, rather than reducing to naive MCQ accuracy. We calculate the acc.@IoU at IoU thresholds of 0.1, 0.2, 0.3, 0.4, and 0.5 to determine the final result.

Black-Box Evaluation aims to evaluate the model’s ability to seek out clues implicitly. Understanding long videos involves the retrieval of clues distributed across various spatiotemporal locations within the entire video. Therefore, an effective model for long videos should naturally focus on capturing human-annotated clue intervals in its hidden states. However, beyond the explicitly annotated clue intervals, there are likely hidden clues scattered throughout the video that can also help to determine the correct answer. Thus, a model with access to the full video should yield higher accuracy compared to solely relying on the clue interval. In other words, the accuracy of Long-Video MCQ ( long-acc.) should be greater than or equal to the accuracy of Clue-based MCQ (clue-acc.).

With this insight, for the black box evaluation, we define a new metric called Clue Recovery Rate (CRR). This metric evaluates the model’s robustness to context dilution, i.e., how stable a model can find related clues from long but diluted video context. CRR is calculated by:

A CRR of less than 100% suggests that the MLLM’s ability to retrieve short clues from long video representations is not optimal.

3.3 Clue-aided Open-Ended QA Evaluation

Finally, CG-Bench supports open-ended QA evaluation for more comprehensive assessment results. Previous works such as MM-Vet (Yu et al., 2023) and MMBench-Video (Fang et al., 2024) use LLMs to evaluate open-ended QA for images and short videos. In contrast, long videos typically contain more complex information, thus user-generated questions tend to be ambiguous. As a result, the correct answer can be in many distinct forms, which can cause discrepancies between the LLM-evaluated score and the real model’s QA ability, as shown in Figure 6.

To address this, we leverage a low-hallucination MLLM to evaluate the similarity between the text output and the visual information. We choose GPT4o (OpenAI, 2024) as the multimodal evaluator because it ranks among the top in several well-known benchmarks, such as OpenCompass (Contributors, 2023), Lmsys leaderboard (Chiang et al., 2024), etc., and shows relatively lower hallucinations than other MLLMs. Since directly using GPT-4o for multimodal judging can still introduce hallucination errors and incur high API costs, we propose a heuristic evaluation method to mitigate biases and reduce costs. First, GPT-4o assesses whether the output can be evaluated based solely on the text answer. If feasible, it outputs either yes or no; otherwise, it requests visual cues by stating “I need visual clues”. This prompts the inclusion of supplementary visual data in the prompt to aid GPT-4o in its evaluation process. By using pre-annotated time intervals with question clues, we sample frames as visual aids, further reducing hallucination errors and costs. We quantitatively analyze this evaluation method in Sec 4.3. More details can be found in the supplementary materials.

Experiments

In this section, we evaluate a wide range of MLLMs using CG-Bench. We first introduce the evaluation setup, followed by quantitative results for both closed-source and open-source models. Finally, we analyze some key factors in the evaluation.

We first briefly describe the settings used in our experiments. The supplementary material provides more detailed settings.

Models. We evaluate the performance of three mainstream commercial models on our CG-Bench: GPT4o (OpenAI, 2024), Gemini-1.5 (Anil et al., 2023), and Claude-3.5, including their different versions. Also, we assess the representative open-source video models such as LLaVA-OneVision (Li et al., 2024), Qwen2-VL (Wang et al., 2024a) and InternVL2 (Chen et al., 2024d), among others.

Frame Sampling. For long video understanding, the frame sampling strategy significantly impacts evaluation results. For open-source MLLMs, we make the best use of our computational resources to use as many frames as possible. For closed-source MLLMs, since the local computational resource is no longer a bottleneck, we can use even more frames. We uniformly sample 128 frames for Long-video MCQ, and use 32 frames as the for Clue-based MCQ.

Modality. We also explore other modalities: subtitles and audio. For subtitles, we employ a uniform sampling method. If the timestamp of a sampled frame falls within the time interval of a subtitle, that subtitle will be included in the analysis. Each subtitle is considered only once to avoid redundancy.

Prompt. For MCQ tasks, the model is prompted to provide the uppercase letter corresponding to the correct option. In Open-Ended QA tasks, the model responds freely based on the questions. For the Clue Grounding task, we append the timestamps of each frame and subtitle to enhance the model’s time-awareness, requiring it to return nested lists in the format [[s1, e1], [s2, e2], ...]. For open-ended evaluation, we require the model to assess the correctness between the predictions and the ground truth and respond with yes or no.

2 Main Results

As shown in Table 3, the closed-source MLLM GPT4o (OpenAI, 2024) achieved a significant overall lead, surpassing other MLLMs across all metrics. Notably, GPT4o’s long-acc. approaches 45.2%, significantly higher than Gemini-1.5-Pro (Anil et al., 2023), highlighting its strong capabilities in long video understanding. For open-source MLLMs, Qwen2-VL’s (Wang et al., 2024a) performance is undeniably impressive, achieving comparable results to GPT4o on long-acc. and clue-acc.. Other models achieve sub-optimal performance due to the lack of supporting enough context or sufficient training on videos. Although these MLLMs achieve relatively high accuracy on the MCQ task, they all experienced significant performance degradation when subjected to credibility and open-ended evaluation of CG-Bench. For example, GPT-4o’s long-acc. dropped from 45.2 to 4.38 in Acc@IoU and 39.5 in OE-acc.. Notably, with the same number of sampling frames, GPT-4o achieves a CRR of 77.5, while Gemini1.5-Pro only obtains 74.3. This indicates that Gemini-1.5-Pro has an inferior ability to retrieve short-term clues from long videos. Overall, the current MLLMs do not perform well on our CG-Bench, suggesting that there is still considerable room for improvement in their capability and credibility.

Since it is difficult to input more than 128 frames due to the hardware limitations, we alternatively conducted a human evaluation experiment under constrained visual conditions, to see how severe the “undersampling” issue is for longer video. We uniformly sampled 30 videos from CG-Bench, resulting in 296 questions. For each video, we uniformly sampled 128 frames and asked volunteers to perform an MCQ testing. The resulting accuracy was 59.85% (row 3 in Table 3). This result indicates that our dataset is indeed challenging and that it is difficult to derive solutions from a limited number of frames. It also highlights that even the most advanced models, such as GPT-4o, have ample room for improvement in long video comprehension.

3 Analysis

Furthermore, we perform a comprehensive analysis of the two leading closed-source MLLMs, GPT4o (OpenAI, 2024) and Gemini-1.5 Pro (Anil et al., 2023), as well as the best performing open-source MLLM, Qwen2-VL (Wang et al., 2024a), on our CG-Bench. In this analysis, we use 1000 QAC triplets sampled uniformly from all annotations for fast experiments. We report acc.@IoU with τ=0\tau=0 for a more obvious comparison.

Impact of Prompt & Modality. As shown in Table 4, we conduct the ablation studies on the subset that contains subtitles and explore the impact of different prompts on GPT4o and the effect of the audio modality on Gemini-1.5 Pro. Our findings indicate that all prompt types (FT/S/ST), except video frames (F), provide performance benefits across most metrics. Subtitles contribute more to long-acc. than they do to clue-acc.. Additionally, the inclusion of timestamp information (FT/ST) is critical for interval prediction. Timestamps from both frames and subtitles enhance IoU-related metrics, revealing a complementary effect. When both FT and ST are added simultaneously, mIoU increases from 3.39 to 9.68, and Acc@IoU rises from 10.7 to 26.7. When S, FT, and ST are all used in the prompt, the model achieves the best performance across all metrics. In contrast, our exploration of the audio modality (A) revealed that audio does not yield significant performance gain and, in some cases, even slightly degrades the results, as shown in Table 4. Finally, we conduct experiments using only subtitles from 128 frames versus the full video. The results show that while subtitles offer useful semantic cues, their impact is significantly reduced when visual input is included. This suggests that our benchmark favors visual signals.

Impact of Frame Number. As illustrated in Figure 7, we conducted experiments to analyze the performance across various metrics as the number of frames increases. Overall, the performance of all three MLLMs gradually improves with the addition of more frames, with GPT-4o consistently outperforming the others across all metrics. For long-acc. and OE acc., Qwen2VL achieves performance comparable to GPT-4o. However, compared with Qwen2VL, Gemini excels in terms of mIoU and Acc@IoU. Regarding CRR, GPT-4o demonstrates greater consistency between clue-acc. and long-acc. across more frames, indicating its superior reliability in long video understanding. For open-ended QA, Gemini’s higher refusal rate results in a noticeable decline in performance.

Open-ended Evaluation Quality. To assess the stability and accuracy of various MLLMs as evaluators, we utilized four models—Gemini, Qwen2VL, Claude, and GPT-4o—each of which evaluated GPT-4o’s predictions five times. Human evaluations of GPT-4o’s predictions are also conducted for reference. The results, shown in Figure 8, indicate that GPT-4o has the highest stability and the smallest deviation from human-assigned scores. Furthermore, Table 5 explores the impact of different evaluation methods. When evaluators were provided only with ground truth (col. “GT”) or visual information (col. “Vis”), the scoring bias (absolute difference) between human and model-based evaluation increased. While fully leveraging visual information (col. “GT+Vis”) improved evaluation accuracy, it also significantly increased the time and cost required. Our proposed heuristic evaluation method achieves the lowest evaluation bias. Additionally, we manually annotated 200 evaluation samples to determine the necessity of visual request triggers. From the bottom block in Table 5, the statistics show that our method achieved a visual request trigger rate (the probability that the model triggers “visual clues required”) of 14%. The recall rate of this triggering achieves 88%. This proves that our approach effectively balances cost and performance.

Performance grouped by Video Duration. We grouped videos by duration and evaluated the long-acc. performance of GPT-4o-0806 using 128 frames. Figure 9 shows that the model struggles with undersampling, especially for longer videos.

Impact of Frame Sampling Strategy. We investigate how different frame sampling strategies affect performance. To expedite testing, we primarily evaluated GPT4o-0806 using 50 uniformly sampled frames, focusing on the long-acc metric. The experiment consists of three parts: 1) low resolution, 2) high resolution, and 3) keyframe extraction (via FFmpeg) combined with low resolution. As shown in Table 6, higher resolution offers some improvement, while keyframe extraction has no significant impact.

Conclusion and Future Work

In this paper, we introduce CG-Bench, a novel benchmark designed to evaluate clue-grounded question answering capabilities in long video understanding. Unlike existing benchmarks that focus on short videos or rely solely on multiple-choice questions, CG-Bench emphasizes the importance of models retrieving and grounding their answers in specific video segments, enhancing evaluation credibility. By incorporating 1,219 manually curated videos categorized into a detailed three-tier system and 12,129 QA pairs spanning perception, reasoning, and hallucination question types, CG-Bench offers a comprehensive and diverse dataset for assessing MLLMs. Our two proposed clue-based evaluation methods—clue-grounded white-box and black-box evaluations—provide novel ways to assess whether models genuinely comprehend video content or merely rely on superficial cues. Through extensive experiments involving various closed-source and open-source MLLMs, we found that current models significantly underperform in long video understanding compared to short videos. We hope that CG-Bench will serve as a valuable resource for the research community, driving the development of more trustworthy and capable MLLMs for long video understanding.

References