VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
Yuxuan Wang, Yueqian Wang, Dongyan Zhao, Cihang Xie, Zilong Zheng
Introduction
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and language generation (Alayrac et al., 2022; Li et al., 2023a; Liu et al., 2023; OpenAI et al., 2024). However, despite their strong performance on standard benchmarks (Antol et al., 2015; Lin et al., 2014; Xu et al., 2017, 2016; Krishna et al., 2017), these models frequently produce incorrect or unsubstantiated responses w.r.t. visual inputs (Li et al., 2023b; Tong et al., 2024; Petryk et al., 2024). This issue, often referred to as “hallucination” (Rohrbach et al., 2018), means that MLLMs can generate irrelevant or nonsensical content that deviates from the original visual context. Given these challenges, a natural question arises: How can we examine the vulnerability to hallucinations in MLLMs? Addressing this need not only reveals the extent of hallucination in these models, but also helps identify underlying causes and develop methods to further enhance models.
To address this need, researchers have begun to benchmark hallucinations related to objects (Rohrbach et al., 2018; Li et al., 2023b), relationships (Han et al., 2024), and attributes (Sun et al., 2023; Wang et al., 2024a; Chen et al., 2023a; Wang et al., 2024b), as well as factual information (Cui et al., 2023; Jing et al., 2023; Chen et al., 2024). However, most existing studies focus on hallucinations involving basic static visual attributes from images in large image-language models. They often overlook potential hallucination issues arising from dynamic content, such as actions, events, and stories, in large video-language models (LVLMs). Furthermore, video-language tasks, such as video summarization (Ghauri et al., 2020; Song et al., 2015), due to their higher complexity, are in lack of careful evaluation in existing datasets. As a result, we are still unsure of the extent, severity, characteristics, and causes of hallucination issues within LVLMs. Additionally, existing benchmarks focus on certain attributes of hallucinations, lacking a comprehensive and robust evaluation.
To tackle these issues, we introduce VideoHallucer, the first comprehensive benchmark designed to assess hallucination in LVLMs. Within VideoHallucer, we establish a clear taxonomy of hallucinations, distinguishing between two primary categories: intrinsic and extrinsic. Specifically, Intrinsic hallucinations involve generated content that directly contradicts information present in the source video, and can be categorised into three subtypes: object-relation, temporal, and semantic detail hallucinations. While extrinsic hallucinations involve content that cannot be verified from the source, and can be classified as either extrinsic factual, aligning with general knowledge but not present in the source video, or extrinsic non-factual, which includes all the others. To avoid confounding factors from LLMs (Li et al., 2023b; Zhang et al., 2023a), our benchmark focuses on identifying hallucinations in video-language grounding using a binary VQA-based method (Li et al., 2023b; Cui et al., 2023; Chen et al., 2023a), Specifically, we introduce an adversarial evaluation(Tong et al., 2024) with paired questions—one basic and one intentionally hallucinated—to rigorously test models. Moreover, we balance ’yes’ and ’no’ responses to reduce language biases and provide clear explanations to minimize misinterpretation. Comparisons with existing multimodal hallucination benchmark are discussed in Table 1 and Sec. 2.
By comprehensively evaluating twelve LVLMs on VideoHallucer, our analysis leads to three significant insights: First, our analysis revealed a widespread issue of hallucinations across LVLMs, and more critically, the performance gaps between humans and models in all VideoHallucer settings are significant (Sec. 4.2). Second, we confirm the benefits brought by scaling on hallucination — increasing the size of the training dataset or/and the model’s parameters improves hallucination detection related to basic visual cues and counterfactuals. However, we found that this approach has a limited impact on the models’ ability to detect extrinsic factual hallucinations (Sec. 4.2). Third, we observed that current models are more proficient at recognizing facts than detecting hallucinations, where the latter requires the models to discern facts within the context of the source material (Sec. 5.1). We also found that these shortcomings in recognizing extrinsic factual hallucinations could be partially mitigated by implementing high-quality explanatory mechanisms (Sec. 5.2).
Building on the above observations, in Sec. 5, we devise Self-PEP, a plug-and-play framework that bolsters the self-improvement capabilities of models through the integration of explanatory processes. By applying Self-PEP, most models demonstrated enhanced performance on the VideoHallucer benchmark, with an average improvement of 5.38%. We believe this framework can streamline future research and development in detecting and mitigating hallucinations in video.
Related Work
Generative models in NLP, particularly Large Language Models (LLMs), have shown remarkable proficiency across various language generation tasks. Despite their capabilities, a significant issue persists: the text they produce can sometimes be irrelevant or nonsensical. This issue is known in the field of NLP as “hallucination.” Hallucination refers to instances where the content generated by these models does not make sense or deviates from the intended meaning of the source material it is based on (Filippova, 2020; Maynez et al., 2020; Parikh et al., 2020; Zhou et al., 2021; Ji et al., 2023). The concept of hallucination in NLG tasks can vary slightly, but it is generally categorized into two types based on the relationship between the generated content and the source material. These types are known as Intrinsic Hallucination and Extrinsic Hallucination (Dziri et al., 2021; Huang et al., 2021; Maynez et al., 2020; Ji et al., 2023). Intrinsic hallucination refers to instances where the content generated by a model directly contradicts the information provided in the source material. On the other hand, extrinsic hallucination occurs when the generated content cannot be confirmed or refuted by the source material. In such cases, the content may either align with or contradict extrinsic knowledge. Building on this, Cao et al. (2022) further categorize extrinsic hallucinations into two types: factual and non-factual. Factual hallucinations produce content that can be verified against real-world knowledge, while non-factual hallucinations generate content that is at odds with what is known about the world. Within specialized research domains, there is a divergence of opinion regarding the value of factual hallucinations. Some studies (Maynez et al., 2020; Thomson and Reiter, 2020), suggest that factual hallucinations can be beneficial. They argue that the additional knowledge introduced by these hallucinations can enhance the informational quality of the generated content.
Advancements in image-language modeling, particularly in generative tasks such as image captioning and image-based question answering, have led researchers to investigate the phenomenon of hallucination in this field. The concept of object hallucination in image captioning was first introduced by (Rohrbach et al., 2018). They highlighted the issue of captions erroneously including objects that are absent from the images. To address this, they developed the CHAIR metric, an automatic evaluation tool designed to measure the accuracy of object references in captions by calculating the precision of the hallucinated objects. Following this development, numerous studies have begun to explore hallucination within image-language models. However, the CHAIR metric itself has come under scrutiny. Li et al. (2023b) identified certain limitations in CHAIR, noting that the metric’s results could be skewed by the way instructions are designed and that the reliance on human-crafted parsing rules could severely restrict its applicability. To overcome these challenges, POPE introduced a binary VQA benchmark specifically tailored for the detection of object hallucination, aiming to provide a more robust and reliable means of evaluation in this area. Recent advancements have been made in the development of evaluation toolkits for object hallucination with the aid of LLMs, as evidenced by research (Wang et al., 2023; Zhai et al., 2023; Petryk et al., 2024; Qiu et al., 2024). Concurrently, there is a growing body of work aimed at expanding the concept of object hallucination to include additional visual features. (Sun et al., 2023; Wang et al., 2024a; Chen et al., 2023a; Wang et al., 2024b; Han et al., 2024) have begun to explore the relationships, attributions, and other visual cues, including counting, OCR and etc. Moreover, (Jing et al., 2023; Guan et al., 2024; Cui et al., 2023; Chen et al., 2024; Liu et al., 2024) are pioneering the detection of factual hallucinations within image-language understanding models, marking a significant step forward in the field. Despite the availability of benchmarks for image-language models, there is a notable lack of a clear and comprehensive framework specifically designed to evaluate LVLMs. To address this gap, our work introduces a thorough benchmark dedicated to assessing the performance of LVLMs, with a particular focus on their susceptibility to hallucination. Moreover, Table 1 presents a comparison of VideoHallucer with other existing datasets designed for the detection of hallucinations in vision-language models. VideoHallucer stands out as the first and most extensive benchmark specifically tailored for LVLM hallucination detection.
The 𝒱ideoℋallucer𝒱𝑖𝑑𝑒𝑜ℋ𝑎𝑙𝑙𝑢𝑐𝑒𝑟\mathcal{V}ideo\mathcal{H}allucer Benchmark
We will provide a detailed introduction to VideoHallucer in the following sections. Sec. 3.1 will introduce the construction process, including the different types of questions. Then, in Sec. 3.2, we will display the statistical information of VideoHallucer. Finally, in Sec. 3.3, we will demonstrate how we evaluate LVLMs on VideoHallucer.
To evaluate hallucination issues in detail, we split our benchmark into intrinsic and extrinsic categories, resulting in five settings. In the following section, we introduce how we construct question-answer pairs for each of these five settings separately. The detailed annotation procedure is discussed in Sec. A.2.
Our work creates the object-relation hallucination setting that concentrates on the objects and their interactions over time, which we organize into three categories: subject, relation, and object. We construct this setting from existing datasets VidOR (Shang et al., 2019; Thomee et al., 2016) and VidVRD (Shang et al., 2017). We follow Li et al. (2023b) and use templates to generate basic questions about objects and their relations. Annotators then generate semantically distinct yet visually analogous alternatives for these questions to identify hallucinated content. Through this semi-automated annotation process, we produced question-answer pairs across videos.
To benchmark the temporal hallucination issue in LVLMs, we design a setting to evaluate models’ hallucination issues from three dimensions: absolute temporal, relative temporal, and event length. Utilizing the ActivityNet dataset (Krishna et al., 2017), we create question-answer pairs spanning videos. For absolute temporal, we select events located in the first or last of videos and asking if they occur at the beginning or end. For relative temporal, we choose pairs of events with clear temporal separation and ask which occurs first. For event length, we compare the durations of pairs of events, questioning which is longer.
Recent research (Sun et al., 2023; Wang et al., 2024a; Chen et al., 2023a; Wang et al., 2024b; Han et al., 2024) underscores the importance of hallucinating object attributes in image-language models, such as OCR, object counting, and scene details, which we summarize as "semantic details". To benchmark this in LVLMs, we developed a setting focused on detecting hallucinations related to these details. We’ve employed a contrastive learning-inspired method using the HawkEye dataset (Wang et al., 2024c), which breaks down long videos into short video segments by semantic similarity. By computing the CLIP score of these video segments, we select segment pairs with a score above as the source video. Annotators then identify semantic differences between these pairs and create basic-hallucinated question-answer pairs from different perspective, resulting in a collection of 400 pairs and corresponding videos.
1.2 Extrinsic Hallucination
Our research focuses on detecting extrinsic factual hallucinations—facts consistent with reality but not verifiable from the source material—which can be contextually beneficial or detrimental. In text summarization, hallucinations are unwanted due to the need for source accuracy, while in conversational agents, they may enhance creativity and are more acceptable. We do not assess the overall value of such hallucinations but aim to identify them. At VideoHallucer, we curate instructional videos and course lectures, selecting content that stands alone clearly. Specifically, we select videos from YouCook (Zhou et al., 2018), COIN (Tang et al., 2019), and EDUVSUM (Ghauri et al., 2020) datasets as our source videos. We prefer shorter videos to minimize complexity. For instructional videos, we analyze tutorial steps and create questions about whether certain steps should be taken to finish a task. To prevent ambiguity, we focus on the final steps to detect hallucinations. We edit videos to exclude these steps and pose similar questions as factual hallucinated questions. For course lectures, annotators summarize the content and create questions to determine if the video contains the summary’s content. To create hallucinated questions, annotators modify the summary or add unrelated knowledge, ensuring the summary remains factual but not directly sourced from the video. They also provide explanations for these hallucinated questions to aid further research and improve clarity. As a result, we’ve created a dataset of paired basic and hallucinated questions and answers, totaling pairs from videos— instructional videos and course lectures.
Hallucinations that are not based on factual information can be particularly harmful due to their distorted content. Recently, there has been a growing interest among researchers in addressing this issue. At VideoHallucer, we specifically develop a setting to aid in the detection of non-factual hallucinations. To ensure consistency with the factual hallucinations, we select the same set of videos and corresponding basic questions used in the factual hallucination setting. For the creation of hallucinated content, we instructed annotators to manually alter the basic questions or infuse them with counterfactual information. The format of these questions remains identical to that used in the factual context. As a result, we created a comprehensive setting consisting of question-answer pairs linked to videos.
2 Dataset Statistics
There are five settings in VideoHallucer, each corresponding to a different type of hallucination in video understanding. For each setting, we create question-answer pairs, including 200 basic questions and hallucinated questions. To ensure a fair evaluation, the basic questions in both factual and non-factual settings are identical. As a result, we obtain question-answer pairs with an average length of words. Our dataset comprises videos, ranging from seconds to seconds of different settings, with an average length of seconds. These videos encompass most existing common video understanding benchmarks (Xu et al., 2017, 2016; Jang et al., 2017; Yu et al., 2019) as well as long video benchmarks (Xiao et al., 2021; Mangalam et al., 2023).
To present VideoHallucer more intuitively, we display the word cloud of our benchmark in Fig. 2 and the sunburst chart in Fig. 3. As shown in these figures, the questions within our benchmark are informative and contain three types: “Does”, “Is”, and “Should”. We believe our benchmark can directly and effectively reveal potential hallucination problems within LVLMs.
3 Evaluation
In this work, we opt for the VQA-based benchmark for the following reasons: (i) Influence of External Factors: Similar to metrics such as BLEU (Papineni et al., 2002) and ROUGE (Lin, 2004), the values in caption-based benchmarks can be affected by factors such as caption prompt and length (Li et al., 2023b). (ii) Complexity: Methods like CHAIR (Rohrbach et al., 2018) require intricate, human-crafted parsing rules. (iii) LLM Hallucinations: The potential hallucinations existing in LLMs’ generation make it unconvincing to use themselves for self-evaluation (Wang et al., 2024a). To ensure the credibility of our benchmark, we show a positive correlation between our QA-based evaluation and caption-based methods. More details are discussed in Sec. 6.
To mitigate biases such as the distribution of answers and language bias, we develop VideoHallucer using an adversarial approach (Tong et al., 2024). Specifically, for each evaluation item, we formulate two types of questions: a basic question and a hallucinated question. The basic question assesses the core capabilities of LVLMs, while the hallucinated question includes deliberately hallucinated content. We then calculate the overall accuracy by considering both the basic and hallucinated questions as a paired set, marking it as a hit only if both questions are answered correctly. We posit that enhancing a model’s ability to recognize and counter hallucinations should not compromise its performance on fundamental tasks. This dual-question structure is designed to ensure that improvements in counter-hallucination do not detract from the model’s original competencies.
In addition to the accuracy, we calculate the Yes Percentage Difference (Pct. Diff) and False Positive Ratio (FP Ratio) (Guan et al., 2024) to reveal the bias of these LVLMs. Specifically, the Yes Percentage Difference is calculated as
where is the set of video question pairs, is the prediction from models, is the ground truth. A smaller indicates the number of “yes” responses from models is closer to the ground truth, revealing less language bias. the False Positive Ratio is calculated as
where is the set of wrongly answered video question pairs. demonstrates the percentage of “yes” in all wrongly predicted answers. A value closer to indicates less bias from the models.
Experiment
In this section, we evaluate the most popular LVLMs on our VideoHallucer. We first present the setups of these models (Sec. 4.1), followed by the main results and performance analysis (Sec. 4.2). Additionally, we compare the effectiveness of current Image-Language Models on the object-relation and semantic details settings of VideoHallucer (Sec. 4.3). Finally, we reveal the human performance on our benchmark (Sec. 4.4).
We assess twelve LVLMs, comprising ten open-source models (7B unspecified), including VideoChatGPT, Valley2, Video-LLaMA2, VideoChat2, Video-LLaVA, LLaMA-VID, VideoLaVIT, MiniGPT4-Video, PLLaVA, and LLaVA-NeXT-Video-DPO, and two closed-source models, Gemini-1.5-Pro and GPT-4o. To make a fair comparison, we set all these baselines following their original setting including the number of frames and generation hyper-parameters.
2 Main Benchmark Results
The overall results are delineated in Table 3. We find that although all models demonstrate strong capabilities in answering basic questions, they experience a significant decline in accuracy when confronted with hallucinated questions. The overall accuracy significantly drops compared to the accuracy on basic and hallucinated questions due to a mistake in one of the question pairs. This pattern implies a widespread susceptibility to hallucination issues among the current models. Regarding the “Yes/No Bias”, models with an obvious bias are more likely to have hallucination problems. Specifically, we find that most models tend to generate “Yes” answers, except for VideoChat2, which is more likely to generate “No” answers. Since PLLaVA and VideoChat2 share the same video tuning data, the difference stems from the image data. Therefore, we believe the bias and hallucination issues in VideoChat2 originate from the training image data. Additionally, LLaVA-NeXT-Video-DPO shares similar tuning data with Video-LLaVA, and as a result, the DPO strategy significantly reduces bias and hallucination. Generally, we do not find a huge gap between open-source models and closed-source models, but all models lack significant human-like hallucination detection capabilities.
Here, we will dig deeper into different types of hallucinations in LVLMs. We illustrated the more detailed analysis in Figure 4. We have the following findings.
First, when comparing the various dimensions of the radar chart, we observe that most models exhibit fewer hallucinations in the object-relation setting (ORH) than in other areas. Specifically, the performance of most models in this setting is centered around . This observation points to a deficiency in existing modeling techniques to detect hallucinations beyond elementary visual cues.
Second, regarding semantic detail hallucination (SDH), when comparing LLaMA-VID, LLaVA-NeXT-Video, and PLLaVA with training data increasing, we find models with more training data significantly perform better than others. Moreover, we compare the different scales of these models on non-factual settings (ENFH), and we find larger models markedly surpass smaller models. We ascribe this superiority to the effects of data and model parameter scaling, which appear to bolster the model’s prowess in discerning visual details and retaining world knowledge.
Finally, for extrinsic factual hallucination (EFH), most models demonstrate inability in this setting. To be more specific, most existing models can’t discern hallucination issues that align with world knowledge but contradict the video context. Given the significant differences in model performance in the two settings of extrinsic hallucinations, we will explore this issue in greater depth in the Sec. 5 to uncover the reasons behind this phenomenon.
3 Comparison between Image-Language Models
We carried out a comparative study of two types of multimodal models: image-language models and video-language models. For this analysis, we selected the top-performing LLaVA-1.5 and GPT4V. To facilitate a fair comparison, we used the middle frame from each video as the input for the image-language models. The findings, presented in Table 4, show that open-source image-language models have a superior performance in detecting object-relation hallucination, even though the dataset includes dynamic interactions. We believe this discrepancy in performance is due to two main reasons. First, there is a significantly larger amount of training data available for images than for videos. This means that image-language models benefit more from the scaling law. Second, videos tend to include more noise than images, making them more likely to encounter the hallucination problem. In light of these insights, we suggest that future development of LVLMs could be enhanced by integrating image datasets (Lin et al., 2023) or by building upon the foundations of existing image-language models (Zhang et al., 2024; Xu et al., 2024).
4 Human Evaluation
We recruited three individuals proficient in English to evaluate the VideoHallucer benchmark. Each evaluator possesses basic computer knowledge, which is essential since part of the extrinsic dataset includes computer science courses. To mitigate potential bias from the evaluators, we randomized the sequence of question-answer pairs to ensure that basic and hallucination pairs did not appear consecutively. We then computed the Pearson correlation coefficient among all evaluators’ scores, which resulted in a moderate agreement with a value of .
Self-PEP: Self-improvement with Predict-Explain-Predict
In Sec. 4.2, we find that most models perform worse in the factual hallucination setting compared to others on VideoHallucer. In this section, we aim to understand the reasons behind this by comparing the fact detection and hallucination detection abilities of these models (Sec. 5.1). Based on our findings, we propose a simple yet effective method with explanation to mitigate the models’ hallucination issues (Sec. 5.2).
One of our key findings suggests that these models struggle to identify external factual hallucinations within video content. As a result, we have redirected our focus in this section to assess the models’ capability to discern factual knowledge. To this end, we crafted additional QA pairs derived from questions in the extrinsic hallucination setting. These questions aim to determine the factual accuracy of statements in the original extrinsic hallucination questions.
In our experiment, we utilized course videos from this context due to their rich factual content. For instance, we paraphrased extrinsic factual hallucination questions to “Does the following course summary include any non-factual information? {summary}”, and non-factual hallucination questions to “Does the following course summary encompass all essential factual information? {summary}”. By setting the correct answer to these questions as "no", we sought to counteract any language biases present in language models. The experimental results, illustrated in Fig. 5, reveal that most models are more adept at detecting factual knowledge than at detecting hallucinations. This indicates that current methods can effectively recognize counterfactual content. However, their ability to detect hallucinations in video data is markedly limited, highlighting a significant potential for improvement in this area. Among these models, we find the LLaVA-NeXT-Video-DPO performs better on hallucination detection than fact detection, we believe this gain comes from the DPO. Therefore, we believe that human feedback could mitigate hallucination issues in LVLMs.
2 Self-PEP Framework
Advances in LLMs have seen significant improvements through costly human feedback (Ouyang et al., 2022; Sun et al., 2023). To reduce expenses, recent studies (Madaan et al., 2023; Shinn et al., 2023; Ye et al., 2023; Yan et al., 2023; Pan et al., 2023; Chen et al., 2023b; Zhang et al., 2023c; You et al., 2023; Lightman et al., 2023) focus on self-improvement of LLMs for better performance and clarity. Our research in Sec. 5.1 shows LLMs are better at identifying facts than hallucination, with the latter demanding a firm grasp of factual context. Building on these insights, we investigate further in this section. We examine how explanation can affect a model’s ability to detect hallucinations. Our methodology involves prompting the model to generate a prediction and then provide an explanation for its output. We then use the explanation to refine the model’s predictive accuracy. This experiment is conducted within the context of extrinsic factual hallucination detection, and the results are illustrated in Fig. 7. Although the self-generated explanations are less impactful compared to those derived from ground truth explanations, our results demonstrate that they do contribute to improving the model’s capability in hallucination detection. This improvement is related to the model’s proficiency in fact detection.
Building on these insights, as depicted in Fig. 7, we devise an innovative framework called Self-Improvement with Predict-Explain-Predict (Self-PEP). This framework is designed to enhance the model’s resilience against the tendency to produce hallucinations. It capitalizes on the model’s established strength in fact detection over hallucination identification. The Self-PEP framework operates in two phases: self-improvement and self-explanation. In the self-improvement phase, the model autonomously extracts visual knowledge (Yin et al., 2023), while the self-explanation phase involves a three-step process: predict, then explain, and finally refine the prediction using the explanation. The implementation details of the Self-PEP framework and qualitative results are provided in Appendix A.3.3.
The efficacy of this framework is demonstrated in Table 5, where we note that Self-PEP significantly boosts the performance of most models on the VideoHallucer benchmark, yielding substantial improvements. In addition, we find the improvements on hallucinated questions are much more significant than the basic questions. When comparing all different settings, we find our methods could consistently improve all models’ performance on extrinsic factual hallucination. Our method, for instance, could potentially harm PLLaVA in other settings while improving its performance solely on the factual setting. Consequently, we believe our method could reduce hallucination issues within existing LVLMs, particularly factual hallucinations, which are consistently present in LVLMs.
Conclusion and Discussion
In this work, we introduce VideoHallucer, a novel and comprehensive benchmark for detecting hallucinations in LVLMs. Our adversarial approach, which challenges models with paired questions, ensures a thorough evaluation of a model’s ability to discern hallucinations. Through rigorous testing of 12 LVLMs, we have identified the pervasive nature of spurious hallucinations and the limitations of model scaling in addressing certain types of hallucinations. Our innovative Self-PEP framework has demonstrated the potential to significantly enhance model performance against hallucinations.
Adversarial attacks are deliberately crafted inputs designed to provoke erroneous outputs from a model, potentially leading to various security breaches such as unauthorized data access, fraud, system intrusion, malware deployment, content manipulation, and service disruption. In contrast, our tool, VideoHallucer, is specifically tailored to assess perceptual challenges. It evaluates whether a model can produce accurate and reliable responses based on the provided context, without being misled by superficially plausible but incorrect information.
VideoHallucer is primarily designed to evaluate issues related to hallucinated content and the faithfulness of generated text from LVLM. This approach is relatively direct and specific. In comparison, other benchmarks may concentrate on basic visual understanding or more abstract cognitive tasks such as logical reasoning or advanced scene comprehension.
Existing hallucination benchmarks primarily encompass two types of benchmarks and evaluation pipelines found in existing research: VQA-based and Caption-based. The VQA-based benchmark, exemplified by methods like POPE (Li et al., 2023b), employs binary question-answering to detect instances of object hallucination. In contrast, caption-based benchmarks (Rohrbach et al., 2018; Wang et al., 2023; Zhai et al., 2023) assess the precision or recall of hallucinated objects within captions, with evaluations conducted either through rule-based parsing or LLMs. In our study, we opt for the VQA-based benchmark for several reasons. Firstly, akin to established evaluation metrics for content generation such as BLEU (Papineni et al., 2002) and ROUGE (Lin, 2004), these values can be influenced by extrinsic factors, including the nature of the caption prompt and the caption’s length (Li et al., 2023b). Secondly, methods like CHAIR (Rohrbach et al., 2018) depend heavily on intricate, human-crafted rules for parsing, which increases complexity. Thirdly, although recent advancements have incorporated LLMs for evaluation, employing a tool that itself is prone to have hallucination issues for assessment purposes is problematic. To establish the credibility of our benchmark, we demonstrate a positive correlation between our QA-based evaluation and the caption-based methods. Further details on this analysis are available in Appendix A.1.
References
Appendix A Appendix
We provide appendices and supplementary materials as follows:
In Section A.1, we show the detailed analysis of two ways of evaluation methods for vision hallucination benchmark.
Section A.2 demonstrates the detailed annotation procedures.
Section A.3 outlines the implementation details of our study
In A.3.1, we show baseline configurations
In A.3.2, we demonstrate the evaluation prompts
In A.3.3, we reveal the implementation details of Self-PEP
Section A.4, we show the analysis of the coherence of the self-generated explanation and the groud-truth explanation.
In Section A.5.1, we present the comprehensive results for the VideoHallucer benchmark.
Section A.5.2 details the performance of Self-PEP on VideoHallucer
A.6. Example Questions. We provide examples from VideoHallucer that cover various types of hallucinations: object, spatial relation, temporal relation, absolute temporal, relative temporal, and semantic detail (including attribution, event, count, OCR, camera, and scene), as well as extrinsic factual (instruction, course) and non-factual (instruction, course) hallucinations.
We show the limitations and ethic states in Sec. A.7.
In this study, we investigate the relationship between VideoHallucer (QA-based mehtod) and the caption-based method. To determine the effectiveness of our QA-based method in comparison to the caption-based approach, we use the Pearson correlation coefficient to measure the coherence. Our focus is on two well-established caption-based evaluation metrics: the CHAIR score [Rohrbach et al., 2018] and the Coverage Score [Zhai et al., 2023]. Specifically, we randomly select a sample of 100 images from the COCO dataset to investigate the correlation between our newly introduced accuracy on VideoHallucer and the traditional metrics employed for image caption. We apply the experiments on LLaVA-1.5. To prevnet the issue of hallucination in LLMs, we implement a rule-based method to calculate the Coverage Score similar to CHAIR. Our findings reveal a moderate positive correlation with the Coverage Score () and a weak negative correlation with the CHAIR score (). We illustrate the correlation between the overall accuracy on VideoHallucer and the Coverage Score in Fig. 8. These results indicate that overall accuracy on VideoHallucer correlates positively with the caption-based methods, suggesting that VideoHallucer is a reliable and robust alternative to caption-based evaluation methods for assessing hallucinations in generative models.
A.2 Annotation Details
To enhance the quality and reduce the annotation costs of our benchmark, we have developed a semi-automatic pipeline for constructing VideoHallucer. The process begins with the use of readily available tools and templates to create or collect initial pairs of basic and hallucinated content. Subsequently, we engage annotators to meticulously review and refine the question-answer pairs, ensuring they are well-grounded in the video content. This review focuses on three key aspects: correctness, ambiguity, and fluency. To guarantee the high standard of VideoHallucer, we implement a two-stage human verification process. In the following paragraphs, we will give a detailed introduction to different sets of VideoHallucer.
In the domain of vision-language modeling, object hallucination is a persistent challenge, particularly within the context of image captioning tasks. These tasks demand that the model accurately identifies and describes objects in an image. A significant body of research has been dedicated to addressing this issue, encompassing various approaches such as the development of benchmarks, the refinement of evaluation metrics, and the exploration of both analytical and practical solutions. More recently, the focus of research has expanded from object hallucination to include relation hallucination, with a particular emphasis on spatial relationships between objects. Building on this foundation, our work takes a novel step by applying these concepts to video content. We aim to capture not only the objects but also the dynamic interactions between them, considering both spatial and temporal dimensions. We posit that the ability to detect basic visual concepts is a fundamental skill for LVLMs.
We posit that objects and their relations are fundamentally intertwined in the basic visual comprehension of videos. To reflect this in our benchmark, we have constructed an object-relation hallucination set. This set is divided into three distinct categories: subject, relation, and object. To elaborate, there is a wealth of prior research that has focused on annotating objects and relations within video content. In order to leverage this existing body of work and avoid duplicating efforts, we have developed our object-relation hallucination sets based on these pre-existing datasets. Specifically, we have utilized the VidOR [Shang et al., 2019, Thomee et al., 2016] and VidVRD [Shang et al., 2017] datasets as the foundation for our set, ensuring that we build upon the annotations they provide without unnecessary redundancy in human annotation labor. For subject and object, we following the paradigm of POPE [Li et al., 2023b], use the template “Is there a/an {object} in the video?”. For relation, for spatial relation, we use the template “Is {subject} {relation} {object} in the video?”, for temporal relation, or action, we use the template “Does/Do {subject} {relation} {object} in the video?”. For hallucinated questions, instead of randomly selecting from existing object pools, we ask annotator to write semantic different but visually similar object, and anatomy of the relation word. After obtaining the question automatically, we manually check the question answer pairs. As a result, we get 200 basic-hallucination question-answer pairs, totally 400 question-answer pairs and 183 videos.
Temporal information is the most fundamental difference between video input and image input. However, most recent video-text LLMs mainly focus on factual contents in short videos, ignoring the temporal information of multiple events in long videos. Recent studies [Lei et al., 2023, Xiao et al., 2023] found that they have almost no ability to perform time-related tasks such as video grounding.
The understanding of temporal information in videos contains two aspects: where in the video timeline is an event located, and how long it lasts. Therefore, we plan to test 3 time-related capabilities of video-text models: (1) 50 examples for absolute temporal understanding, where the model is asked to answer the position of an event in the video, such as whether it is located at the beginning or at the end of the video, (2) 75 examples for relative temporal understanding, where the model is asked to determine the order of two events, and (3) 75 examples for event length understanding, where the models is asked to compare the duration length of two events.
We construct the temporal hallucination benchmark based on the val set of ActivityNet Captions [Krishna et al., 2017], a popular temporal video grounding dataset with event-level video annotations: each event is labeled with a text description and a time span (start sec, end sec) in the video. For absolute temporal understanding examples, we sampled 50 events from 50 different videos that are entirely located in the first or last third of its video. For events in the first third of its video, we use “Does {event} happen at the beginning of the video?” as the basic question, and “Does {event} happen at the end of the video?” as the hallucinated question, and vise versa for the events in the last thirds of its video. For relative temporal understanding examples, we sampled 75 pairs of events from 75 different videos which satisfies that the end time of the first event is at least earlier than the start time of the second event, where is the duration of the video. We randomly choose “Does {event1} happen earlier than {event2}?” or “Does {event2} happen later than {event1}?” as the basic question, and “Does {event1} happen later than {event2}?” or “Does {event2} happen earlier than {event1}?” as the hallucinated question. For event length understanding examples, we also sampled 75 pairs of events from 75 different videos which satisfies that the duration of the first event is at least 2 times longer than the duration of the second event. We randomly choose “Is the event {event1} longer than {event2}?” or “Is the event {event2} shorter than {event1}?” as the basic question, and “Is the event {event1} shorter than {event2}?” or “Is the event {event2} longer than {event1}?” as the hallucinating question. As a result, we get 200 basic-hallucination question-answer pairs, totally 400 question-answer pairs and 165 videos.
Recent studies increasingly emphasize the importance of object attributes in visual analysis. A number of these works [Sun et al., 2023, Wang et al., 2024a, Chen et al., 2023a, Wang et al., 2024b, Han et al., 2024] delve deeper, attempting to uncover a broader spectrum of visual details such as optical character recognition (OCR), object counting, and environmental context. Given the challenge of creating a definitive classification for the diverse range of visual details, our work consolidates these attributes under the umbrella term “other semantic details.” This encompasses OCR, attribute recognition, viewpoint identification, and scene understanding, among others. It has been observed that most vision-language models exhibit deficiencies in accurately identifying these nuanced details. To address this gap, we have specifically developed a dataset aimed at enhancing the detection of hallucinations related to these other semantic details.
To enhance our understanding of the detailed semantic content in videos, we have adopted a technique inspired by contrastive learning, which involves identifying distinctions by drawing comparisons. In particular, we utilize video clips from the HawkEye dataset [Wang et al., 2024c], which segments lengthy videos into shorter clips based on PySceneDetect, and then merging clips that share high semantic similarity. These clips are designed to encapsulate distinct semantic meanings. Once we have these video clips, we compute the CLIP score among them to measure their semantic relatedness. We then select pairs of video clips with a CLIP score exceeding for further analysis. Subsequently, we engage annotators to pinpoint the semantic disparities between these selected pairs of clips. Based on these differences, the annotators are tasked with creating a set of basic-hallucinated question-answer pairs. Throughout this process, we instruct the annotators to pay particular attention to various aspects of the semantic details within the clips. As a result, we get 200 basic-hallucination question-answer pairs, totally 400 question-answer pairs and 400 videos.
Extrinsic factual hallucinations, which are unverifiable against the original source material yet remain factually consistent, may sometimes be considered beneficial. Certain studies suggest that these hallucinations could enhance the richness and informativeness of the generated content by providing additional background knowledge. However, the usefulness of such hallucinations is highly context-dependent. In the field of text summarization, any form of hallucination is undesirable, as accuracy and fidelity to the source are paramount. Conversely, in conversational agents, there is an expectation for the generation of varied and imaginative responses, where factual hallucinations might be more acceptable. Therefore, in our research, we do not seek to determine the overall value of factual hallucinations. Instead, our focus is on developing models that can detect and identify these hallucinations.
At VideoHallucer, we meticulously curate two distinct categories of videos: instructional videos and course lectures. Our selection criteria prioritize content that is rich in knowledge and can be understood independently of extrinsic information, thereby minimizing potential ambiguity. For the instructional category, we extract samples from the YouCook [Zhou et al., 2018] and COIN [Tang et al., 2019] datasets. These include cooking recipes and guides for various activities. In the realm of course lectures, our focus is on the EDUVSUM dataset [Ghauri et al., 2020], which features online courses in fields like computer science, the history of science, and engineering. To simplify matters, we intentionally choose shorter videos from these sources. When dealing with instructional videos, our process begins by identifying the steps outlined in each tutorial. To formulate basic questions, we adopt a template that reads: “Based on the video, should we {step} when we {target}?” For hallucinated questions, we pay special attention to the final steps, as they typically contain detailed information and dropping these steps won’t make the instructional video incomplete. We select these as the focal point for VideoHallucer. Specifically, we trim videos and remove the content related the these stepsand pose questions in the same format.
For course lectures, we first instruct annotators to condense the content into a summary that captures the key points. We then generate basic questions following the pattern: “Does the video provide information that allows us to learn that {summary}?” For hallucinated questions, we challenge annotators to alter the summary or incorporate knowledge not covered in the video. This ensures that while the summary remains factual, it cannot be solely derived from the video content. Moreover, we ask the annotator to write an explanation for these hallucinated questions. The explanation point out the reason and the location of the hallucinated question. Through this meticulous process, we have compiled a dataset of 200 paired basic and hallucination questions and answers, resulting in a total of 400 question-answer pairs and 200 videos, where 150 videos are instructional videos and 50 videos are online courses.
Hallucinations that are not based on factual information can be particularly harmful due to their distorted content. Recently, there has been a growing interest among researchers in addressing this issue. At VideoHallucer, we have specifically developed a dataset to aid in the detection of non-factual hallucinations.
To ensure consistency with the detection of factual hallucinations, we have selected the same set of videos and corresponding basic questions used in the factual hallucination dataset. For the creation of hallucinated content, we instructed annotators to manually alter the basic questions or infuse them with counterfactual information. The format of these questions remains identical to that used in the factual context. As a result of this process, we have compiled a comprehensive dataset consisting of 400 question-answer pairs linked to 200 videos.
A.3 Implementation Details
In our experiment, we choose the 7B-level video language models for fair comparison, including VideoChatGPT [Maaz et al., 2023], Valley2 [Luo et al., 2023], Video-LLaMA2 [Zhang et al., 2023b], VideoChat2 [Li et al., 2023c], VideoLLaVA [Lin et al., 2023], LLaMA-VID [Li et al., 2023d], VideoLaVIT [Jin et al., 2024], and MiniGPT4-Video [Ataallah et al., 2024]. To assure a fair comparison, we take the default hyper-parameter of these models, including “max_new_tokens”, “do_sample”, “temperature”, “num_beams”, and “num_of_frames”. To further reveal the potential of scaling law, we additionally add the existing best-performed closed-source model Gemini-1.5-Pro [Reid et al., 2024] as the current upbound of existing LVLMs. For the Gemini-1.5-Pro, we take the fps as 1, and to improve the evaluation efficiency, we set the max number of frames as 128.
A.3.2 Prompt for the Overall Results
To evaluate responses, we append the prompt “Answer the question using ’yes’ or ’no’.” to the end of each question. We then compare the answer to “yes” or “no” to calculate accuracy.
A.3.3 Implementation Detail of the Self-PEP
As illustrated in Fig. 7, there are two main components of the Self-PEP: the self-improvement and the self-explain. For the self-improvement, we ask the model to extract the visual knowledge [Yin et al., 2023]. In our implementation, we ask the model to generate the caption for simplicity, the prompt is “Describe the video: ”. After obtaining the caption, we ask the model to predict the answer the question with the input of both the video and the self-generated caption. The prompt is “Description: {description} Please provide a clear response to the question below by watching the video. If necessary, you can also use the accompanying Description to help refine your answer. Your response should be a simple ’yes’ or ’no’. Question: {question} Answer the question using ’yes’ or ’no’: ”. We further apply the self-explain to the model, by asking the model to explain-then-predict to get the final answer. The prompt is “Description: {description} Please offer a detailed explanation for your answer to the following question. After explaining, verify the accuracy of the information you’ve used in your explanation. Once you’ve confirmed the facts, please respond to the question with a simple ’yes’ or ’no’. Question: {question} Answer: {predict} Answer the question using ’yes’ or ’no’: ”. Furthermore, we illustrate the qualitative study of the Self-PEP framework on Gemini-1.5-Pro in Fig. 9, we find that with the explanation, the model is able to correct its answer.
A.4 Explanation Analysis
Upon comparing Fig. 7 with Fig. 5, it becomes evident that there is a correlation between the model’s self-explanation capability and its ability to detect factual information. To further explore this relationship, we evaluate how well the self-generated explanations align with the established ground truth explanations. For this purpose, we employ GPT-4 to conduct a consistency analysis, utilizing the prompt outlined in Fig. 10. The results, presented in Table 6, indicate that the quality of the explanations has a significant impact on the model’s proficiency in identifying instances of hallucination. Consequently, our goal is to enhance the model’s capacity to produce accurate and reliable explanations as a means to mitigate the issue of hallucination.
A.5 Detailed Quantitative Results
In this section, we provide the detailed quantitative results on VideoHallucer, including the overall results (Sec. A.5.1) and the results with Self-PEP (Sec. A.5.2).
In this section, we present the comprehensive results of various LVLMs across different settings of VideoHallucer. Table 7 details the outcomes for object-relation hallucination; Table 8 displays the findings for temporal hallucination; Table 9 outlines the results for semantic detail hallucination; Table 10 depicts the results for extrinsic factual hallucination; and Table 11 provides the results for extrinsic non-factual hallucination.
A.5.2 The Detailed Results of Self-PEP
In this section, we further show the detailed results of Self-PEP on different settings of VideoHallucer. Table Table 12 details the outcomes for object-relation hallucination; Table Table 13 displays the findings for temporal hallucination; Table Table 14 outlines the results for semantic detail hallucination; Table Table 15 depicts the results for extrinsic factual hallucination; and Table Table 16 provides the results for extrinsic non-factual hallucination.
A.6 Example Questions of VideoHallucer
In this section, we provide more cases from VideoHallucer.
A.7 Limitations and Ethic States
Although we take a lot of strategies to make sure the quality of VideoHallucer, there is noise introduced by human annotations. In addition, currently, the scalability of this benchmark is still limited.
VideoHallucer is designed to counter hallucination in LVLMs, however, the hallucinated questions in this benchmark could lead to misinterpretation in related research area.
We bear all responsibility in case of violation of rights and our dataset is under the license of CC BY-NC-SA (Attribution-NonCommercial-ShareAlike).