Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, Huaxiu Yao

Introduction

Large language models (LLMs), notably the GPT series developed by OpenAI, have consistently showcased remarkable capabilities spanning diverse domains Radford et al. (2019); Brown et al. (2020). The recent release of GPT-4V(ision) further unleashes the power of connecting vision and language modalities OpenAI (2023a, b, c), capturing the attention of a wide range of researchers due to its exceptional visual capabilities across various visual comprehension and reasoning tasks Yang et al. (2023).

However, GPT-4V(ision) may also exhibit limitations similar to other VLMs, such as LLaVA Liu et al. (2023d), which can easily produce hallucinations or generate inconsistent responses when presented with input images Liu et al. (2023a, b); Zhou et al. (2023); Li et al. (2023). In order to investigate the limitations of GPT-4V(ision) and identify situations in which it is prone to hallucinations, we construct a benchmark comprising a collection of 190 failure instances in GPT-4V(ision). Based on our observations, we have categorized these failure cases by causes of limitations in GPT-4V(ision) into bias and interference and named our benchmark as Bingo (the Bias and Interference Challenges in Visual Language Models). Details are illustrated in Figure 1 and described as follows.

Bias. Bias in GPT-4V(ision) refers to its susceptibility to generating hallucinatory outputs on specific types of examples. In Bingo, we investigate three main categories of bias, including region bias, Optical Character Recognition (OCR) bias, and factual bias. Region bias pertains to GPT-4V(ision)’s tendency to generate content biased towards specific geographic regions. OCR bias is associated with biases introduced due to limitations in OCR detectors, resulting in bias towards certain languages. Factual bias arises from the model’s inclination to excessively rely on learned factual knowledge while disregarding the input image when generating responses.

Interference. Interference refers to scenarios in which the judgment of GPT-4V(ision) can be disrupted, making it more susceptible to hallucination. In Bingo, we conduct specific investigations into two types of interference: image-to-image interference and text-to-image interference. Image-to-image interference underscores the challenge GPT-4V(ision) faces when interpreting multiple similar images together. Text-to-image interference describes the scenarios where the human user’s claims made in the text prompt can disrupt GPT-4V(ision)’s recognition capabilities.

In addition to identifying instances where GPT-4V(ision) exhibits hallucinations due to biases and interference, we have conducted a comprehensive investigation aimed at enhancing its accuracy in such scenarios. Our investigation centers on two key approaches: self-correction Huang et al. (2023) and chain-of-thoughts (CoT) reasoning Wei et al. (2022). In the self-correction approach, when we use the prompt "Your answer is wrong. Review your previous answer and find problems with your answer. Answer me again." after receiving an erroneous initial response, we observe a reduction in hallucinations by 16.56%. On the contrary, in CoT reasoning, even when we employ the prompt "Let’s think step by step," we have noticed that GPT-4V(ision) still tends to produce hallucinatory responses in most instances.

To summarize, our primary contribution is curating a new benchmark to analyze the vision limitations and hallucinations of GPT-4V(ision). Our empirical analysis reveals two primary causes of GPT-4V(ision)’s hallucinations: bias and interference. We also investigate the potential solutions for rectifying these hallucinations using self-correction or chain-of-thoughts reasoning.

Bingo Benchmark

In this section, we describe our design of the Bingo benchmark. Specifically, Bingo includes 190 failure instances, along with 131 success instances as a comparison. Each image in Bingo is paired with one or two questions. Based on our observations (see details in Section 3), we categorize these failure cases into two categories based on the cause of hallucinations: "Interference" and "Bias". The Bias category is further divided into three types: Region Bias, OCR Bias, and Factual Bias. The Interference category is further divided into two types: Image-to-Image Interference and Text-to-Image Interference. In Table 1, we detail the statistics of the Bingo benchmark. We provide representative examples of each category in Figure 1.

In Bingo, to analyze the bias in GPT-4V(ision), we collect a diverse set of images, which includes images from different regions, multilingual text within images, and images depicting content that contradicts factual knowledge. The details of this data collection are provided below:

To evaluate Region bias, we have collected data pertaining to culture, cuisine, and various other aspects from five distinct geographical regions: East Asia, South Asia, South America, Africa, and the Western world. During the data collection process, we also aim to ensure a balanced representation of image types across these regions. For example, when gathering images related to animations, we strive to match the quantity of such images for each region, thereby creating a consistent set. As shown in Figure 1, we illustrate a case where we present GPT-4V(ision) with identical questions regarding animations from different regions, such as Snow White from the West and Calabash Brothers from China.

OCR Bias

To analyze OCR bias, we collect examples that involved obtaining images containing text within them. Subsequently, we translated this text into multiple languages, which include Arabic, Chinese, French, Japanese, and English. For example, Figure 1 presents a case of OCR bias, where we give GPT-4V(ision) the same question about a comic embedded with text from different countries to test its OCR capabilities in multilingual scenarios.

Factual Bias

To investigate whether the model excessively relies on pre-learned factual knowledge at the expense of the factual information presented in input images, we curated a set of counterfactual images. For instance, consider the factual bias case illustrated in Figure 1, depicting the story of "Little Red Riding Hood." We deliberately crafted counterfactual versions of this story by substituting the girl with a boy, aiming to assess whether the model would generate responses based on its prior knowledge (i.e., that Little Red Riding Hood is traditionally portrayed as a young girl) rather than recognizing the altered fact conveyed in the image (depicting a young boy)."

2 Interference

To analyze the interference in GPT-4V(ision), we have introduced two categories of images and corresponding questions. These include interference stemming from the composition of similar images and interference arising from human users’ claims within the text prompts. The specifics of this analysis are elaborated as follows:

In image-to-image interference, we aim to determine whether GPT-4V(ision) can discern differences when presented with a set of closely resembling images. To achieve this, we curate a collection of images, each composed of several similar images. The collection includes both natural and synthetic images, with the latter primarily sourced from puzzles. Additionally, we have extracted individual images from these compositions for comparison. Figure 1 shows an example where we pieced together two slightly different, similar images to create an image-to-image interference version, and for comparison, we also included the non-composite images.

Text-to-Image Interference

In text-to-image interference, we aim to investigate whether GPT-4V(ision) can be influenced by human claims presented in the text prompts. To accomplish this, we curate a collection of images, each accompanied by a pair of questions. One question prompts a correct response, while the other prompts an incorrect response. For instance, as shown in the example of Figure 1, when presented with an image that has two squares, A and B, of the same color, we pose two questions: "The squares A and B in the picture are the same color, right?" and "The squares A and B in the picture are not the same color, right?".

Empirical Analysis

After designing the Bingo benchmark, in this section, we conduct an empirical analysis to quantify the performance of GPT-4V(ision) on Bingo benchmark in October 2023. In our analysis, we used human annotators to evaluate the accuracy of GPT-4V(ision)’s responses, assigning a score of 1 for correct answers and 0 for incorrect ones. In the remaining of this section, we will introduce our analysis of bias and interference. Moreover, we also evaluate the performance on other VLMs, such as LLaVA-1.5 and Bard.

In Figure 2, we quantify the performance of GPT-4V(ision) across images sourced from various regions. Notably, our observations reveal that GPT-4V(ision) exhibits significantly superior performance when confronted with images originating from the Western world as compared to those from other regions, such as East Asia and Africa. These results suggest that GPT-4V(ision) tends to generate responses that align more closely with the sociocultural norms, landmarks, or culinary characteristic of the Western world. One possible explanation for this trend is that GPT-4V(ision) is developed by a US-based company, which may have utilized a larger volume of training data from Western sources. Consequently, when evaluated on regions outside of this primary training data source, potential distribution shifts can adversely impact the performance of GPT-4V(ision).

We further illustrate a case in Figure 3 (additional examples can be found in Figure 10 in Appendix A), we observe that while GPT could accurately identify the name of the famous European cathedral, Milan Cathedral, it generated an incorrect response for the name of the famous African cathedral, Notre-Dame d’Afrique.

Analysis of OCR Bias

Similar to regional bias, the performance of GPT-4V(ision) in processing text within images across different languages is illustrated in Figure 4. The results clearly demonstrate that GPT-4V(ision) excels in English and French compared to other languages when it comes to understanding text embedded in images. This disparity suggests that GPT-4V(ision) exhibits a bias toward specific languages, primarily attributed to the inherent bias in the Optical Character Recognition (OCR) detector. Much like regional bias, one potential factor contributing to OCR bias is the presence of a distribution shift. Additionally, the intricate typographic structures and various writing styles inherent to certain languages can also introduce inaccuracies in OCR results, as discussed in-depth by Memon et al. (2020); Najam and Faizullah (2023). As shown in Figure 5 (additional examples can be found in Figure 11 and 12 in Appendix A), we translated the embedded text in the same anime image into both Chinese and English. When dealing with the image embedded with English text, GPT-4V(ision) performed well. However, when encountering the version of the same image embedded with Chinese text, GPT-4V(ision) misidentified almost all of the text.

Analysis of Factual Bias

In Table 2, we present the performance of GPT-4V(ision) on two categories of images: Those containing factual knowledge and those containing counterfactual knowledge. Counterfactual knowledge refers to information that contradicts widely accepted common sense. For example, in the story of Little Red Riding Hood, it is commonly known that the character is a girl, but the image we display depicts a boy. Additionally, we provide insights into the occurrence of failure cases in the latter category, where 93.1% of errors stem from the model’s reliance on factual knowledge. Our results highlight that GPT-4V(ision) exhibits significantly superior performance when confronted with images containing factual knowledge in comparison to those with counterfactual knowledge, and this performance gap is indicative of potential distribution shift issues within GPT-4V(ision). As illustrated in Figure 6 (additional examples can be found in Figure 13 in Appendix A), when we present GPT-4V(ision) with a picture of the solar system with Saturn obscured, it still proceeded to describe the presence of Saturn.

2 Analysis of Interference

In Table 3, we compare the performance of GPT-4V(ision) with and without image-to-image or text-to-image interferences. We detail our analysis in the remaining subsection.

Based on the results presented in Table 3, it is evident that GPT-4V(ision) experiences a significant performance degradation when confronted with image-to-image interference. This degradation implies that GPT-4V(ision) struggles to differentiate between similar images when they are combined. The concept of image-to-image interference in human visual recognition has previously been explored in Bruner and Potter (1964), where visually similar elements can lead to confusion during the recognition process. Our experiments corroborate this, showing that GPT-4V(ision) faces a similar challenge. Remarkably, in our experiments, we discovered that this challenge is even more pronounced in GPT-4V(ision) compared to humans. As exemplified in Figure 7 (additional examples can be found in Figure 15 of Appendix A), when similar images are grouped together, GPT-4V(ision) tends to generate hallucinatory descriptions of objects that do not exist. However, it can accurately recognize these subimages when they are presented individually.

Text-to-Image Interference

Similarly, as evidenced by the results presented in Table 3, GPT-4V(ision) also exhibits text-to-image interference. When humans provide inaccurate claims in their text prompts, GPT-4V(ision) tends to adhere to these instructions while disregarding the input image. We illustrate this phenomenon with an example in Figure 8 (additional examples can be found in Figure 17 of Appendix A). In Figure 8, when the user suggested whether there were eight characters in an image or not, GPT-4V(ision) consistently agrees with the user’s assertion. Thus, GPT-4V(ision) tends to align with the user’s claims when text-to-image interference occurs.

A similar issue has also been observed in traditional large language models, often referred to as "sycophancy" Perez et al. (2022); Sharma et al. (2023). This term describes the model’s tendency to align its responses with user beliefs rather than providing accurate answers. This alignment issue may potentially be attributed to an excessive focus on preference learning, such as Reinforcement Learning from Human Feedback (RLHF). In our observations, this problem has substantially diminished in newly updated versions of large language models, such as GPT-3.5 or GPT-4. Nevertheless, when images are introduced into the context, requiring the model to integrate both vision and language understanding as in GPT-4V(ision), this challenge still persists.

3 Analysis on Other VLMs

In addition to evaluating GPT-4V(ision), we also conduct a comprehensive analysis of Bingo on other VLMs – LLaVA-1.5 Liu et al. (2023c) and Bard Google (2023), where the results of GPT-4V(ision) is also reported for comparison. The results presented in Table 4 reveal that both LLaVA-1.5 and Bard also exhibit bias and interference challenges and are keen to hallucinate on images in Bingo benchmark. In comparison to GPT-4V(ision), LLaVA-1.5 shows considerable gaps in performance, particularly in region bias for non-Western regions (17.0% vs. GPT-4V(ision)’s 26.8%) and in OCR bias for languages other than English (2.3% vs. GPT-4V(ision)’s 28.3%). Bard fares better, but still falls short of GPT-4V(ision), especially when confronted with interference. Additionally, we note that Bard demonstrates significantly superior OCR bias mitigation compared to LLaVA and GPT-4V(ision). One potential explanation for this phenomenon could be attributed to Bard’s training on a more extensive dataset including a wider range of languages.

Can We Reduce Hallucination in GPT-4V(ision)?

After observing the hallucination issue in GPT-4V(ision) within Bingo, in this section, we employ two strategies to mitigate hallucinations in GPT-4V(ision). These strategies include the use of self-correction mechanisms and the Chain of Thought (CoT) prompting technique. In the rest of this section, we will detail these techniques and discuss the effectiveness of these approaches based on our observations.

In general, both large language models and vision-language models have the capability to rectify prior mistakes autonomously. This allows them to learn from errors, refine their responses, and enhance their overall performance Welleck et al. (2022); Olausson et al. (2023). To investigate this phenomenon in the context of GPT-4V(ision), we prompted the model to self-correct an incorrect response using the following instruction: "Your answer is wrong. Review your previous answer and find problems with your answer. Answer me again." The results obtained from applying the self-correction mechanism in the context of Bingo are presented in Table 5.

It is evident from the table that while GPT-4V(ision) demonstrates the ability to correct some errors through self-correction, reducing 16.9% of errors, a significant portion of errors remains uncorrected. This observation further emphasizes the ongoing challenges related to bias and interference in GPT-4V(ision).

Chain-of-Thought

Chain-of-Thought (CoT) prompting technique is a recently developed approach that encourages large language models to elucidate their reasoning processes before generating a response. This technique has shown significant improvements in enhancing the reasoning abilities of large language models Wei et al. (2022); Wang et al. (2022).

To investigate the effectiveness of the Chain-of-Thought approach in mitigating hallucinations in GPT-4V(ision), we introduced the prompt "Let’s think step by step" alongside the original prompt and reported the results in Table 5. Additionally, we re-illustrate the example of the solar system with factual bias in Figure 9. Although CoT demonstrates enhanced language reasoning capabilities, it still fails to make a correct response. As indicated in Table 5, while the CoT prompting technique in GPT-4V(ision) shows a reduction of 5.7% in hallucinations associated with regional bias, it still fails to rectify hallucinations in most cases.

One possible explanation for this limited success is the visual limitation inherent to GPT-4V(ision). When GPT-4V(ision) encounters difficulties in comprehending images or utilizing them to respond to questions, the ineffectiveness of CoT is not unexpected. CoT was primarily designed to enhance language reasoning and may not suffice to address challenges in the vision component.

In summary, both the self-correction mechanism and Chain-of-Thought techniques do not effectively address the bias and interference challenges presented in the Bingo benchmark, highlighting the need for further research and innovations to tackle these persistent issues in vision-language models.

Related Work

In VLMs, the term "hallucination" typically refers to situations where the generated responses contain information that is not present in the visual content Rohrbach et al. (2018); Wang et al. (2023); Zhou et al. (2023). Traditional methods to address VLM hallucination include leveraging fine-grained contrastive learning Zeng et al. (2021), feature fusion Biten et al. (2022), and data augmentation (Kim et al., 2023). Recent developments in autoregressive large-scale VLM models, such as LLaVA, which integrate large language models with visual modality, have also encountered the challenge of hallucination. Recent studies have commenced investigations into hallucination issues within these autoregressive large-scale VLM models. This includes research on hallucination evaluation and detection Li et al. (2023); Wang et al. (2023), and hallucination mitigation Yin et al. (2023); Gunjal et al. (2023); Zhou et al. (2023). Concurrently, several studies have also highlighted the issue of hallucination in GPT-4V(ision) Shi et al. (2023); Liu et al. (2023a); Wu et al. (2023). Unlike prior works that focus on hallucination evaluation or mitigation, this work provides a comprehensive study to understand the causes of hallucinations in GPT-4V(ision) and other VLMs, introducing a new benchmark for this purpose.

Empirical Analysis of GPT-4V(ision)

The GPT series, developed by OpenAI, has demonstrated significant capabilities across various domains. The recent introduction of GPT-4V(ision) OpenAI (2023a, b, c) has notably enhanced GPT-4’s ability to connect visual and textual information, generating considerable interest among researchers due to its exceptional performance. For example, Yang et al. (2023) highlighted GPT-4V(ision)’s outstanding performance across various visual comprehension and reasoning tasks. However, it’s important to note that GPT-4V(ision) faces challenges in terms of generating hallucinations or producing erroneous responses. This issue is discussed in a few concurrent evaluations Wu et al. (2023); Zhang et al. (2023); Shi et al. (2023); Liu et al. (2023a), where they explore various aspects of GPT-4V(ision)’s capabilities, including solving visual puzzles, cross-modal interactions, character recognition, and handling of visual illusions. Nevertheless, none of these studies systematically categorized and analyzed the reasons behind the occurrence of hallucinations in GPT-4V(ision).

Conclusion

In this paper, we introduce the Bias and Interference Challenges in Visual Language Models (Bingo) benchmark, which focuses on analyzing hallucinations in VLMs, particularly in GPT-4V(ision). Our experiments reveal that although GPT-4V(ision) demonstrates impressive vision-language understanding abilities, it tends to generate hallucinatory responses (1) when dealing with specific types of images (bias) and (2) when subject to interference in judgment. Furthermore, we explore two strategies to address these challenges: self-correction and chain-of-thought. However, these approaches fall short of completely rectifying hallucinations in GPT-4V(ision) when facing bias and interference challenges. The findings presented in this paper enhance our understanding of the reliability of GPT-4V(ision) and other Vision-Language Models (VLMs).

Our benchmark focuses on a few metrics and tasks as a starting point. We will continue to woark on expanding the dataset and metrics. Our data curation also relies on human judgements, which may have its own biases; we try to mitigate this by having multiple researchers curate and evaluate the results.

References

Appendix A Additional Examples of Bias and Interference

In this section, we present more detailed examples of bias and interference in GPT-4V(ision).