A Survey of Hallucination in Large Foundation Models

Vipula Rawte, Amit Sheth, Amitava Das

Introduction

Foundation Models (FMs), exemplified by GPT-3 Brown et al. (2020) and Stable Diffusion Rombach et al. (2022), marks the commencement of a novel era in the realm of machine learning and generative artificial intelligence. Researchers introduced the term “foundation model” to describe machine learning models that are trained on extensive, diverse, and unlabeled data, enabling them to proficiently handle a wide array of general tasks. These tasks encompass language comprehension, text and image generation, and natural language conversation.

Foundation models refer to massive AI models trained on extensive volumes of unlabeled data, typically through self-supervised learning. This training approach yields versatile models capable of excelling in a diverse range of tasks, including image classification, natural language processing, and question-answering, achieving remarkable levels of accuracy.

These models excel in tasks involving generative abilities and human interaction, such as generating marketing content or producing intricate artwork based on minimal prompts. However, adapting and implementing these models for enterprise applications can present certain difficulties Bommasani et al. (2021).

2 What is Hallucination in Foundation Model?

Hallucination in the context of a foundation model refers to a situation where the model generates content that is not based on factual or accurate information. Hallucination can occur when the model produces text that includes details, facts, or claims that are fictional, misleading, or entirely fabricated, rather than providing reliable and truthful information.

This issue arises due to the model’s ability to generate plausible-sounding text based on patterns it has learned from its training data, even if the generated content does not align with reality. Hallucination can be unintentional and may result from various factors, including biases in the training data, the model’s lack of access to real-time or up-to-date information, or the inherent limitations of the model in comprehending and generating contextually accurate responses.

Addressing hallucination in foundation models and LLMs is crucial, especially in applications where factual accuracy is paramount, such as journalism, healthcare, and legal contexts. Researchers and developers are actively working on techniques to mitigate hallucinations and improve the reliability and trustworthiness of these models. With the recent rise in this problem Fig. 2, it has become even more critical to address them.

3 Why this survey?

In recent times, there has been a significant surge of interest in LFMs within both academic and industrial sectors. Additionally, one of their main challenges is hallucination. The survey in Ji et al. (2023) describes hallucination in natural language generation. In the era of large models, Zhang et al. (2023c) have done another great timely survey studying hallucination in LLMs. However, besides not only in LLMs, the problem of hallucination also exists in other foundation models such as image, video, and audio as well. Thus, in this paper, we do the first comprehensive survey of hallucination across all major modalities of foundation models.

The contributions of this survey paper are as follows:

We succinctly categorize the existing works in the area of hallucination in LFMs, as shown in Fig. 1.

We offer an extensive examination of large foundation models (LFMs) in Sections 2, 3, 4 and 5.

We cover all the important aspects such as i. detection, ii. mitigation, iii. tasks, iv. datasets, and v. evaluation metrics, given in LABEL:tab:big-table.

We finally also provide our views and possible future direction in this area. We will regularly update the associated open-source resources, available for access at https://github.com/vr25/hallucination-foundation-model-survey

3.2 Classification of Hallucination

As shown in Fig. 1, we broadly classify the LFMs into four types as follows: i. Text, ii. Image, iii. video, and iv. Audio.

The paper follows the following structure. Based on the above classification, we describe the hallucination and mitigation techniques for all four modalities in: i. text (Section 2), ii. image (Section 3), iii. video (Section 4), and iv. audio (Section 5). In Section 6, we briefly discuss how hallucinations are NOT always bad, and hence, in the creative domain, they can be well-suited to producing artwork. Finally, we give some possible future directions for addressing this issue along with a conclusion in Section 7.

Hallucination in Large Language Models

As shown in Fig. 4, hallucination occurs when the LLM produces fabricated responses.

SELFCHECKGPT Manakul et al. (2023), is a method for zero-resource black-box hallucination detection in generative LLMs. This technique focuses on identifying instances where these models generate inaccurate or unverified information without relying on additional resources or labeled data. It aims to enhance the trustworthiness and reliability of LLMs by providing a mechanism to detect and address hallucinations without external guidance or datasets. Self-contradictory hallucinations in LLMs are explored in Mündler et al. (2023). and addresses them through evaluation, detection, and mitigation techniques. It refers to situations where LLMs generate text that contradicts itself, leading to unreliable or nonsensical outputs. This work presents methods to evaluate the occurrence of such hallucinations, detect them in LLM-generated text, and mitigate their impact to improve the overall quality and trustworthiness of LLM-generated content.

PURR Chen et al. (2023) is a method designed to efficiently edit and correct hallucinations in language models. PURR leverages denoising language model corruptions to identify and rectify these hallucinations effectively. This approach aims to enhance the quality and accuracy of language model outputs by reducing the prevalence of hallucinated content.

Hallucinations are commonly linked to knowledge gaps in language models (LMs). However, Zhang et al. (2023a) proposed a hypothesis that in certain instances when language models attempt to rationalize previously generated hallucinations, they may produce false statements that they can independently identify as inaccurate. Thus, they created three question-answering datasets where ChatGPT and GPT-4 frequently provide incorrect answers and accompany them with explanations that contain at least one false assertion.

HaluEval Li et al. (2023b), is a comprehensive benchmark designed for evaluating hallucination in LLMs. It serves as a tool to systematically assess LLMs’ performance in terms of hallucination across various domains and languages, helping researchers and developers gauge and improve the reliability of these models.

Using interactive question-knowledge alignment, Zhang et al. (2023b) presents a method for mitigating language model hallucination Their proposed approach focuses on aligning generated text with relevant factual knowledge, enabling users to interactively guide the model’s responses to produce more accurate and reliable information. This technique aims to improve the quality and factuality of language model outputs by involving users in the alignment process. LLM-AUGMENTER Peng et al. (2023) improves LLMs using external knowledge and automated feedback. It highlights the need to address the limitations and potential factual errors in LLM-generated content. This method involves incorporating external knowledge sources and automated feedback mechanisms to enhance the accuracy and reliability of LLM outputs. By doing so, the paper aims to mitigate factual inaccuracies and improve the overall quality of LLM-generated text. Similarly, Li et al. (2023d) introduces a framework called “Chain of Knowledge” for grounding LLMs with structured knowledge bases. Grounding refers to the process of connecting LLM-generated text with structured knowledge to improve factual accuracy and reliability. The framework utilizes a hierarchical approach, chaining multiple knowledge sources together to provide context and enhance the understanding of LLMs. This approach aims to improve the alignment of LLM-generated content with structured knowledge, reducing the risk of generating inaccurate or hallucinated information.

Smaller, open-source LLMs with fewer parameters often experience significant hallucination issues compared to their larger counterparts Elaraby et al. (2023). This work focuses on evaluating and mitigating hallucinations in BLOOM 7B, which represents weaker open-source LLMs used in research and commercial applications. They introduce HALOCHECK, a lightweight knowledge-free framework designed to assess the extent of hallucinations in LLMs. Additionally, it explores methods like knowledge injection and teacher-student approaches to reduce hallucination problems in low-parameter LLMs.

Moreover, the risks associated with LLMs can be mitigated by drawing parallels with web systems Huang and Chang (2023). It highlights the absence of a critical element, “citation,” in LLMs, which could improve content transparency, and verifiability, and address intellectual property and ethical concerns.

“Dehallucinating” refers to reducing the generation of inaccurate or hallucinated information by LLMs. Dehallucinating LLMs using formal methods guided by iterative prompting is presented in Jha et al. (2023). They employ formal methods to guide the generation process through iterative prompts, aiming to improve the accuracy and reliability of LLM outputs. This method is designed to mitigate the issues of hallucination and enhance the trustworthiness of LLM-generated content.

2 Multilingual LLMs

Large-scale multilingual machine translation systems have shown impressive capabilities in directly translating between numerous languages, making them attractive for real-world applications. However, these models can generate hallucinated translations, which pose trust and safety issues when deployed. Existing research on hallucinations has mainly focused on small bilingual models for high-resource languages, leaving a gap in understanding hallucinations in massively multilingual models across diverse translation scenarios.

To address this gap, Pfeiffer et al. (2023) conducted a comprehensive analysis on both the M2M family of conventional neural machine translation models and ChatGPT, a versatile LLM that can be prompted for translation. The investigation covers a wide range of conditions, including over 100 translation directions, various resource levels, and languages beyond English-centric pairs.

3 Domain-specific LLMs

Hallucinations in mission-critical areas such as medicine, banking, finance, law, and clinical settings refer to instances where false or inaccurate information is generated or perceived, potentially leading to serious consequences. In these sectors, reliability and accuracy are paramount, and any form of hallucination, whether in data, analysis, or decision-making, can have significant and detrimental effects on outcomes and operations. Consequently, robust measures and systems are essential to minimize and prevent hallucinations in these high-stakes domains.

The issue of hallucinations in LLMs, particularly in the medical field, where generating plausible yet inaccurate information can be detrimental. To tackle this problem, Umapathi et al. (2023) introduces a new benchmark and dataset called Med-HALT (Medical Domain Hallucination Test). It is specifically designed to evaluate and mitigate hallucinations in LLMs. It comprises a diverse multinational dataset sourced from medical examinations across different countries and includes innovative testing methods. Med-HALT consists of two categories of tests: reasoning and memory-based hallucination tests, aimed at assessing LLMs’ problem-solving and information retrieval capabilities in medical contexts.

ChatLaw Cui et al. (2023), is an open-source LLM specialized for the legal domain. To ensure high-quality data, the authors created a meticulously designed legal domain fine-tuning dataset. To address the issue of model hallucinations during legal data screening, they propose a method that combines vector database retrieval with keyword retrieval. This approach effectively reduces inaccuracies that may arise when solely relying on vector database retrieval for reference data retrieval in legal contexts.

Hallucination in Large Image Models

Contrastive learning models, employing a Siamese structure Wu et al. (2023), have displayed impressive performance in self-supervised learning. Their success hinges on two crucial conditions: the presence of a sufficient number of positive pairs and the existence of ample variations among them. Without meeting these conditions, these frameworks may lack meaningful semantic distinctions and become susceptible to overfitting. To tackle these challenges, we introduce the Hallucinator, which efficiently generates additional positive samples to enhance contrast. The Hallucinator is differentiable, operating in the feature space, making it amenable to direct optimization within the pre-training task and incurring minimal computational overhead.

Efforts to enhance LVLMs for complex multimodal tasks, inspired by LLMs, face a significant challenge: object hallucination, where LVLMs generate inconsistent objects in descriptions. This study Li et al. (2023e) systematically investigates object hallucination in LVLMs and finds it’s a common issue. Visual instructions, especially frequently occurring or co-occurring objects, influence this problem. Existing evaluation methods are also affected by input instructions and LVLM generation styles. To address this, the study introduces an improved evaluation method called POPE, providing a more stable and flexible assessment of object hallucination in LVLMs.

Instruction-tuned Large Vision Language Models (LVLMs) have made significant progress in handling various multimodal tasks, including Visual Question Answering (VQA). However, generating detailed and visually accurate responses remains a challenge for these models. Even state-of-the-art LVLMs like InstructBLIP exhibit a high rate of hallucinatory text, comprising 30 percent of non-existent objects, inaccurate descriptions, and erroneous relationships. To tackle this issue, the study Gunjal et al. (2023)introduces MHalDetect1, a Multimodal Hallucination Detection Dataset designed for training and evaluating models aimed at detecting and preventing hallucinations. M-HalDetect contains 16,000 finely detailed annotations on VQA examples, making it the first comprehensive dataset for detecting hallucinations in detailed image descriptions.

Hallucination in Large Video Models

Hallucinations can occur when the model makes incorrect or imaginative assumptions about the video frames, leading to the creation of artificial or erroneous visual information Fig. 5.

The challenge of understanding scene affordances is tackled by introducing a method for inserting people into scenes in a lifelike manner Kulal et al. (2023). Using an image of a scene with a marked area and an image of a person, the model seamlessly integrates the person into the scene while considering the scene’s characteristics. The model is capable of deducing realistic poses based on the scene context, adjusting the person’s pose accordingly, and ensuring a visually pleasing composition. The self-supervised training enables the model to generate a variety of plausible poses while respecting the scene’s context. Additionally, the model can also generate lifelike people and scenes on its own, allowing for interactive editing.

VideoChat Li et al. (2023c), is a comprehensive system for understanding videos with a chat-oriented approach. VideoChat combines foundational video models with LLMs using an adaptable neural interface, showcasing exceptional abilities in understanding space, time, event localization, and inferring cause-and-effect relationships. To fine-tune this system effectively, they introduced a dataset specifically designed for video-based instruction, comprising thousands of videos paired with detailed descriptions and conversations. This dataset places emphasis on skills like spatiotemporal reasoning and causal relationships, making it a valuable resource for training chat-oriented video understanding systems.

Recent advances in video inpainting have been notable Yu et al. (2023), particularly in cases where explicit guidance like optical flow can help propagate missing pixels across frames. However, challenges arise when cross-frame information is lacking, leading to shortcomings. So, instead of borrowing pixels from other frames, the model focuses on addressing the reverse problem. This work introduces a dual-modality-compatible inpainting framework called Deficiency-aware Masked Transformer (DMT). Pretraining an image inpainting model to serve as a prior for training the video model has an advantage in improving the handling of situations where information is deficient.

Video captioning aims to describe video events using natural language, but it often introduces factual errors that degrade text quality. While factuality consistency has been studied extensively in text-to-text tasks, it received less attention in vision-based text generation. In this research Liu and Wan (2023), the authors conducted a thorough human evaluation of factuality in video captioning, revealing that 57.0% of model-generated sentences contain factual errors. Existing evaluation metrics, mainly based on n-gram matching, do not align well with human assessments. To address this issue, they introduced a model-based factuality metric called FactVC, which outperforms previous metrics in assessing factuality in video captioning.

Hallucination in Large Audio Models

Automatic music captioning, which generates text descriptions for music tracks, has the potential to enhance the organization of vast musical data. However, researchers encounter challenges due to the limited size and expensive collection process of existing music-language datasets. To address this scarcity, Doh et al. (2023) used LLMs to generate descriptions from extensive tag datasets. They created a dataset known as LP-MusicCaps, comprising around 2.2 million captions paired with 0.5 million audio clips. They also conducted a comprehensive evaluation of this large-scale music captioning dataset using various quantitative natural language processing metrics and human assessment. They trained a transformer-based music captioning model on this dataset and evaluated its performance in zero-shot and transfer-learning scenarios.

Ideally, the video should enhance the audio, and in Li et al. (2023a), they have used an advanced language model for data augmentation without human labeling. Additionally, they utilized an audio encoding model to efficiently adapt a pre-trained text-to-image generation model for text-to-audio generation.

Hallucination is not always harmful: A different perspective

Suggesting an alternative viewpoint, Wiggers (2023) discusses how hallucinating models could serve as “collaborative creative partners,” offering outputs that may not be entirely grounded in fact but still provide valuable threads to explore. Leveraging hallucination creatively can lead to results or novel combinations of ideas that might not readily occur to most individuals.

“Hallucinations” become problematic when the statements generated are factually inaccurate or contravene universal human, societal, or particular cultural norms. This is especially critical in situations where an individual relies on the LLM to provide expert knowledge. However, in the context of creative or artistic endeavors, the capacity to generate unforeseen outcomes can be quite advantageous. Unexpected responses to queries can surprise humans and stimulate the discovery of novel idea connections.

Conclusion and Future Directions

We concisely classify the existing research in the field of hallucination within LFMs. We provide an in-depth analysis of these LFMs, encompassing critical aspects including 1. Detection, 2. Mitigation, 3. Tasks, 4. Datasets, and 5. Evaluation metrics.

Some possible future directions to address the hallucination challenge in the LFMs are given below.

In the context of natural language processing and machine learning, hallucination refers to the generation of incorrect or fabricated information by AI models. This can be a significant problem, especially in applications like text generation, where the goal is to provide accurate and reliable information. Here are some potential future directions in the automated evaluation of hallucination:

Researchers can work on creating specialized evaluation metrics that are capable of detecting hallucination in generated content. These metrics may consider factors such as factual accuracy, coherence, and consistency. Advanced machine learning models could be trained to assess generated text against these metrics.

Combining human judgment with automated evaluation systems can be a promising direction. Crowdsourcing platforms can be used to gather human assessments of AI-generated content, which can then be used to train models for automated evaluation. This hybrid approach can help in capturing nuances that are challenging for automated systems alone.

Researchers can develop adversarial testing methodologies where AI systems are exposed to specially crafted inputs designed to trigger hallucination. This can help in identifying weaknesses in AI models and improving their robustness against hallucination.

Fine-tuning pre-trained language models specifically to reduce hallucination is another potential direction. Models can be fine-tuned on datasets that emphasize fact-checking and accuracy to encourage the generation of more reliable content.

2 Improving Detection and Mitigation Strategies with Curated Sources of Knowledge

Detecting and mitigating issues like bias, misinformation, and low-quality content in AI-generated text is crucial for responsible AI development. Curated sources of knowledge can play a significant role in achieving this. Here are some future directions:

Incorporating knowledge graphs and curated knowledge bases into AI models can enhance their understanding of factual information and relationships between concepts. This can aid in both content generation and fact-checking.

Develop specialized models that focus on fact-checking and content verification. These models can use curated sources of knowledge to cross-reference generated content and identify inaccuracies or inconsistencies.

Curated sources of knowledge can be used to train AI models to recognize and reduce biases in generated content. AI systems can be programmed to check content for potential biases and suggest more balanced alternatives.

Continuously update and refine curated knowledge sources through active learning. AI systems can be designed to seek human input and validation for ambiguous or new information, thus improving the quality of curated knowledge.

Future directions may also involve the development of ethical guidelines and regulatory frameworks for the use of curated knowledge sources in AI development. This could ensure responsible and transparent use of curated knowledge to mitigate potential risks.

In summary, these future directions aim to address the challenges of hallucination detection and mitigation, as well as the responsible use of curated knowledge to enhance the quality and reliability of AI-generated content. They involve a combination of advanced machine learning techniques, human-AI collaboration, and ethical considerations to ensure AI systems produce accurate and trustworthy information.

References