CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models
Junho Kim, Hyunjun Kim, Yeonju Kim, Yong Man Ro
Introduction
With recent advancements of Large Language Models (LLMs) , Large Multi-modal Models (LMMs), sometimes referred as Large Vision-Language Models , have been drawn great attention for their natural multi-modal interaction with users through back-and-forth conversations. Leveraging their robust generation capabilities, various pioneering tasks in pre-LMM era such as image captioning , visual question answering , object detection , etc., have been integrated into a single task rather than treated as sub-tasks and achieved significant milestones . However, at the same time, the hallucination issue has become one of the emerging problems when adopting LMMs into real-world applications due to their potential spurious generation in critical areas.
Here, unlike hallucination studies in LLMs mainly focusing on factuality hallucination originated from the language knowledge, the hallucination problem in LMMs refers cross-modal inconsistency between the given visual contents and the generated responses for the user instructions. After the seminal works giving eyes to LLMs to understand visual contents with visual instruction tuning, numerous cutting-edge LMMs actively have been proposed. Albeit the scaling laws following more stronger versatile vision models , higher resolution , deeper alignment layers , larger model sizes, etc., LMMs still suffer from generating responses that seem plausible but are factually incorrect for the given visual contents.
The origin of LMM hallucination is an intertwined problem for their inherent training paradigm, which involves alignment projection matching during the pre-training, followed by fine-tuning with the limited instruction-following data. Several approaches, aimed for mitigating hallucinatory effects, have addressed the inconsistent issues in the context of data-associated solution , scaling model architectures , or additional RL-based training . Among them, reactive methods intervene the decoding phase of LMMs’ inference and alleviate undesired responses. Motivated by Li et al. that have proposed contrastive decoding (CD) method between expert and amateur language models, recent CD-based approaches in LMMs have proposed several ways of contrasting model responses from visual inputs with their counterparts (e.g., visual contamination , image-biased models , or fine-grained visual information ).
Our research question begins with "How effectively do contemporary LMMs capture visual evidences in their descriptive responses, and what information must be curbed to produce informative and consistent responses?". As illustrated in Fig 1 (bottom-left), when asking LMMs to generate a comprehensive description for visual content, the output seemingly generates detailed description effectively, but a closer examination often reveals missed fined-grained information or hallucinatory instances in the responses. By recursively referring these incomplete descriptions generated by the models themselves, we aim to restrict the incorrect information flow during the generation phase and enhance the alignment of the model responses grounded in true visual evidences.
In this paper, we introduce a novel training-free contrastive decoding method, COuntering DEscription Contrastvie Decoding (CODE), designed to use self-generated descriptions as a contrastive reference to mitigate hallucination issues in LMMs. The core idea of our proposed method is on harnessing the self-generated descriptions which possibly encompass both factual evidence and hallucinatory information from visual contents as a look-up reference for response correction. Specifically, within our contrastive framework as illustrated in Fig. 1, the comprehensive descriptions from model itself alternatively propagate to visual input tokens and contrast the discrepancy with the logits from actual visual contents to enhance next-token prediction. In addition, we introduce a dynamic restriction strategy that enables adaptive control of information flow during the auto-regressive decoding phase, taking into account both token-level predictions and their distribution within the vocabulary set.
By conducting extensive experiments and analyses on prevailing cutting-edge LMMs , we corroborate the effectiveness of our method in reducing hallucination and enhancing the coherence and informativeness in various benchmarks . Our decoding method can be seamlessly integrated into existing LMMs by simply substituting the image tokens with self-generated descriptions in a training-free manner.
Our contribution can be summarized into three-fold as follows:
We introduce COuntering DEscription Contrastive Decoding (CODE), a training-free decoding strategy that employs self-generated descriptions to minimize hallucinations in LMMs. By contrasting logit information from descriptions with actual visual contents, CODE enhances visual consistency and coherence in the model responses.
Our approach incorporates dynamic restriction strategies within the contrastive decoding phase. It selectively regulates the information flow by adjusting token-level predictions based on their distribution in the vocabulary, thus ensuring more contextual responses.
We validate the effectiveness of our decoding method across various benchmarks using cutting-edge LMMs. The results demonstrate that CODE significantly reduces hallucination while enhancing the relevance and informativeness in the responses.
Related Work
After the emergence of large-scaled LLMs that can interact with users with question-answer chat, various of vision+LLM studies— i.e., LMMs, have been proposed to integrate the robust linguistic capability into the visual understanding and reasoning in diverse vision-language task. As earlier works such as LLaVA , Instruct-BLIP , and MiniGPT-4 , which utilized visual instruction tuning, have bridged two modalities of vision and language through fine-tuning with a learnable query, exemplified by Q-Former or projection layer-based alignments . To enhance cross-modal consistency in vision-language representation, recent works have proposed several solutions to address the underlying weaknesses in both modalities: (i) utilizing higher-resolution visual inputs , (ii) deploying Mixture-of-Expert (MoE) concepts integrating versatile vision models , (iii) improving weak alignment interface , or (iv) adopting larger LLMs to scale up the language model prior .
Hallucination Issue, Harming Cross-modal Consistency
Despite of the endeavor developments of LMMs, they cannot be free from cross-modal inconsistency between the visual contents and their generated responses, so-called hallucination . This not only leads to performance degradation but also provokes an over-reliance issue, resulting in incorrect model responses that are not grounded in true visual evidence. This critical concern regarding response trustworthiness and model reliability hinders the adoption of LMMs in real-world applications. To mitigate the hallucination problem, diverse works have been proposed employing additional training on curated datasets or reinforcement learning under feedback systems . Among them, by intervening during the response generation, decoding-based approaches are introduced to encourage models to represent more precise responses. We refer to readers for more comprehensive survey papers addressing hallucination in LMMs. Our work is in line with CD-based approaches that utilize logit discrepancy from counterpart outputs to enhance coherence. Unlike the previous works that focus on twisting visual information, we utilize self-generated description as contrasting visual counterpart and correct hallucinatory responses based on the model understanding.
Proposed Method
After the seminal works in natural language processing have introduced Contrastive Decoding (CD) mechanism, which considering information disparities between expert and amateur models for more coherence and informativeness, various works have deployed this strategy into LMMs by twisting visual contents or model information for the contrastive approach. The next-token probability from CD can be generally formulated as follows:
where and indicates visual counterparts and sub-optimal amateur model, respectively— note that can be regarded as self-correction. Intuitively, the objective of CD is amplifying model outputs by reflecting information deviation between top candidate log-probabilities. Therefore, the selection of logit counterparts for referring is the key challenge for high-quality and consistent responses during the contrastive decoding frameworks.
1 Comprehensive Image Description as Visual Counterpart
As the visual counterpart, we deploy comprehensive image descriptions generated from the model as contrasting reference, which are inevitably less informative than the visual contents themselves. Our motivation is on the innate difference of information density between vision and language . While vision information exhibits relatively less redundancy for spatial signals— e.g., human can visually recognize objects with a few masked patches on images, languages contain high-entropy information, resulting in spans that are more semantic and information-dense than visual signals. That is, it is difficult to infer blanked-out words— e.g., "I went to the store to buy some ____". Building upon the property of each modality, we first delve into the self-generated model responses with specific instruction for the visual contents to elicit the encompassed visual evidences in the representation space and assess its potential subject role as a visual counterpart for contrastive decoding.
As illustrated in Fig. 1, we input a query instruction into the model to generate a comprehensive visual description for the given visual content (please see the detailed instruction for the self-generated description in Appendix. A). Then, we exploit the generated description as recursive visual inputs, replacing the position of image tokens in the model input sequence (e.g.,
2 COuntering DEscription Contrastive Decoding
Based on our analysis in sec. 3.1, we can obtain a pair of the visual content and its comprehensive description , such that corresponds to . By contrasting the logit variation between the paired information into the model response generation, we can formulate the next-word prediction using our proposed method, COuntering DEscription Contrastive Decoding (CODE):
Here, unlike the previous approaches that restrict the logit variations with fixed value as in Eqn. 1, we present a dynamic restriction for the logit variations by comparing the information between visual contents and their comprehensive descriptions. Revisiting the role of , it determines whether to promote or curb information from logit variation, thus directly influencing next-token generation— higher value results in more aggressive adjustment for the variations. However, when confronting that both and yield similar logit score on the correct token, the variation gets closer to zero, thereby the next-token prediction can be unexpectedly reversed if other tokens get rewarded than the correct token with a fixed on a token-by-token basis. Although this aligns with the initial intent of CD, a more robust selector is necessary to effectively restrict the logit information flow.
Accordingly, our method predicts next-token not only at the individual token-level but also considering its distribution across the entire vocabulary set, enabling dynamic control of the information flow. To measure the relative entropy between the token distributions from visual contents and its comprehensive description at time step , we deploy Bounded Divergence () , which is a type of statistical distance that ensures symmetric and bounded measure:
where and equals , if and only if , and denotes a smoothing parameter. Here, the upper-bound of the divergence apparently exists, such that , due to the following condition .
We define the dynamic restriction as , where it enables a token-wise feedback control that adjusts the information weighting with respect to the closeness of the two distributions. The major role of the restriction term is maintaining a balance in the logit variation for the observed prediction disparities between the and distributions. That is, when the distributions are close enough (i.e., ), the value of approaches zero, indicating minimal divergence. That is, approaches 1, allowing for higher amplification of logit variations in predicting the next-token outputs. This adjustment reflects the increased reliability of predictions when the two distributions from and are closely aligned. On the other hand, for the dissimilar distributions, decreases towards zero, compelling the model to restrict information flow from the variation. This reduction limits the potential for introducing erroneous or less probable predictions by focusing more on visual information, thereby maintaining coherence in the output when the model’s understanding of the visual content significantly deviates from its textual description.
3 Adaptive Information Constraint
One major challenge in contrastive-based decoding is the scenario where implausible tokens are rewarded, even when predictions are made with low confidence. This issue can also arise in our method, particularly when token distributions derived from textual descriptions provide more confidence than the visual content, ironically undermining the most predictive tokens. To address it, Li et al. have introduced an adaptive plausibility constraint, which filters out less plausible tokens by truncating them based on the maximum token confidence from the expert model. While this approach simply penalizes false positive tokens in the candidate pool, it may also have unintended side effects by prematurely applying a cutoff threshold to lower-confidence tokens. Specifically, early threshold settings can sometimes eliminate the possibility of identifying correct token predictions, which might otherwise be dismissed in a pool considered to contain mostly false negatives.
Improving the previous constraint , we present adaptive information constraint () designed to dynamically retain tokens that may be informative despite their lower confidence. By comparing prediction distributions between and , we filter out less relevant tokens from the candidate pool as follows:
where dynamically regulate the token candidate pool utilizing the divergence term in Eqn. 3, defined as . This strategy can expand the token searching pool when the next-token prediction, derived from both visual content and comprehensive description, shows a similar distribution yet uncertainty in selecting the candidate token (i.e., false negatives). Finally, we only consider the next-token prediction within , and for the tokens satisfying , we set their logits to to filter out from the candidate pool. Please see comprehensive Algorithm. 1 in Appendix. B.
Experiments
To validate the efficacy of our method over various LMM families and sizes, we implemented our method on contemporary LMMs: LLaVA-1.5 (13B) , Emu2-Chat (14B) , InternLM-XComposer2 (7B) , LLaVA-NeXT (34B) , Yi-VL (34B) , and InternVL 1.5 (26B) . We compared our method with five baseline decoding strategies. For the regular decoding strategies, we used greedy decoding, Nucleus sampling , and beam search decoding. Additionally, we selected OPERA and VCD for contrastive decoding method, which designed to mitigate hallucinations with contrastive frameworks. We used the default parameter settings for all methods, where top-p value and temperature for Nucleus sampling, the number of window size for searching is (i.e., num-beams ) for both beam search decoding and OPERA, and CD-, CD- for VCD, and for our method. Note that OPERA inference requires too much memory especially for LLaVA-NeXT (34B), so that we excluded OPERA results for this model.
Benchmarks and Evaluation Metrics
The benchmarks for evaluating hallucinations in LMMs can be broadly categorized into discriminative and generative streams. The discriminative type assesses hallucinations by evaluating the predicted answer among given options (e.g., multiple choice or yes/no question), while generative benchmarks typically employ more advanced language models (e.g., GPT-aided evaluation) to rate the subject model descriptions. Within this taxonomy, we carefully select benchmarks to test baselines. Please see Appendix. C for benchmark details.
As discriminative benchmarks, we utilize mainly three datasets for detailed evaluation. Specifically, POPE is a commonly used benchmark for detecting object hallucination by converting object annotations sourced from MSCOCO . Under the three different subsets: random, popular, and adversarial, the metric for POPE measures binary classification performance for simple yes/no questions. MMVP aims to evaluate the understanding of visual details for different visual patterns using paired classification accuracy. Due to its evaluation design, which involves comparing two similar CLIP-blind image pairs, MMVP requires LMMs to capture subtle visual differences. RealworldQA is the most recent dataset tailored to assess the capability of LMMs in basic real-world spatial understanding, using the accuracy metric within multiple-choice questions.
We use three benchmarks for generative benchmarks, extending the evaluation scope to include open-ended captioning tasks beyond merely assessing classification within given answer options. Generally, ChatGPT is used to score the quality of the model-generated sentences. The metric for both LLaVA-QA90 and LLaVA-Bench (In-the-Wild) is score ratio, where model responses rated from GPT-4 are divided by GPT-4 answers such that , where all scores are rated by GPT-4. It has three types of questions: conversation, detailed description, and complex reasoning. MMHal-Bench evaluates the degree of hallucination for the various question types: object attribute, adversarial object, comparison, counting, spatial relation, environment, holistic description, and others. GPT-4 measures the severity of hallucination in a range of to and the higher score denotes less hallucination.
1 Evaluation Results
We summarize our comprehensive experimental results across six LMMs, six decoding methods, on six benchmarks in a spider chart format for visibility of the improvements in Fig. 3. As in the figure, CODE generally shows competent performance and consistent results among different measurements and benchmarks. We delve into each result in detail in the subsections below.
Results on Discriminative Benchmarks
Previous CD-based methods require heuristic choices to control the degree of amplification for logit variation and penalizing parameter that filter out the implausible next-tokens in adaptive plausibility constraint. To tackle it, we proposed two regulation methods that can dynamically control information flow in CODE and token candidate pool, respectively: (i) dynamic restriction (DR), in Eqn. 2 and (ii) adaptive information constraint (AIC), in Eqn. 4. To validate the effectiveness of such adaptive regulations built in CODE, we conducted ablation study on them.
We implement baselines with same CODE framework using self-generated descriptions as visual counterparts, but with default settings of and . As in Table. 3, either use of DR or AIC can enhance the benchmarks than the fixed and . Our CODE implementation that utilizes both DR and AIC to dynamically restrict information flow exhibits the best results among the baselines.
Computational Analysis
We utilize comprehensive descriptions from models as additional information for contrasting with visual contents, thereby leading to computational loads similar to other decoding methods. To analyze the computation, we compare the token throughput (token/s) and decoding latency (ms/token) with other CD-based methods on NVIDIA RTX A6000 GPUs as in Table. 4.
Token-level Case Study
As illustrated in Fig. 4, to verify whether the proposed CODE effectively mitigates object hallucination, we analyze the output logit values of LMM at the token-level case study, with a greedy search as the baseline. The first two rows in the table indicate the original greedy decoded tokens which are elected based on high from the visual content and CODE output tokens, respectively. As in the figure, we can observe that visual hallucination occurs at the "Yoplait" token highlighted in red. For relatively easy tokens at the beginning of sentence, produces identical decisions maintaining consistency with , which indicates the amplification of logit variation is effectively adjusted due to similar prediction distributions from visual contents and description-only information. However, at the hallucination-occurred time step, logit scores are deviated between the two information, resulting in a more confusing state to identify between GT token "Fage" and hallucinatory "Yoplait". In our framework, "Fage" is relatively more amplified from to than "Yoplait", which changes from to . By simultaneously considering both token-level and distributional prediction over the vocabulary, CODE changes the wrong next-token output to correct one, mitigating hallucination. For more case studies, please refer Appendix. D.
Additional Experiments on In-the-Wild
Contemporary open-sourced LMMs are fine-tuned with various combinations of vision-language datasets , mostly composed of COCO-sourced visual images and their curated instruction. Although the existing hallucination benchmarks intentionally convert question queries to assess model robustness against inconsistency, the visual contents in benchmarks are limited to in-distribution COCO images. To validate our method in more challenging and real-world scenarios, we compared baselines on LLaVA-Bench (In-the-wild) and RealworldQA as in Fig. 5 and achieved competent performance (case studies in Appendix. F).
Albeit the computational analysis in Table. 4, as one of limitations, our contrastive decoding method requires additional computational resources than the use of vanilla decoding. However, considering an essential ongoing research topics and developments aimed at mitigating the negative effects of hallucination problems in both LLMs and LMMs, our work contributes important societal impacts towards more real-world applicability and robust AI system.
Conclusion
We present COuntering DEscription Contrastive Decoding (CODE), a novel and training-free decoding method to mitigate hallucination in Large Multi-modal Models. By utilizing self-generated descriptions as corrective references during the decoding phase, CODE dynamically adjusts the information flow for next-token predictions, enhancing the coherence and informativeness of responses while reducing the cross-modal inconsistency. Extensive experiments demonstrate that CODE effectively decreases hallucinations across various benchmarks and contemporary LMMs, significantly improving contextual relevance and response alignment with visual contents.
Appendix A Instruction for Comprehensive Description
In generating a comprehensive description for the given visual content, we aim to obtain as detailed a description as possible, ensuring that the response fully spans the visual representation space, even though it may not be entirely feasible as discussed in 3.1. The specific prompt instruction used to describe the image contents in detail is described in Table. 5.
Appendix B Detailed Algorithm for CODE
We describe the complete details of CODE implementation for better understanding in Algorithm. 1.
Appendix C Benchmark Details
POPE is a widely used benchmark designed for evaluating object-level hallucination, which can be split into the three subset categories based on how to select object replacements: (i) random, randomly sampled objects (ii) popular, top- frequent objects not existing in the image, and (iii) adversarial, top- objects that have high co-occurrence. The number of images is and each image has questions along with the subsets, making a total of 9000 images. In this work, we only consider adversarial split, which is the most challenging subset in POPE benchmark.
MMVP includes images with different visual patterns that CLIP model struggles to identify the visual differences (CLIP-paired images): Orientation and Direction (\faCompass), Presence of Specific Features (\faSearch), State and Condition (\faSync), Quantity and Count (\faSortNumericUp), Positional and Relational Context (\faMapPin), Color and Appearance (\faPalette), Structural and Physical Characteristics (\faCogs), Text (\faFont), Viewpoint and Perspective (\faCamera). It follows multiple selection tests, but uses GPT-4 to map the model response to the answer options.
RealworldQA is recently introduced benchmarks for evaluating basic real-world understanding for multi-modal models. It consists of total anonymized outdoor (mostly taken from vehicles) and indoor images with multiple selection questions.
Generative Benchmarks
LLaVA-QA90 & LLaVA-Bench (In-the-Wild) consist of three subset response types for each image: (i) Conversation, which is conversation format between the user and assistant answering the vision-related questions for the given images, (ii) Detailed description, which requires detailed description for the given image scene, and (iii) Complex reasoning, which involves in-depth reasoning questions for the image. The former benchmarks sourced from COCO images (total images with questions), while the latter benchmarks are gathered from web for challenging domain situations (total images with questions).
MMHal-Bench is specially focused on penalizing hallucinations. It has total image-question pairs composed with question categories for objects: Object attribute (Attr), Adversarial object (Adv), Comparison (Comp), Counting (Count), Spatial relation (Rel), Environment (Env), Holistic description (Hol), and Others (Other). As like in the above LLaVA-Bench benchmarks, MMHal-Bench also utilize GPT-4 to analyze and rate the model responses and score in a range of to .
Appendix D Additional Token-level Case Study
As additional token-level case studies, we explore how the logit information changes during our CODE decoding phase in other examples. As a first example, LLaVA-NeXT struggles to distinguish between Haleakala National Park and Diamond Head, both located in Hawaii, and predicts the former during inference with the vanilla decoding method. Using our CODE decoding method, as shown in Fig. 6, the information flow of "Haleakala" token is curbed, inducing a token inversion to "Diamond", which matches the ground truth word. This occurs because the logit variation is dynamically adjusted based on both token-level and distributional information.
In another example illustrated in Fig. 7, we show how the adaptive information constraint prevent from rewarding implausible tokens, thus suppress hallucinatory prediction during contrastive decoding. The original tokens from predict the correct answer at the hallucinatory time step, highlighted in bold. In this case, the original prediction should be preserved, and token inversion should not occur. By our CODE decoding method, dynamically controls the adaptive information constraint, so that those hallucination tokens’(i,e., "four" and "dragon") logit values are cut off to and removed from candidate token pool.
Appendix E Further Discussion on Broader Impact
We proposed CODE, which can be seamlessly integrated into LMMs without additional training. There is still a lot of room for progress and mitigation of hallucination issues, as our method cannot assure 100% removal of hallucinations. However, by providing more coherent and contextually accurate responses, our work can potentially be integrated into real-world applications, making user interactions with AI in customer service, education, and personal assistance more effective and satisfying in the near future.
Furthermore, enhanced accuracy and reliability using our method can reduce hallucinations in LMMs, improving the accuracy of AI-generated descriptions in critical fields such as autonomous driving, robotics, healthcare, and augmented reality. This advancement not only enhances practical applications but also significantly benefits the research community working on hallucination, an area that is not yet fully explored. Our contributions can help pave the way for deeper understanding and new research directions to address these challenges, for more trustworthy AI systems.
Appendix F Additional Qualitative Results
Discriminative Capability Case Study: MMVP.
Appendix G Failure Cases
As discussed in Appendix. E, even if CODE shows competent performance along various benchmarks and LMM baselines. The hallucination cannot be eliminated. In this section, we attached some of failure cases to shed some light for future work direction. As illustrated in Table. LABEL:table:failure_cases, the baseline models fail to correct hallucination even with our CODE method. Upon closer examination of the model responses, it is evident that the hallucinatory responses tend to be biased towards language priors such as "strawberry-flavored" or "holding a glass of beer". Our approach mainly uses self-generated description as contrasting reference (i.e., close to the concept of self-correction), thus the strong assumption is on that the amateur model (comprehensive description) should generate not too much deviated responses from true answers.
In the failure examples, we can infer that the reliance on self-generated descriptions may not always suffice, especially when the descriptions themselves are biased or inaccurate. This indicates a need for integrating more robust mechanisms to verify and correct these biases. Additionally, enhancing the model’s understanding and processing of visual content could help mitigate such issues. Future work could explore the integration of external knowledge sources and more sophisticated bias detection techniques to further reduce hallucinations and improve the overall accuracy and reliability of LMMs.