Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models

Ce Zhang, Zifu Wan, Zhehan Kan, Martin Q. Ma, Simon Stepputtis, Deva Ramanan, Russ Salakhutdinov, Louis-Philippe Morency, Katia Sycara, Yaqi Xie

Introduction

Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across various multi-modal tasks, such as image captioning and visual question answering, by extending the capabilities of powerful Large Language Models (LLMs) to incorporate visual inputs (Liu et al., 2023; Li et al., 2023b; Dai et al., 2023; Bai et al., 2023; Ye et al., 2024). Despite their proficiency in interpreting both visual and textual modalities, these models often suffer from hallucinations, where LVLMs erroneously produce responses that are inconsistent with the visual input (Li et al., 2023d; Gunjal et al., 2024; Yin et al., 2023; Wu et al., 2024). This potential for misinformation raises significant concerns, limiting the models’ reliability and restricting their broader deployment in real-world scenarios (Liu et al., 2024b; Bai et al., 2024; Chen et al., 2024b; Zhao et al., 2024).

Recent research has revealed that a major cause of hallucinations in LVLMs is the over-reliance on language priors due to biased training sets, which can override the visual content in response generation (Bai et al., 2024; Liu et al., 2024b; Leng et al., 2024). In response, various strategies have been developed to detect and mitigate these hallucinations by directly introducing additional training (Chen et al., 2024a; Sun et al., 2023; Jiang et al., 2024; Chen et al., 2023; Zhang et al., 2024), demonstrating promising results in reducing over-reliance. However, the need for additional data and costly training processes hinders their deployment in downstream tasks. More recently, a new paradigm of methods has emerged to tackle the hallucination problem in LVLMs by intervening in the decoding process (Huang et al., 2024; Deng et al., 2024; Kim et al., 2024). Among these, recent training-free contrastive decoding-based methods (Li et al., 2023c) have proven effective in mitigating undesired hallucinations by contrasting token predictions derived from original visual input with bias-inducing counterparts, such as no/distorted visual input (Favero et al., 2024; Leng et al., 2024), disturbed instructions (Wang et al., 2024), or premature layers (Chuang et al., 2024).

While these contrastive decoding-based methods effectively mitigate hallucinations arising from language priors, we recognize that hallucinations can also originate beyond language bias, stemming from visual deficiencies in LVLMs (Tong et al., 2024). For instance, in counting hallucinations, language does not imply any count information; instead, miscounts largely arise from visual recognition errors of LVLMs, as complex scenes include numerous, similar objects at ambiguous positions which may confuse the LVLMs, leading to incorrect visual understanding and, consequently, hallucinated answers. Therefore, we argue that current contrastive decoding-based methods may struggle to generalize effectively across different types of hallucinations.

In this work, we explore the potential of leveraging powerful text-to-image generative models (e.g., Stable Diffusion (Rombach et al., 2022; Podell et al., 2024)) to mitigate various types of hallucinations in LVLMs. Our work is based on a simple yet intuitive hypothesis: Given a visual input and a textual prompt to an LVLM, if the generated response conditioned on the original image is accurate and non-hallucinatory, a text-to-image generative model should be capable of reversing this process to produce a similar image from that response. Alternatively, if there is a discrepancy between the original image and the one generated from the response, this difference can serve as valuable self-feedback, guiding the decoding process to correct potential hallucinations in the initial response. To verify this hypothesis, we conduct an empirical study (in Section 3.2), demonstrating that generative models can provide valuable self-feedback for mitigating hallucinations at both the response and token levels.

Building on this insight, we introduce self-correcting Decoding with Generative Feedback (DeGF), a novel training-free decoding algorithm that effectively incorporates feedback from text-to-image generative models to recursively enhance the accuracy of LVLM responses. Specifically, for each instance, we generate a new image based on the initial response, which serves as an auxiliary visual reference to assess and verify the accuracy of the initial output. We propose self-correcting decoding that either enhances or contrasts predictions from the original and this reference based on the auxiliary visual reference, confirming or revising the initial LVLM response based on the degree of divergence between the two predictions. By integrating this additional visual reference and generative feedback, LVLMs can gain enhanced visual insights and verify the initial response to ensure accurate visual details in the text outputs. In Figure 1, we demonstrate that incorporating generative feedback in our approach can reduce various types of hallucinations, including object existence, visual appearance, counting, etc. To the best of our knowledge, we are the first work to explore the use of text-to-image generative feedback as a self-correcting mechanism for mitigating hallucinations in LVLMs.

The effectiveness of DeGF is evaluated on LLaVA-1.5, InstructBLIP, and Qwen-VL across six benchmarks: POPE (Li et al., 2023d), CHAIR (Rohrbach et al., 2018), MME-Hallucination (Fu et al., 2023), MMBench (Liu et al., 2024d), MMVP (Tong et al., 2024), and LLaVA-Bench. Extensive experimental results validate the effectiveness of our DeGF in mitigating various types of hallucinations in LVLMs. Qualitative case studies and GPT-4V-aided evaluation on LLaVA-Bench further demonstrate that our approach enhances both the accuracy and detailedness of the LVLM responses.

The contributions of this paper are summarized as follows:

We investigate the potential of text-to-image generative models in mitigating hallucinations in LVLMs and demonstrate that text-to-image generative models can provide valuable self-feedback for mitigating hallucinations at both the response and token levels.

We propose self-correcting Decoding with Generative Feedback (DeGF), a novel training-free decoding algorithm for LVLMs that recursively enhances the accuracy of responses by integrating feedback from text-to-image generative models with complementary/contrastive decoding.

Extensive experimental evaluations across six benchmarks demonstrate that our DeGF consistently outperforms state-of-the-art approaches in effectively mitigating hallucinations in LVLMs.

Related Work

Hallucination in LVLMs. With advances of autoregressive LLMs (Touvron et al., 2023; Chowdhery et al., 2023; Chiang et al., 2023), researchers have extended these powerful models to process visual inputs, leading to the development of LVLMs (Liu et al., 2023; Dai et al., 2023; Bai et al., 2023; Ye et al., 2024). These models typically train a modality alignment module to project visual tokens into the textual embedding space of the LLM, demonstrating impressive performance in various multi-modal tasks such as visual question answering and image captioning (Liu et al., 2024b; Bai et al., 2024). However, LVLMs are prone to hallucinations, where contradictions arise between the visual content and the generated textual response (Li et al., 2023d; Liu et al., 2024b; Bai et al., 2024).

To mitigate hallucinations in LVLMs, early works have introduced various approaches, including reinforcement learning from human feedback (RLHF) (Gunjal et al., 2024; Sun et al., 2023), applying auxiliary supervision (Jiang et al., 2024; Chen et al., 2023), incorporating negative (Liu et al., 2024a) or noisy data (Yue et al., 2024), and training post-hoc revisors for correction (Zhou et al., 2024; Yin et al., 2023). Despite promising results, these methods often lack practicality due to their reliance on additional data and costly training processes. To address this, another line of work focuses on training-free methods that can be seamlessly integrated into existing LVLMs. Such methods encompass contrastive decoding (Leng et al., 2024; Favero et al., 2024) and guided decoding with auxiliary information (Chen et al., 2024d; Deng et al., 2024; Woo et al., 2024). In this work, we present a novel training-free approach that recursively enhances the accuracy of the LVLM response by incorporating text-to-image generative feedback. To the best of our knowledge, we are the first work to effectively utilize feedback from text-to-image generative models to mitigate hallucinations in LVLMs.

Text-to-Image Synthesis. Text-to-image synthesis aims to create realistic images from textual descriptions (Zhu et al., 2019; Ge et al., 2023). In recent years, significant progress has been achieved in this area, largely due to the advent of deep generative models (Zhan et al., 2023; Goodfellow et al., 2014). These advances include Generative Adversarial Networks (GAN) (Sauer et al., 2023; Kang et al., 2023), autoregressive models (Chang et al., 2023; Yu et al., 2022), and diffusion models (Ho et al., 2020; Karras et al., 2022; Nichol et al., 2022; Saharia et al., 2022; Rombach et al., 2022). Among these, diffusion-based methods have been particularly distinguished due to their ability to generate high-quality, detailed images with fine-grained control over the synthesis process (Yang et al., 2023; Croitoru et al., 2023). Pre-trained on large-scale text-image datasets such as LAION (Schuhmann et al., 2022), diffusion-based methods have demonstrated strong vision-language alignment, making them valuable for downstream tasks such as classification (Li et al., 2023a) and semantic segmentation (Amit et al., 2021; Wolleb et al., 2022).

More recently, Jiao et al. (2024) incorporate text-to-image generative models to enhance fine-grained image recognition in LVLMs by introducing the Img-Diff dataset, which generates pairs of similar images using Stable Diffusion XL (Podell et al., 2024). Their results demonstrate that fine-tuning LVLMs with this additional data leads to improved performance on several VQA tasks. In contrast, in this work, we directly leverage a pre-trained diffusion model to provide valuable self-feedback for refining the generated responses of LVLMs in the decoding process, dynamically improving the accuracy and consistency of the model’s response without modifying the underlying LVLMs.

Method

In this work, we present DeGF, a novel training-free algorithm that recursively improves the accuracy of LVLM responses using text-to-image generative feedback, as illustrated in Figure 2.

We consider an LVLM parameterized by θ\theta, which processes an input image vv and a textual query x\mathbf{x}, aiming to autoregressively generate a fluent sequence of textual responses y\mathbf{y}. The visual input vv is first processed by a vision encoder and then projected into visual tokens within the textual input space using a vision-language alignment module (e.g., Q-Former (Li et al., 2023b) or linear projection (Liu et al., 2023)). These visual tokens, along with the textual query tokens, are then fed into the language encoder for conditioned autoregressive generation. We denote the autoregressive generation process as

where yty_{t} represents the token at time step tt, y<t≜[y0,…,yt−1]\mathbf{y}_{<t}\triangleq[y_{0},\dots,y_{t-1}] denotes the sequence of tokens generated before time step tt, and fθf_{\theta} is the logit distribution (unnormalized log-probabilities) produced by the LVLM over a vocabulary of textual tokens V\mathcal{V}. At each step t∈[0,…,T]t\in[0,\dots,T], the response token yty_{t} is sampled from the probability distribution pθ(yt∣v,x,y<t)p_{\theta}(y_{t}|v,\mathbf{x},\mathbf{y}_{<t}), and this generative process continues iteratively until the response sequence y≜[y0,…,yT]\mathbf{y}\triangleq[y_{0},\dots,y_{T}] is complete.

2 Visual Reference Generation

In our method, we incorporate generative feedback from diffusion models to guide the decoding process. Specifically, given a visual input vv and a textual query x\mathbf{x}, we first prompt the LVLMs to generate an initial response τ\bm{\tau}, which includes relevant descriptions of the visual input with potential hallucinations. Subsequently, we leverage a pre-trained diffusion model G\mathcal{G} to generate a new image v′v^{\prime} based on the initial response:

Here, xTx_{T} denotes a sample from the standard Gaussian distribution, which serves as the initial noisy input to the diffusion model. Starting from this pure noise image xTx_{T}, the diffusion model G\mathcal{G} iteratively applies TT steps of the denoising process to obtain xT,xT−1,…,x0x_{T},x_{T-1},\dots,x_{0}, where the final output x0x_{0} corresponds to the final generated image v′v^{\prime}. Through this diffusion process, the generative model visualizes the initial response, providing a visual reference that helps mitigate potential hallucinations and produce a more accurate and consistent output.

Effectiveness of Text-to-Image Generative Models in Reflecting Hallucinations. We validate the effectiveness of generative models in reflecting hallucinations through an empirical study, as shown in Figure 3.For Figure 3, we evaluate 1,000 CHAIR samples (Left) and 3,000 POPE samples (Right). The experimental results demonstrate that text-to-image generative models can provide valuable self-feedback for mitigating hallucinations at both the response and token levels.

We conduct the following two experiments: (1) We generate an image v′v^{\prime} using diffusion model based on the initial caption provided by LLaVA-1.5 and compute the CLIP image similarities between the original image vv and the generated image v′v^{\prime} using OpenCLIP (Cherti et al., 2023) ViT-H/14 backbone. Following prior work, we use the CHAIR (Rohrbach et al., 2018) benchmark, a rule-based metric on MS-COCO (Lin et al., 2014) for evaluating object hallucination from generated captions. We report the average per-instance metric CHAIRI\text{CHAIR}_{I} within each bin of CLIP similarity, which evaluates the object hallucination rates in the entire initial response. As shown in Figure 3 (Left), a clear negative correlation between hallucination rates and CLIP similarities is observed (with a correlation coefficient of ρ=−0.63\rho=-0.63). This indicates that lower similarity between the original image and generated image corresponds to higher rates of hallucinations at the response level. (2) Similarly, we generate an image v′v^{\prime} based on the initial response given by LLaVA-1.5 for each instance on the POPE (Li et al., 2023d) benchmark. In Figure 3 (Right), we present the density plot of Jensen-Shannon (JS) divergence between the predicted probabilities for both images, i.e., pθ(yt∣v,x,y<t)p_{\theta}(y_{t}|v,\mathbf{x},\mathbf{y}_{<t}) and pθ(yt∣v′,x,y<t)p_{\theta}(y_{t}|v^{\prime},\mathbf{x},\mathbf{y}_{<t}), for hallucinatory and non-hallucinatory tokens.Note that POPE benchmark contains yes-or-no questions about object existence. In this experiment, we evaluate only the first response token (i.e., yes or no) to determine the presence of hallucinations. The results show that the density of JS divergence follows a long-tail distribution, with hallucinatory tokens exhibiting significantly longer tails and higher JS divergence. This shows JS divergence between probabilities derived from the original and the generated image corresponds well to hallucinations at the token level. These observations provide insights into the effectiveness of generative models in reflecting hallucinations, and motivate us to incorporate the generative feedback during the decoding process.

3 Self-Correcting Decoding with Generative Feedback

In this section, we focus on effectively utilizing generative feedback during the decoding process to mitigate potential hallucinations. Specifically, we propose a self-correcting decoding approach that leverages generative feedback to confirm or revise the initial response by selectively enhancing or contrasting the logits for each generated token based on the measured divergence between the two predicted probability distributions.

Specifically, to predict a specific token yty_{t}, we utilize LVLMs to generate two output distributions, each conditioned on either the original image vv or the synthesized visual reference v′v^{\prime}, expressed as:

We define and compute the following distance metric based on Jensen-Shannon (JS) divergence at each timestep tt to quantify the discrepancy between two next-token probability distributions:

We consider two scenarios based on the token-level generative feedback: (1) If the two predictions are aligned and both images agree on a specific token prediction, we confirm the original prediction as correct, and the auxiliary prediction from the generated image can be combined with the original prediction for enhancement (complementary decoding (Woo et al., 2024)). (2) Conversely, if there is a significant discrepancy between the predictions, indicating that the original prediction is likely hallucinatory, we revise the original response by using the generated visual input as a contrasting reference to refine the initial next-token prediction (contrastive decoding (Leng et al., 2024)). To implement this, we introduce a distance threshold γ\gamma and develop two corresponding decoding approaches as follows:

where α1\alpha_{1} and α2\alpha_{2} are hyperparameters that control the influence of the generated visual reference in the final prediction. Note that setting α1=0\alpha_{1}=0 or α2=0\alpha_{2}=0 degrades this process to regular decoding. The final generated token yty_{t} is sampled from the multinomial distribution with probabilities pθ(yt)p_{\theta}(y_{t}).

Experiments

In this section, we evaluate the effectiveness of our method in mitigating hallucinations in LVLMs across a range of benchmarking scenarios, comparing it with existing state-of-the-art approaches.

Evaluated LVLMs. We evaluate the effectiveness of our method on three state-of-the-art open-source LVLMs: LLaVA-1.5 (Liu et al., 2024c), InstructBLIP (Dai et al., 2023), and Qwen-VL (Bai et al., 2023). Both LLaVA-1.5 and InstructBLIP utilize Vicuna-7B (Chiang et al., 2023) as the language encoder, which is instruction-tuned from LLaMA (Touvron et al., 2023). In contrast, Qwen-VL (Bai et al., 2023) is based on the Qwen 7B backbone. Specifically, we implement our approach using weights of the Qwen-VL-Chat model.

Benchmarks. We conduct extensive experiments on six benchmarks: (1) POPE (Li et al., 2023d) is a widely used benchmark for assessing object hallucinations in LVLMs, which tests the models with yes-or-no questions regarding the presence of specific objects, such as, “Is there a {object} in the image?” (2) CHAIR (Rohrbach et al., 2018) evaluates object hallucinations in open-ended captioning tasks. It prompts the LVLMs to describe specific images selected from a random sample of 500 images from the MSCOCO validation set; (3) MME-Hallucination (Fu et al., 2023) is a comprehensive benchmark for LVLMs consisting of four subsets: existence and count for object-level hallucinations, and position and color for attribute-level hallucinations; (4) MMBench (Liu et al., 2024d) serves as a comprehensive benchmark designed to assess the multi-modal understanding capabilities of LVLMs across 20 dimensions; (5) MMVP (Tong et al., 2024) collects CLIP-blind pairs and evaluates the fine-grained visual recognition capabilities of LVLMs. It consists of 150 image pairs, each accompanied by a binary-option question; (6) LLaVA-Bench provides 24 images featuring complex scenes, memes, paintings, and sketches, along with 60 challenging questions.

Baselines. As a simple baseline, we include results from regular decoding, where the next token is sampled directly from the post-softmax probability distribution. Additionally, we compare the performance of our method with three state-of-the-art decoding approaches: VCD (Leng et al., 2024), M3ID (Favero et al., 2024), and RITUAL (Woo et al., 2024). For evaluations on the CHAIR (Rohrbach et al., 2018) and MME-Hallucination (Fu et al., 2023) benchmark, we further include comparisons with Woodpecker (Chen et al., 2024d), HALC (Chen et al., 2024d), DoLa (Chuang et al., 2024) and OPERA (Huang et al., 2024). We report the performance of these baselines based on our re-implementation using their released code bases.

Implementation Details. In our experiments, we adhere to the default query format for the input data used in both LLaVA-1.5 (Liu et al., 2024c) and InstructBLIP (Dai et al., 2023). Additionally, we set α1=3\alpha_{1}=3, α2=1\alpha_{2}=1, and γ=0.1\gamma=0.1 by default in our decoding process. We follow VCD (Leng et al., 2024) to implement adaptive plausibility constraints (Li et al., 2023c), where we set β=0.1\beta=0.1 in open-ended CHAIR benchmark and β=0.25\beta=0.25 for other tasks. To ensure the reliability of our results, we conduct MME experiments three times with different initialization seeds and report the mean accuracy along with the standard deviation. All experiments are conducted on a single 48GB NVIDIA RTX 6000 Ada GPU. More implementation details are provided in Section B of the Appendix.

2 Results and Discussions

Results on POPE. In Table 1, we compare the performance of our method against other baselines on the POPE benchmark under three different negative sampling settings, across three datasets. As shown in the table, our method consistently outperforms other decoding methods on both LVLMs, achieving state-of-the-art accuracies across all 18 settings, with improvements of up to 5.24% in accuracy, 6.33% in precision, and 2.79% in F1 score compared to the second-best approach. This suggests that incorporating a generative reference enables the LVLMs to perceive more fine-grained visual details, thereby effectively addressing object hallucinations. Moreover, while most decoding methods tend to be overconfident in their responses, the self-correcting decoding mechanism in our method makes it more conservative in responding Yes, as evidenced by significantly higher precision across all settings. This highlights its enhanced performance in filtering out false positives and suppressing misinformation.

Another notable finding is that our method shows significantly improved performance in the popular and adversarial settings, which are more challenging than the random setting. In the popular and adversarial settings, non-existent negative objects frequently appear and co-occur with other objects (Li et al., 2023d), making them more susceptible to hallucination by LVLMs, as evidenced by the varying degrees of performance degradation across all baselines. However, our method exhibits a lower performance drop compared to other baselines, demonstrating its effectiveness in addressing hallucinations arising from object co-occurrence.

Results on CHAIR. We also compare the performance of our methods and other state-of-the-art methods in the open-ended captioning task and report the CHAIR scores, recall, and the average length of responses in Table 2, Table C1, and Table C2. The results, evaluated across two different LVLMs, consistently demonstrate performance improvements achieved by our method over the compared approaches. Specifically, our method outperforms the second-best approach by 3.0% and 2.6% on the CHAIRS metric, while also enhancing the detailedness of generated responses compared to regular decoding, as indicated by the higher recall and increased response length. These results demonstrate that by incorporating generative feedback into the decoding process of LVLMs, our method effectively mitigates object hallucinations in open-ended captioning tasks.

Results on MME-Hallucination and MMBench. Beyond object hallucinations, we further compare the performance of our method with other approaches using the more comprehensive MME-Hallucination benchmark, which includes both object-level and attribute-level hallucinations. The results in Table 3 and Table C3 demonstrate that our method significantly outperforms the compared methods, with substantial margins in the total score metric (e.g., +18.19 on LLaVA-1.5 and +21.11 on InstructBLIP) and consistently superior performance across various evaluation settings, achieving the best results in 6 out of 8 settings. Moreover, our method shows notable improvements on the attribute-level color subset, which is particularly challenging as it requires models to accurately capture subtle attribute information. This further illustrates the effectiveness of our approach in addressing a wide range of hallucinations, both at the object existence level and in finer-grained attribute recognition. Additionally, our proposed DeGF enhances the general multi-modal understanding capabilities of LVLMs, as evidenced by its superior performance on the MMBench benchmark.

Results on MMVP. We conduct experiments on the MMVP benchmark to assess the fine-grained visual recognition capabilities of LVLMs. As shown in Figure 4, applying our self-correcting decoding approach to LLaVA-1.5 significantly improves performance from 22.67% to 27.33%. Our approach also demonstrates notable advantages over other hallucination mitigation baselines, further showcasing its superiority in handling nuanced visual recognition tasks. These results suggest that our approach significantly enhances the model’s capacity to discern and correctly interpret fine-grained distinctions between images with similar appearances but different contents. By integrating generative feedback, our approach effectively reduces misinterpretations and improves the precision of visual recognition tasks, contributing to more reliable and accurate performance in complex scenarios.

Results on LLaVA-Bench. In Figure 5, we present a case study on LLaVA-Bench comparing our method’s response with the response generated by regular decoding using the LLaVA-1.5 model. Specifically, regular decoding often leads to hallucinated or inaccurate content, such as describing “the island below the mountain”. Besides, the response generated by regular decoding tends to focus on elements like the “cloudy sky” and “cohesive and captivating island landscape” without providing specific information about the central features of the image. In contrast, our response is more detailed, mentioning the volcano, the road, the surrounding greenery, and the inhabited areas, which gives a clearer understanding of the image’s content. The GPT-4V-aided evaluation shown in Table 4 further confirms that our method enhances both the accuracy and detailedness of the generated response, outperforming other hallucination mitigation approaches such as VCD and M3ID. Due to the page limit, please refer to Section D of the Appendix for more case studies.

3 Ablation Studies

Analysis of Distance Threshold γ\gamma. In Section 3.3, we introduce a distance threshold γ\gamma to determine the appropriate decoding algorithm for each generated token. Table 4.2 presents an analysis of our method’s performance with various values of γ\gamma across three benchmarks. For simplicity, we report the performance on the MS-COCO dataset with random setting for all POPE results in the ablation studies. Notably, when γ\gamma is set to either 0 or 1—corresponding to the exclusive use of contrastive or complementary decoding for all tokens—the performance exhibits a significant decline, by 0.6% and 1.1% in POPE accuracy, respectively. Moreover, our default setting of γ=0.1\gamma=0.1 achieves the optimal performance in 3 out of 4 evaluated metrics. Additional sensitivity analyses for other hyperparameters are provided in Section C of the Appendix.

Effects of Different Generative Models. Table 4.2 presents the performance of various variants of our method that incorporate different generative models (i.e., different versions of Stable Diffusion) while using the same LLaVA-1.5 backbone. The results indicate that the effectiveness of our DeGF is robust to the choice of generative models, as performance remains largely unaffected by the specific model used, and all variants demonstrate consistent improvements over the original regular decoding approach. Although utilizing SD-XL-v1.0 (Podell et al., 2024) yields slightly better performance, we opt for SD-v1.5 as the default due to its faster image generation speed (3.8 s/image vs. 11.3 s/image).

4 Efficiency Comparison

In Table 7, we compare the efficiency of our approach with other methods on the CHAIR benchmark using the LLaVA-1.5 model, with the maximum token length set to 128. Our approach involves two queries and incorporates a text-to-image generative model to mitigate hallucinations, resulting in a 4.04×\times increase in latency and a 1.21×\times increase in GPU memory usage. Specifically, our method consists of three stages: initial response generation, image generation, and response self-correction, which take an average of 3.4 seconds, 3.8 seconds, and 6.6 seconds per instance, respectively. Compared to other approaches, while our method is slower than regular decoding and contrastive decoding-based methods, it demonstrates efficiency advantages over OPERA and HALC. Note that our approach also achieves the lowest hallucination rates among all compared methods. In Appendix C.9, we discuss several strategies to accelerate our approach, such as limiting the length of the initial response and reducing the number of inference steps in the diffusion process.

Conclusion

In this work, we present self-correcting Decoding with Generative Feedback (DeGF), a novel training-free approach that leverages feedback from text-to-image generative models to recursively improve the accuracy of generated responses. Specifically, we generate a new image based on the initial response given by LVLMs, which serves as a visual reference and provides token-level feedback for mitigating hallucinations. Building on this, we propose a corresponding self-correcting decoding algorithm that measures the discrepancy between next-token predictions conditioned on the original and generated images, selecting either contrastive or complementary decoding to reduce the likelihood of hallucinatory responses. Extensive experimental results across six benchmarks demonstrate that our DeGF consistently outperforms state-of-the-art methods in mitigating hallucinations in LVLMs.

Acknowledgements

This work has been funded in part by the Army Research Laboratory (ARL) award W911NF-23-2-0007, DARPA award FA8750-23-2-1015, and ONR award N00014-23-1-2840. MM and LPM are partially supported by Meta and National Institutes of Health awards R01MH125740, R01MH132225, and R21MH130767. RS is supported in part by the ONR grant N00014-23-1-2368.

Ethics Statement

Our work focuses on developing methods to mitigate hallucinations in large vision-language models, aiming to enhance the reliability of AI-generated content. Our research does not involve human subjects, sensitive data, or any practices that pose privacy or security concerns. Additionally, we discuss the broader ethical and societal implications of this work in Section A of the Appendix.

Reproducibility Statement

The large vision-language models utilized in our experiments, such as LLaVA and InstructBLIP, are open-source and publicly available. We have detailed our experimental setup, including hyperparameter configurations, prompts, and other key design choices, in Section 4 of the main paper and Section B of the Appendix to ensure reproducibility. Code is publicly available at https://github.com/zhangce01/DeGF.

References

Appendix A Limitations and Broader Impacts

Limitations. Although our method effectively mitigates hallucinations in LVLMs, it relies on pre-trained text-to-image generative models, which introduces additional computational complexity. The process of generating images also adds time, potentially slowing down LVLM response generation and making it less suitable for real-time applications. However, our method is training-free, reducing the overhead typically associated with fine-tuning large models, and offering broader applicability across various tasks. Moreover, the use of generative feedback improves the model’s ability to verify and correct responses, particularly in complex scenarios. Thus, while the computational trade-offs may limit real-time performance, our method excels in settings where accuracy and reliability are prioritized over speed. We also hope that advances in efficient diffusion-based models will improve the feasibility of our approach in real-world applications in the future.

Broader Impacts. In this work, our goal is to develop more reliable large vision-language models (LVLMs) by incorporating feedback from generative models. By using this feedback mechanism, we aim to address a critical issue faced by current multi-modal models: hallucinations, where models produce responses that are inconsistent with the visual input. Hallucinations not only degrade model performance but also pose risks in real-world applications by generating inaccurate or misleading information. Our approach leverages the strengths of generative models to detect and mitigate these hallucinations, improving the overall accuracy and reliability of LVLMs. In doing so, we contribute to enhancing trustworthiness and reducing the spread of misinformation in systems that rely on multi-modal AI, making them safer and more effective for a wide range of applications.

Appendix B More Experimental Details

We conduct extensive experiments on the following benchmarks:

POPE (Li et al., 2023d) is a widely used benchmark for assessing object hallucinations in LVLMs. It tests the models with yes-or-no questions regarding the presence of specific objects, such as, “Is there a {object} in the image?” The benchmark draws data from three existing datasets: MSCOCO (Lin et al., 2014), A-OKVQA (Schwenk et al., 2022), and GQA (Hudson & Manning, 2019), and comprises three distinct subsets—random, popular, and adversarial—based on how the negative samples are generated. For each dataset setting, the benchmark provides 6 questions per image, resulting in 3,000 test instances. We evaluate the performance of different methods using four metrics: accuracy, precision, recall, and F1 score.

CHAIR (Rohrbach et al., 2018) evaluates object hallucinations in open-ended captioning tasks. It prompts the LVLMs to describe specific images selected from a random sample of 500 images from the MSCOCO validation set and assesses performance based on two metrics:

Additionally, we assess the recall and the average length of the generated responses.

MME-Hallucination (Fu et al., 2023) is a comprehensive benchmark for LVLMs consisting of four subsets: existence and count for object-level hallucinations, and position and color for attribute-level hallucinations. Each subset includes 30 images and 60 questions, with two questions per image. Similar to POPE (Li et al., 2023d), these questions are structured as yes-or-no queries, and performance is assessed based on binary accuracy. Following the official implementation, the reported score is calculated by combining accuracy and accuracy+, where accuracy is based on individual questions, and accuracy+ is based on images where both questions are answered correctly.

MMBench (Liu et al., 2024d) is a comprehensive evaluation benchmark designed to assess the multimodal understanding and reasoning capabilities of AI models. It focuses on tasks requiring the integration of visual and textual information, testing a model’s ability to handle diverse, real-world scenarios. In particular, MMBench employs a hierarchical ability taxonomy, designating Perception and Reasoning as Level-1 (L-1) abilities. It further refines the taxonomy by incorporating more detailed ability dimensions, organizing them into six Level-2 (L-2) and twenty Level-3 (L-3) dimensions.

MMVP (Tong et al., 2024) collects CLIP-blind pairs and evaluates the fine-grained visual recognition capabilities of LVLMs. It consists of 150 image pairs, each accompanied by a binary-option question. Each image is queried independently, and for a given pair, the LVLM’s response is considered correct only if both associated questions are answered accurately.

LLaVA-Benchhttps://huggingface.co/datasets/liuhaotian/llava-bench-in-the-wild. provides 24 images featuring complex scenes, memes, paintings, and sketches, along with 60 challenging questions. We select examples from this dataset to provide qualitative comparisons between the responses generated by different decoding methods. We also follow Yin et al. (2023) to evaluate the accuracy and detailedness of generated responses of different methods using the advanced LVLM, GPT-4Vhttps://openai.com/index/gpt-4v-system-card..

B.2 More Implementation Details

In our experiments, we adhere to the default query format for the input data used in both LLaVA-1.5 (Liu et al., 2023) and InstructBLIP (Dai et al., 2023). Additionally, we set α1=3\alpha_{1}=3, α2=1\alpha_{2}=1, and γ=0.1\gamma=0.1 by default in our decoding process. We follow VCD (Leng et al., 2024) to implement adaptive plausibility constraints (Li et al., 2023c):

Here, V\mathcal{V} is the whole vocabulary of LVLM, and hyperparameter β∈\beta\in controls the truncation of the next token distribution. A larger β\beta indicates more aggressive truncation, keeping only the high-probability tokens. In our implementation, we set the logits for yt∉V(y<t)y_{t}\notin\mathcal{V}(y_{<t}) to −∞-\infty. By default, we set β=0.1\beta=0.1 in the open-ended CHAIR benchmark and β=0.25\beta=0.25 for other tasks. All experiments are conducted on a single 48GB NVIDIA RTX 6000 Ada GPU.

Recall that in our method, we use a text-to-image generative model to reverse the image-to-text response generation process by producing a new image from the initial response. To ensure the new image is both high-quality and relevant, we aim to generate specific descriptions for the given visual content. Thus, we slightly modify the initial query prompt for each evaluated benchmark:

POPE (Li et al., 2023d), MME-Hallucination (Fu et al., 2023), and MMVP (Tong et al., 2024). In POPE, MME-Hallucination, and MMVP benchmarks, models are tested with yes-or-no/binary selection questions, such as, “Is there a {object} in the image?” To obtain more detailed explanations and descriptions of the original image, we modify the prompt by adding, “Briefly describe relevant details.” This encourages the model to provide not only a yes-or-no answer but also additional visual information.

CHAIR (Rohrbach et al., 2018). For the CHAIR benchmark, we retain the original prompt, “Please describe this image in detail.” as it effectively prompts the model to provide comprehensive visual details from the original image.

Note that for the second query, where both the original and generated images are used as input, we apply the original prompt to ensure a fair comparison.

B.3 Details of Other Baselines

In this work, we mainly compare the performance of our DeGF with three state-of-the-art approaches: VCD (Leng et al., 2024), M3ID (Favero et al., 2024), and RITUAL (Woo et al., 2024). The method and implementation details for these approaches are provided below:

VCD (Leng et al., 2024) contrasts output distributions derived from original and distorted visual inputs. Specifically, given a textual query x{x} and a visual input v{v}, the model generates two distinct output distributions: one conditioned on the original v{v} and the other on the distorted visual input v′{v^{\prime}}, which is derived by applying pre-defined distortions (i.e., Gaussian noise mask) to v{v}. Then, a new contrastive probability distribution is computed by:

In our implementation, we follow the default setting in VCD (Leng et al., 2024) and set α=1\alpha=1 for reproduction. To generate v′v^{\prime}, we use a total of 500 noise steps.

M3ID (Favero et al., 2024) contrasts output distributions derived from original visual inputs and pure text inputs without visual information. The final probability distribution is

Similarly, we follow their recommended best practice and set the hyperparameter λ\lambda, which balances the conditioned model and unconditioned model, to 0.020.02.

RITUAL (Woo et al., 2024) applies common image transformations (e.g., crop, flip, color jitter, etc.) to the original visual input vv, This results in a transformed version of the visual input, v(T)v^{(T)}. Then, RITUAL utilizes both the original and transformed images to generate the response and this dual-input approach significantly reduces the likelihood of hallucinatory outputs. The probability distribution is calculated as follows:

Here, κ\kappa is a balancing hyperparameter, adjusting the contribution of the transformed input relative to the original. We follow their official implementation to set κ=3\kappa=3 in default.

B.4 Dataset and Code Licensing

Datasets. We list the known license information for the datasets below: POPE (Li et al., 2023d) and MMVP (Tong et al., 2024) benchmarks are licensed under MIT License. CHAIR (Rohrbach et al., 2018) is made available under the BSD 2-Clause License. LLaVA-Bench is available under Apache-2.0 License. MME-Hallucination (Fu et al., 2023) benchmark dataset is collected by Xiamen University for academic research only.

Code. In this work, we also use some code implementations from the existing codebase: LLaVA (Liu et al., 2023) and VCD (Leng et al., 2024) are licensed under the Apache-2.0 License. InstructBLIP (Dai et al., 2023) is under BSD-3-Clause License. RITUAL (Woo et al., 2024) is licensed under MIT License.

Appendix C More Experimental Results and Analysis

In Table C1 and Table C2, we present performance comparisons on the CHAIR benchmark with maximum number of tokens set to 128 and 256. The results indicate that our approach also achieves competitive performance across two LVLMs in mitigating hallucinations during long-sequence generation scenarios.

C.2 Full Results on MME-Hallucination

In Table C3, we present the full results on the MME-Hallucination benchmark across three LVLMs.

C.3 Full Results on MMBench

In Table C4, we present the overall performance on the MMBench benchmark, as well as the detailed performance across six Level-2 abilities: Logical Reasoning (LR), Attribute Reasoning (AR), Relation Reasoning (RR), Fine-grained Perception - Single Instance (FP-S), Fine-grained Perception - Cross Instance (FP-C), and Coarse Perception (CP). We follow VCD (Leng et al., 2024) to conduct experiments on the MMBench-dev set.

C.4 Results on MM-Vet

In Table C5 and Table C6, we present the overall performance on the MM-Vet (Yu et al., 2024) benchmark with random sampling decoding and greedy decoding strategies, respectively. We use LLaVA-1.5 as the LVLM backbone. From the results, we observed that our method consistently outperforms others on the MMVet benchmark. Notably, it significantly excels in the OCR, spatial awareness, and math subsets.

C.5 Results on POPE Using Greedy Decoding

In Table C7, we present performance comparisons on the POPE benchmark with random sampling from the MS-COCO dataset. The experiment is conducted using the LLaVA-1.5 backbone.

C.6 Effects of α1\alpha_{1} and α2\alpha_{2} in Self-Correcting Decoding

In Section 3, we present two decoding approaches: complementary decoding and contrastive decoding. We also introduce two balancing hyperparameters, α1\alpha_{1} and α2\alpha_{2}, which control the relative influence of the original and generated images in next-token prediction. In Table C8 and Table C9, we analyze the effect of varying α1\alpha_{1} or α2\alpha_{2} while keeping all other hyperparameters at their default settings. The results indicate that our default choice of α1=3\alpha_{1}=3 and α2=1\alpha_{2}=1 consistently yields the best performance across two benchmarks. Moreover, compared to setting these hyperparameters to 0, which effectively reduces complementary/contrastive decoding to standard decoding, the performance improvements demonstrate that our proposed decoding approaches significantly contribute to the overall effectiveness of DeGF in mitigating hallucinations in LVLMs.

C.7 Effect of β\beta in Adaptive Plausibility Constraint

We further conduct an ablation study on β\beta introduced in Equation (7), where we vary β\beta from 0 to 0.5 while keeping all other hyperparameters fixed. The results in Table C10 show that setting β=0\beta=0, which imposes no constraint, results in suboptimal performance across both benchmarks. Additionally, in the POPE benchmark, where LVLMs handle yes-or-no questions, a more aggressive truncation with β=0.25\beta=0.25 yields the best performance. In contrast, for the open-ended CHAIR benchmark, a lower value of β=0.1\beta=0.1 leads to the best results.

C.8 Scaling Up the LVLMs

We further extend our evaluation to larger-scale 13B variants of the LLaVA-1.5 model to assess the scalability of our approach. Table C11 compares our experimental results with other state-of-the-art approaches across all three subsets of the POPE benchmark using the 13B-sized LLaVA-1.5 model. We observe that scaling up the LLaVA-1.5 model does not alleviate the hallucination issues, as evidenced by the comparable performance of both the 7B and 13B models. Using the 13B-sized model, our DeGF consistently achieves improved performance across all subsets compared to other approaches, demonstrating its general effectiveness and scalability.

C.9 Speeding Up Our Approach

In this section, we propose two strategies to accelerate our approach: limiting the length of the initial response and reducing the number of inference steps in the diffusion process.

Reducing Diffusion Inference Steps. By default, we set the number of diffusion inference steps to 50 to ensure high-quality image generation. To improve the response generation speed, we can reduce the number of diffusion steps. In Table C.9, we report the performance on the CHAIR benchmark after reducing the diffusion inference steps in the model. By reducing the diffusion inference steps from 50 to 10, the average latency decreases by 2.85 seconds per instance, while the performance on CHAIR remains robust. This demonstrates that reducing the inference steps of the diffusion model is an effective way to speed up our approach.

Restricting Length of Initial Response. Our method involves two queries to the LVLM for self-correcting decoding. To enhance efficiency, we can limit the length of the initial response. In Table C.9, we present the efficiency and CHAIR performance results after decreasing the maximum token limit for the initial response. We can see that reducing the maximum number of tokens in the initial response from 128 to 96 decreases the latency by 0.72 seconds per instance while maintaining competitive performance. However, further reductions result in performance degradation, as a shorter initial response fails to adequately cover the entire scene, limiting its ability to generate an image that effectively reflects and mitigates hallucinations.

Note that these two strategies are not conflicting; instead, they are complementary. Setting the diffusion steps to 10 and limiting the maximum number of tokens in the initial response to 96 further reduces the inference latency to 10.21 seconds per instance while maintaining robust performance.

C.10 Quantitative Assessment of Generated Image Quality

Our approach incorporates a text-to-image generation model to mitigate hallucinations. We evaluate the quality of the generated images on all 4 subsets on the MME benchmark using CLIPScore (Hessel et al., 2021). Specifically, we utilize the CLIP backbone with ViT-B/32 backbone for our evaluation. We list the results in Table C14. As we can see from the table, our text-to-image generative model (specifically, SD-v1.5) achieves an average CLIPScore of over 30 across all subsets. For comparison, the advanced DALL-E 3 model achieves a score of 32.0, while DALL-E 2 achieves 31.4.These results are sourced from the technical report on DALL-E 3, available at: https://cdn.openai.com/papers/dall-e-3.pdf. These results highlight the capability of our model to generate high-quality images that closely align with the initial response.

Appendix D More Case Studies

Following VCD (Leng et al., 2024), we use GPT-4V to evaluate responses in open-ended generation scenarios, scoring them based on accuracy and detailedness. Leveraging the strong human-like capabilities of GPT-4V, it can detect incorrect colors, positions, and relationships, providing a comprehensive evaluation of the responses. Specifically, we apply the prompt provided in Table D15 to instruct GPT-4V to rate the two responses on a scale of 1 to 10 for both accuracy and detailedness:

Accuracy measures the consistency between the responses/descriptions generated by the LVLMs and the given image. A lower score is assigned if GPT-4V detects any inconsistencies in the content of the responses.

Detailedness evaluates the depth and specificity of the responses provided by the LVLMs. A higher score is awarded if the response includes comprehensive descriptions, captures fine-grained details of the image, and provides well-elaborated explanations. Conversely, a lower score is given if the response is vague or lacks sufficient detail.

D.2 More Qualitative Results

In Figure D1 and Figure D2, we provide additional case studies on LLaVA-Bench to qualitatively demonstrate the effectiveness of our methods in mitigating hallucinations. We also included GPT-4V evaluations of accuracy and detailedness scores for each instance.

In Figure D3-D6, we provide qualitative evaluations of the images generated by the generative model, including both success and failure cases, across all four subsets of the MME benchmark to better understand the effectiveness of the generative models. Our results show that, despite occasional failure cases, the generative model consistently produces high-quality and realistic images that accurately visualize the initial response, providing effective self-feedback.

Appendix E Future Work

In future work, we aim to extend the evaluation of our method to a broader range of LVLMs, such as Mini-GPT4 (Zhu et al., 2024) and mPLUG-Owl2 (Ye et al., 2024), as well as additional benchmarks, including R-Bench (Wu et al., 2024), which focuses on relation hallucination, and ROPE (Chen et al., 2024c), which addresses multiple-object hallucination. This expanded evaluation will allow us to more comprehensively assess the generalizability and effectiveness of our approach across diverse models and tasks.

Furthermore, we plan to investigate integrating generative feedback directly into the instruction tuning phase. This integration has the potential to eliminate the computational overhead associated with applying our method during inference, thereby significantly improving efficiency without compromising performance. By pursuing these directions, we hope to further enhance the practical applicability and scalability of our approach.