Aligning Large Multimodal Models with Factually Augmented RLHF

Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, Kurt Keutzer, Trevor Darrell

Introduction

Large Language Models (LLMs; Brown et al. (2020); Chowdhery et al. (2022); OpenAI (2023)) can delve into the multimodal realm either by further pre-training with image-text pairs (Alayrac et al., ; Awadalla et al., 2023) or by fine-tuning them with specialized vision instruction tuning datasets (Liu et al., 2023a; Zhu et al., 2023), leading to the emergence of powerful Large Multimodal Models (LMMs). Yet, developing LMMs faces challenges, notably the gap between the volume and quality of multimodal data versus text-only datasets. Consider the LLaVA model (Liu et al., 2023a), which is initialized from a pre-trained vision encoder (Radford et al., 2021) and an instruction-tuned language model (Chiang et al., 2023). It is trained on just 150K synthetic image-based dialogues, which is much less in comparison to the text-only models (Flan (Longpre et al., 2023) utilizing over 100M examples spanning 1800 tasks. Such limitations in data can lead to misalignment between the vision and language modalities. Consequently, LMMs may produce hallucinated outputs, which are not accurately anchored to the context provided by images.

To mitigate the challenges posed by the scarcity of high-quality visual instruction tuning data for LMM training, we introduce LLaVA-RLHF, a vision-language model trained for improved multimodal alignment. One of our key contributions is the adaptation of the Reinforcement Learning from Human Feedback (RLHF) (Stiennon et al., 2020; Ouyang et al., 2022; Bai et al., 2022a), a general and scalable alignment paradigm that shows great success for text-based AI agents, to the multimodal alignment for LMMs. By collecting human preferences with an emphasis on detecting hallucinationsWe instructed crowdworkers to prioritize the responses that exhibit better multimodal alignment and minimize hallucinations. That is, if two responses are free of hallucinations, the crowdworkers were asked to choose/create a more helpful one., and utilizes those preferences in reinforcement learning for LMM fine-tuning (Ziegler et al., 2019; Stiennon et al., 2020). This approach can improve the multimodal alignment with a relatively low annotation cost, e.g., collecting 10K human preferences for image-based conversations with $3000. To the best of our knowledge, this approach is the first successful adaptation of RLHF to multimodal alignment.

A potential issue with the current RLHF paradigm is called reward hacking, which means achieving high scores from the reward model does not necessarily lead to improvement in human judgments. To prevent reward hacking, previous work (Bai et al., 2022a; Touvron et al., 2023b) proposed to iteratively collect “fresh” human feedback, which tends to be costly and cannot effectively utilize existing human preference data. In this work, we propose a more data-efficient alternative, i.e., we try to make the reward model capable of leveraging existing human-annotated data and knowledge in larger language models. Firstly, we improve the general capabilities of the reward model by using a better vision encoder with higher resolutions and a larger language model. Secondly, we introduce a novel algorithm named Factually Augmented RLHF (Fact-RLHF), which calibrates the reward signals by augmenting them with additional information such as image captions or ground-truth multi-choice option, as illustrated in Fig. 1.

To improve the general capabilities of LMMs during the Supervised Fine-Tuning (SFT) stage, we further augment the synthetic vision instruction tuning data (Liu et al., 2023a) with existing high-quality human-annotated multi-modal data in the conversation format. Specifically, we convert VQA-v2 (Goyal et al., 2017a) and A-OKVQA (Schwenk et al., 2022) into a multi-round QA task, and Flickr30k (Young et al., 2014b) into a Spotting Captioning task (Chen et al., 2023a), and train the LLaVA-SFT+ models based on the new mixture of data.

Lastly, we look into assessing the multimodal alignment of LMMs in real-world generation scenarios, placing particular emphasis on penalizing any hallucinations. We create a set of varied benchmark questions that cover the 12 main object categories in COCO (Lin et al., 2014) and include 8 different task types, leading to MMHal-Bench. Our evaluation indicates that this benchmark dataset aligns well with human evaluations, especially when scores are adjusted for anti-hallucinations. In our experimental evaluation, as the first LMM trained with RLHF, LLaVA-RLHF delivers impressive outcomes. We observed a notable enhancement on LLaVA-Bench, achieving 94%, an improvement by 60% in MMHal-Bench, and established new performance benchmarks for LLaVA with a 52.4% score on MMBench (Liu et al., 2023b) and an 82.7% F1 on POPE (Li et al., 2023d). We have made our code, model, and data publicly available at https://llava-rlhf.github.io.

Method

Reinforcement Learning from Human Feedback (RLHF) (Ziegler et al., 2019; Stiennon et al., 2020; Ouyang et al., 2022; Bai et al., 2022a) has emerged as a powerful and scalable strategy for aligning Large Language Models (LLMs) with human values. In this work, we use RLHF to align LMMs. The basic pipeline of our multimodal RLHF can be summarized into three stages:

Multimodal Preference Modeling

Reinforcement Learning

where β\beta is the hyper-parameter to control the scale of the KL penalty.

2 Augmenting LLaVA with High-Quality Instruction-Tuning

Recent studies (Zhou et al., 2023; Touvron et al., 2023b) show that high-quality instruction tuning data is essential for aligning Large Language Models (LLMs). We find this becomes even more salient for LMMs. As these models traverse vast textual and visual domains, clear tuning instructions are crucial. Correctly aligned data ensures models produce contextually relevant outputs, effectively bridging language and visual gaps.

For example, LLaVA synthesized 150k visual instruction data using the text-only GPT-4, where an image is represented as the associated captions on bounding boxes to prompt GPT-4. Though careful filtering has been applied to improve the quality, the pipeline can occasionally generate visually misaligned instruction data that can not be easily removed with an automatic filtering script, as highlighted in Table 1.

In this work, we consider enhancing LLaVA (98k conversations, after holding out 60k conversations for preference modeling and RL training) with high-quality instruction-tuning data derived from existing human annotations. Specifically, we curated three categories of visual instruction data: “Yes” or “No” queries from VQA-v2 (83k) (Goyal et al., 2017b), multiple-choice questions from A-OKVQA (16k) (Marino et al., 2019), and grounded captions from Flickr30k (23k) (Young et al., 2014a). Our analysis revealed that this amalgamation of datasets significantly improved LMM capabilities on benchmark tests. Impressively, these results surpassed models (Dai et al., 2023; Li et al., 2023a; Laurençon et al., 2023) trained on datasets an order of magnitude larger than ours, as evidenced by Table 7 and 4. For a comprehensive breakdown of each dataset’s influence, refer to Section 3.5.

3 Hallucination-Aware Human Preference Collection

Inspired by the recent RLHF studies that collect helpfulness and harmlessness preferences (Bai et al., 2022b; Touvron et al., 2023b) separately, in this study, we decide to differentiate between responses that are merely less helpful and those that are inconsistent with the images (often characterized by multimodal hallucinations). To achieve this, we provide crowdworkers with the template illustrated in Table 2 to guide their annotations when comparing two given responses. With our current template design, we aim to prompt crowdworkers to identify potential hallucinations in the model’s responses.

Nonetheless, our training process integrates a single reward model that emphasizes both multimodal alignment and overall helpfulnessWe are considering the development of a distinct Honest reward model, inspired by the approach in Touvron et al. (2023b). This introduces the possibility of constructing a piecewise Honesty-prioritized reward model. We earmark this direction for future exploration.. We collect human preferences on 10k hold-out LLaVA data by re-sampling the last response with our SFT model and a temperature of 0.70.7. The reward model is initialized from the SFT model to obtain the basic multimodal capabilities.

4 Factually Augmented RLHF (Fact-RLHF)

We conduct multimodal RLHF on 50k hold-out LLaVA conversations, with additional 12k multi-choice questions from A-OKVQA and 10k yes/no questions subsampled from VQA-v2. Due to the concerns of existing hallucinations in the synthetic multi-round conversation data of LLaVA, we only use the first question in each conversation for RL training, which avoids the pre-existing hallucinations in the conversational context.

In preliminary multimodal RLHF experiments, we observe that due to the intrinsic multimodal misalignment in the SFT model, the reward model is weak and sometimes cannot effectively detect hallucinations in the RL model’s responses. In the text domain, previous work (Bai et al., 2022a; Touvron et al., 2023b) proposed to iteratively collect “fresh” human feedback. However, this can be quite costly and cannot effectively utilize existing human-annotated data and there is no guarantee that more preference data can significantly improve the discriminative capabilities of the reward model for multimodal problems.

Facutual Augmentation

To augment the capability of the reward model, we propose Factually Augmented RLHF (Fact-RLHF), where the reward model has access to additional ground-truth information such as image captions to calibrate its judgment. In original RLHF (Stiennon et al., 2020; OpenAI, 2022), the reward model needs to judge the quality of the response only based on the user query (i.e., the input image and prompt):

In Factually Augmented RLHF (Fact-RLHF), the reward model has additional information about the textual descriptions of the image:

This prevents the reward model hacked by the policy model when the policy model generates some hallucinations that are clearly not grounded by the image captions. For general questions with COCO images, we concatenate the five COCO captions as the additional factual information, while for A-OKVQA questions, we use the annotated rationals as the factual information.

The factually augmented reward model is trained on the same binary preference data as the vanilla reward model, except that the factual information is provided both during the model fine-tuning and inference.

Symbolic Rewards: Correctness Penalty & Length Penalty

In some of our RL data, certain questions come with a predetermined ground-truth answer. This includes binary choices (e.g., “Yes/No”) in VQA-v2 and multiple-choice options (e.g., “ABCD”) in A-OKVQA. These annotations can also be regarded as additional factual information. Therefore, in the Fact-RLHF algorithm, we further introduce a symbolic reward mechanism that penalizes selections that diverge from these ground-truth options.

Furthermore, we observed that RLHF-trained models often produce more verbose outputs, a phenomenon also noted by Dubois et al. (2023). While these verbose outputs might be favored by users or by automated LLM-based evaluation systems (Sun et al., 2023b; Zheng et al., 2023), they tend to introduce more hallucinations for LMMs. In this work, we follow Sun et al. (2023a) and incorporate the response length, measured in the number of tokens, as an auxiliary penalizing factor.

Experiments

We adopt the same network architecture as LLaVA (Liu et al., 2023a). Our LLM is based on Vicuna (Touvron et al., 2023a; Chiang et al., 2023), and we utilize the pre-trained CLIP visual encoder, ViT-L/14 (Radford et al., 2021). We use grid features both before and after the final Transformer layer. To project image features to the word embedding space, we employ a linear layer. It’s important to note that we leverage the pre-trained checkpoints of the linear projection matrix from LLaVA, concentrating on the end-to-end fine-tuning phase for multi-modal alignment in our study. For LLaVA-SFT+-7b, we use a Vicuna-V1.5-7b LLM and ViT-L/14 with image resolution 256×256256\times 256. For LLaVA-SFT+-13b, we use a Vicuna-V1.5-13b LLM and ViT-L/14 with image resolution 336×336336\times 336.

RL Models: Reward, Policy, and Value

The architecture of the reward model is the same as the base LLaVA model, except that the embedding output of the last token is linearly projected to a scalar value to indicate the reward of the whole response. Following Dubois et al. (2023), we initialize the value model from the reward model. Therefore, when training an LLaVA-7B-based policy model with an LLavA-13B-based reward model, the value model is also of 13B size. To fit all the models (i.e., police, reward, value, original policy) into one GPU, we adopt LoRA (Hu et al., 2021) for all the fine-tuning processes in RLHF. We use Proximal Policy Optimization (PPO; Schulman et al. (2017)) with a KL penalty for the RL training. Without further notice, both LLaVA-RLHF-7b and LLaVA-RLHF-13b are trained with a LLaVA-SFT+-13b initialized reward model. More details can be found in Appendix F.

2 MMHal-Bench Data Collection

To quantify and evaluate the hallucination in LMM responses, we have created a new benchmark MMHal-Bench. There are two major differences between MMHal-Bench and previous VLM benchmarks: 1) Speciality: In contrast to prevalent LMM benchmarks Liu et al. (2023a; b); Li et al. (2023d) that evaluate the response quality in the general sense (e.g., helpfulness, relevance), we focus on determining whether there hallucination exists in the LMM responses. Our evaluation metrics are directly developed on this main criterion. 2) Practicality: Some previous LMM benchmarks Li et al. (2023d); Rohrbach et al. (2018) also examine hallucination, but they have limited the questions to yes/no questions, which we found the results may sometimes disagree with the detailed description generated by LMM. Instead of over-simplifying the questions, we adopt general, realistic, and open-ended questions in our MMHal-Bench, which can better reflect the response quality in practical user-LMM interactions.

In MMHal-Bench, we have meticulously designed 96 image-question pairs, ranging in 8 question categories ×\times 12 object topics. More specifically, we have observed that LMM often make false claims about the image contents when answering some types of questions, and thus design our questions according to these types:

Object attribute: LMMs incorrectly describe the visual attributes of invididual objects, such as color and shape.

Adversarial object: LMMs answers questions involving something that does not exist in the image, instead of pointing out that the referred object cannot be found.

Comparison: LMMs incorrectly compare the attributes of multiple objects.

Counting: LMMs fail to count the number of the named objects.

Spatial relation: LMMs fail to understand the spatial relations between multiple objects in the response.

Environment: LMMs make wrong inference about the environment of the given image.

Holistic description: LMMs make false claims about contents in the given image when giving a comprehensive and detailed description of the whole image.

Others: LMMs fail to recognize the text or icons, or incorrectly reason based on the observed visual information.

We create and filter the questions in an adversarial manner. More specifically, we design the image-question pairs to ensure that the original LLaVA\textsc13Bx336{}_{\textsc{13Bx336}} model hallucinates when answering these questions. While these questions are initially tailored based on LLaVA\textsc13Bx336{}_{\textsc{13Bx336}}’s behavior, we have observed that they also have a broader applicability, causing other LMMs to hallucinate as well.

To avoid data leakage or evaluation on data that LMMs have observed during training, we select images from the validation and test sets of OpenImages (Kuznetsova et al., 2020) and design all brand-new questions. Our image-question pairs cover 12 common object meta-categories from COCO (Lin et al., 2014), including “accessory”, “animal”, “appliance”, “electronic”, “food”, “furniture”, “indoor”, “kitchen”, “outdoor”, “person”, “sports”, and “vehicle”.

When evaluating LMMs on MMHal-Bench, we employ the powerful GPT-4 model (OpenAI, 2023) to analyze and rate the responses. Currently, the publically available GPT-4 API only supports text input, so it cannot judge directly based on the image contents. Therefore, to aid GPT-4’s assessment, we also provide category names of the image content, and a standard human-generated answer in the prompt, in addition to the question and LMM response pair. Consequently, GPT-4 can determine whether hallucination exists in the LMM response by comparing it against the image content and the thorough human-generated answer. When provided with adequate information from MMHal-Bench, GPT-4 can make reasonable decisions aligned with human judgments. For example, when deciding whether hallucination exists in responses from LLaVA\textsc13Bx336{}_{\textsc{13Bx336}} and IDEFICS\textsc80B{}_{\textsc{80B}}, GPT-4 agrees with human judgments in 94% of the cases. Please see the Appendix for the example image-question pairs and GPT-4 prompts we used for MMHal-Bench evaluation.

3 Results

We use LLaVA-Bench (Liu et al., 2023a) and our MMHal-Bench as our main evaluation metrics for their high alignment with human preferences. In addition, we conducted tests on widely-recognized Large Multimodal Model benchmarks. We employed MMBench (Liu et al., 2023b), a multi-modal benchmark offering an objective evaluation framework comprising 2,974 multiple-choice questions spanning 20 ability dimensions. This benchmark utilizes ChatGPT to juxtapose model predictions against desired choices, ensuring an equitable assessment of VLMs across varying instruction-following proficiencies. Furthermore, we incorporated POPE (Li et al., 2023d), a polling-based query technique, to offer an evaluation of Large Multimodal Model object perception tendencies.

By delving into the specific performances for the capability benchmarks (i.e., MMBench and POPE), we observe a notable improvement in capabilities brought by high-quality instruction-tuning data (LLaVA-SFT+) in Tables 4 and 7. LLaVA-SFT+\textsc7B{}_{\textsc{7B}} model exemplifies this with an impressive performance of 52.1% on MMBench and an 82.7% F1 score on POPE, marking an improvement over the original LLaVA by margins of 13.4% and 6.7% respectively. However, it’s worth noting that LLaVA-SFT+ does trail behind models like Kosmos and Shikra. Despite this, LLaVA-SFT+ stands out in terms of sample efficiency, utilizing only 280k fine-tuning data—a 5% fraction of what’s employed by the aforementioned models. Furthermore, this enhancement isn’t confined to just one model size. When scaled up, LLaVA-SFT+\textsc13Bx336{}_{\textsc{13Bx336}} achieves commendable results, attaining 57.5% on MMBench and 82.9% on POPE. Comparatively, the effect of RLHF on the capability benchmarks is more mixed. LLaVA-RLHF shows subtle degradations at the 7b scale, but the 13b LLaVA-RLHF improves over LLaVA-SFT+ by 3% on MMBench. This phenomenon is similar to the Alignment Tax observed in previous work (Bai et al., 2022a). Nonetheless, with our current empirical scaling law of LLaVA-RLHF, we believe RLHF alignment would not damage in general capabilities of LMMs for models of larger scales.

RLHF improves human alignment benchmarks further.

From another angle, even though high-quality instruction data demonstrates large gains in capability assessment, it does not improve much on human-alignment benchmarks including LLaVA-Bench and MMHal-Bench, which is also evident in recent LLM studies (Wang et al., 2023). LLaVA-RLHF show a significant improvement in aligning with human values. It attains scores of 2.05 (7b) and 2.53 (13b) on MMHal-Bench and improves LLaVA-SFT+ by over 10% on LLaVA-Bench. We also presented qualitative examples in Table 1, which shows LLaVA-RLHF produces more reliable and helpful outputs.

4 Ablation Analysis

We conduct ablation studies on LLaVA\textsc7B{}_{\textsc{7B}} and evaluate over the four aforementioned benchmarks.

5 Ablation on High-Quality Instruction-Tuning Data

In Table 5, we evaluate the impact of individual instruction-tuning datasets. For the sake of simplicity, we did not adjust the mixture rate, earmarking that consideration for future research. Our findings indicate that A-OKVQA (Schwenk et al., 2022) contributes significantly to performance enhancements, boosting results by +9.8% on MMBench and a more modest +3.8% on POPE. In contrast, VQA-v2 (Goyal et al., 2017a) is particularly influential on POPE, where it leads to a 6% improvement, while only having a slight impact on MMBench. This differential can possibly be attributed to the overlapping “Yes/No” format in VQA and the multiple-choice structure of A-OKVQA. Flickr30k notably enhances the performance in LLaVA-Bench and MMHal-Bench — a likely consequence of the inherently grounded nature of the task. Furthermore, amalgamating these three datasets results in compounded performance gains across various capability benchmarks.

6 Ablation on Fact-Augmented RLHF

We compare the performance of Fact-Augmented RLHF (Fact-RLHF) with standard RLHF in Table 5. Our findings indicate that while the conventional RLHF exhibits improvement on LLaVA-Bench, it underperforms on MMHal-Bench. This can be attributed to the model’s tendency, during PPO, to manipulate the naive RLHF reward model by producing lengthier responses rather than ones that are less prone to hallucinations. On the other hand, our Fact-RLHF demonstrates enhancements on both LLaVA-Bench and MMHal-Bench. This suggests that Fact-RLHF not only better aligns with human preferences but also effectively minimizes hallucinated outputs.

7 Data Filtering v.s. RLHF

In our preliminary tests, we employed the Fact-RLHF reward model to filter out 70%, 50%, and 30% of LLaVA data. Subsequently, we finetuned an LLaVA model on this filtered data, yielding scores of 81.2, 81.5, and 81.8 on LLaVA-Bench. However, performance on MMHal-Bench , POPE, and MMBench remained largely unchanged. We believe this stagnation can be attributed to two factors: the absence of a negative feedback mechanism preventing the model from identifying hallucinations in its output, and the potential limitations of our Fact-RLHF reward model, especially when compared against the high-capacity oracle models in previous successful studies (Touvron et al., 2023b).

Related Work

Recent success in Large Language Models (LLMs) such as GPTs (Brown et al., 2020; OpenAI, 2023), PaLM (Chowdhery et al., 2022; Anil et al., 2023), BLOOM (Scao et al., 2022; Muennighoff et al., 2022), LLaMA (Touvron et al., 2023a; b), Alpaca (Taori et al., 2023) and Vicuna (Chiang et al., 2023) has spurred significant improvements in multi-modal models. Flamingo (Alayrac et al., ) pioneered integrating LLMs into vision-language pretraining, utilizing gated cross-attention dense blocks to adapt to visual features; its open-source variant is OpenFlamingo (Awadalla et al., 2023) and IDEFICS (Laurençon et al., 2023). PaLI (Chen et al., 2022; 2023b) studies the scaling factor of V&L components across a wide range of tasks. PaLM-E(Driess et al., 2023) further extends LMM to the embodied domain. BLIP-2 (Li et al., 2023c) introduced the Querying Transformer (Q-former) to bridge the gap between image and language encoders, which was further improved by InstructBLIP (Dai et al., 2023). Otter (Li et al., 2023b; a) focuses on enhancing OpenFlamingo’s instruction-following capability. MiniGPT-4 (Zhu et al., 2023) suggests GPT4’s prowess is due to sophisticated LLMs and recommends using a single project layer to align visual and linguistic models. It showcases abilities akin to GPT4 but is computationally efficient. mPLUG-Owl (Ye et al., 2023) offers a new approach: initially aligning visual features and then fine-tuning the language model using LoRA (Hu et al., 2021). Recently, QWen-VL (Bai et al., 2023) scales the pre-training of LMM to 1.4B data and achieves impressive results across benchmarks. Among them, LLaVA (Liu et al., 2023a; Lu et al., 2023) pioneered LMM work by harnessing GPT4 (OpenAI, 2023) for generating vision-language tuning datasets similar to text instruction efforts (Wei et al., 2021; Chung et al., 2022; Longpre et al., 2023; Sanh et al., 2021; Mukherjee et al., 2023; Taori et al., 2023; Köpf et al., 2023). However, due to the syntactic nature of these generated datasets, misalignments between image and text modalities are prevalent. Our research is the first to address this misalignment through RLHF.

Hallucination

Prior to the advent of LLMs, the NLP community primarily defined “hallucination” as the generation of nonsensical content or content that deviates from its source (Ji et al., 2023). The introduction of versatile LLMs has expanded this definition, as outlined by (Zhang et al., 2023) into: 1) Input-conflicting hallucination, which veers away from user-given input, exemplified in machine translation (Lee et al., 2018; Zhou et al., 2020); 2) Context-conflicting hallucination where output contradicts prior LLM-generated information (Shi et al., 2023); and 3) Fact-conflicting hallucination, where content misaligns with established knowledge (Lin et al., 2021). Within the LMM realm, “object hallucination” is well-documented (Rohrbach et al., 2018; MacLeod et al., 2017; Li et al., 2023d; Biten et al., 2022), referring to models producing descriptions or captions including objects that don’t match or are missing from the target image. We expand on this, encompassing any LMM-generated description unfaithful to image aspects, including relations, attributes, environments, and so on. Consequently, we present MMHal-Bench, aiming to holistically pinpoint and measure hallucinations in LMMs.

Discussions & Limitations

Hallucination phenomena are observed in both Large Language Models (LLMs) and Large Multimodal Models (LMMs). The potential reasons are two-fold. Firstly, a salient factor contributing to this issue is the low quality of instruction tuning data for current LMMs, as they are typically synthesized by more powerful LLMs such as GPT-4. We expect our proposed high-quality vision instruction-tuning data and future efforts on manually curating high-quality vision instruction tuning data can alleviate this problem.

Secondly, the adoption of behavior cloning training in instruction-tuned LMMs emerges as another fundamental cause (Schulman, 2023). Since the instruction data labelers lack insight into the LMM’s visual perception of an image, such training inadvertently conditions LMMs to speculate on uncertain content. To circumvent this pitfall, the implementation of reinforcement learning-based training provides a promising avenue, guiding the model to articulate uncertainties more effectively (Lin et al., 2022; Kadavath et al., 2022). Our work demonstrates a pioneering effort in this direction. Figure 3 illustrates the two sources of hallucination in current behavior cloning training of LLMs.

However, while LLaVA-RLHF enhances human alignment, reduces hallucination, and encourages truthfulness and calibration, applying RLHF can inadvertently dampen the performance of small-sized LMMs. Balancing alignment enhancements without compromising the capability of LMM and LLM is still an unresolved challenge. Furthermore, though we’ve demonstrated the effective use of linear projection in LLaVA with top-tier instruction data, determining an optimal mixture and scaling it to bigger models remains intricate. Our research primarily delves into the fine-tuning phase of VLMs, leaving the issues of misalignment in other modalities and during pre-training yet to be explored.

Finally, while MMHal-Bench emphasizes the evaluation of LMMs with an aim to curtail hallucinations, it is noteworthy that short or evasive responses can inadvertently attain high scores on MMHal-Bench. This underlines an intrinsic trade-off between honesty and helpfulness (Bai et al., 2022a). Consequently, for a more comprehensive assessment of alignment with human preferences, we advocate for the evaluation of prospective LMMs using both MMHal-Bench and LLaVA-Bench.

Conclusion

We proposed several strategies to tackle the multimodal misalignment problems, particularly for vision language models (VLMs), which often produce text inconsistent with the associated images. First, we enrich GPT-4 generated vision instruction tuning data from LLaVA with existing human-authored image-text pairs. Next, we adopt the Reinforcement Learning from Human Feedback (RLHF) algorithm from the text domain to bridge vision-language gaps, wherein human evaluators discern and mark the more hallucinated output. We train the VLM to optimize against simulated human preferences. Moreover, we introduce the Factually Augmented RLHF, leveraging additional factual information such as image captions to enhance the reward model, countering reward hacking in RLHF, and boosting model performance. For tangible real-world impact assessment, we have devised MMHal-Bench, an evaluation benchmark targeting the penalization of hallucination. Remarkably, LLaVA-RLHF, being the first VLM trained with RLHF, shows a notable surge in performance across benchmarks. We opensource our code, and data and hope our findings could help the future development of more reliable and human-aligned LLMs and LMMs.

References

Appendix A Source of Multimodal Hallucination

Appendix B Detailed Evaluation Results on MMHal-Bench

We include Table 6 for the full evaluation results on MMHal-Bench.

Appendix C Detailed Evaluation Results on POPE

We include Table 7 for the full evaluation results on POPE.

Appendix D Amazon Mechanical Turk Design for Human Feedback Data Collection

The instruction we gave to the crowdworkers is shown in Table 2. Here, we demonstrate the few-shot examples we provided to the crowdworkers.

Appendix E Example Questions of MMHal-Bench

In this section, we showcase some example questions of MMHal-Bench. As mentioned in the main paper, MMHal-Benchcovers 12 common object categories, and 8 types of questions where LMMs usually incorrectly hallucinate:

Object attribute: LMMs incorrectly describe the visual attributes of invididual objects, such as color and shape. See example Table 11.

Adversarial object: LMMs answers questions involving something that does not exist in the image, instead of pointing out that the referred object cannot be found. See example Table 12.

Comparison: LMMs incorrectly compare the attributes of multiple objects. See example Table 13.

Counting: LMMs fail to count the number of the named objects. See example Table 14.

Spatial relation: LMMs fail to understand the spatial relations between multiple objects in the response. See example Table 15.

Environment: LMMs make wrong inference about the environment of the given image. See example Table 16.

Holistic description: LMMs make false claims about contents in the given image when giving a comprehensive and detailed description of the whole image. See example Table 17.

Others: LMMs fail to recognize the text or icons, or incorrectly reason based on the observed visual information. See example Table 18.

Appendix F Details on implementations and hyperparameters

For LoRA-based fine-tuning during the RLHF stage, we use a low-rank r=64r=64 for both attention modules and feed-forward network modules. We follow Dubois et al. (2023) on the implementation of the PPO algorithm, which is a variant of (Ouyang et al., 2022)https://github.com/openai/lm-human-preferences. Specifically, we normalize the advantage across the entire batch of rollouts obtained for each PPO step and initialize the value model from the reward model.

We used a batch size of 512 for each PPO step. This comprised two epochs of gradient steps, each having 256 rollouts. We applied a peak learning rate of 3×10−53\times 10^{-5} with cosine decay. We clipped the gradient by its Euclidean norm at a limit of 11. Our training spanned 44 complete rounds on our held-out RL data, equaling around 500 PPO steps. For generalized advantage estimation (GAE; Schulman et al. (2015)), both λ\lambda and γ\gamma were set at 1. We opted for a constant KL regularizer coefficient of 0.10.1.

For symbolic rewards, the length penalty is set as the number of response tokens divided by the maximum response length (set to 896896) times the length penalty coefficient. We set the length penalty coefficient to −10.0-10.0 for general questions, −40.0-40.0 for detailed description questions in LLaVA data, and 2.52.5 for complex reasoning questions in LLaVA data. The correctness penalty is set to for incorrect responses (or irrelevant responses), and to 22 for correct responses. A penalty of −8.0-8.0 is also applied to incomplete responses.

Appendix G GPT-4 Examplers and Prompt for MMHal-Bench

We leverage GPT-4 OpenAI (2023) to evaluate the model responses to the image-question pairs in MMHal-Bench. To this end, we first explain the concept of “hallucination” in the context of LMM and list several examples, and request GPT-4 to analyze and rate the response by LMMs. Finally, we instantiate the query by providing the image contents (extracted from OpenImages annotations), question, standard human-generated answer, and the LMM response to evaluate. We use the following template prompt as the input to GPT-4, and extract its output to quantify the quality of each response.