Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander Hauptmann, Yonatan Bisk, Yiming Yang

Introduction

This paper addresses the challenge of aligning LMMs, particularly in tasks that involve video instruction following. Despite recent advancements in reinforcement learning (RL) (Ouyang et al., 2022; Bai et al., 2022; Lee et al., 2023; Sun et al., 2023b) and DPO (Rafailov et al., 2024; Chen et al., 2024b; Hosseini et al., 2024), which have been effective in guiding LLMs towards generating more honest, helpful, and harmless content, their effectiveness in multimodal contexts remains limited. The critical obstacle lies in developing a robust reward system capable of distinguishing preferred responses from less preferred ones, especially when such responses are generated based on video inputs. The challenge is further complicated by the presence of hallucinations in generated content, stemming from the scarcity of alignment data across different modalities (Liu et al., 2023b; Sun et al., 2023a).

While human preference data is valuable, it is challenging to scale due to its cost and labor-intensive nature, as highlighted by the LLaVA-RLHF (Sun et al., 2023a) paper, which collected 10k human-evaluated instances at a considerable cost of $3000. Existing approaches for distlling preferences, such as those for image data using GPT-4V (Li et al., 2023d), encounter scalability issues, especially for video inputs that require analyzing multiple frames. While Ahn et al. (2024) leverage a supervised finetuning (SFT) model for self-evaluation, the efficacy of the SFT model remains uncertain, particularly in accurately assessing the factuality of responses in relation to their corresponding videos.

To tackle the aforementioned challenges, we introduce a cost-effective reward mechanism aimed at reliably evaluating the quality of responses generated by video (LLMs), serving as a basis for further preference optimization. We propose the use of detailed video captions as a proxy for video content, enabling a language model analyze video content and assess the accuracy of an LMM’s response to a related question and determine the presence of hallucinations. The language model provides natural language feedback as a chain-of-thought step, and generates a numerical score for reward, facilitating a cost-effective feedback system.

However, high-quality video captions are essential for this process. To mitigate the shortage of high-quality video captions, we have developed a comprehensive video caption dataset, ShareGPTVideo, using a novel prompting technique with the GPT-4V model, comprising 900k captions that encompass a wide range of video content, including temporal dynamics, world knowledge, object attributes, and spatial relationships. With this video caption dataset available, we verify that our reward mechanism, which utilizes video captions as a proxy, is well-aligned with evaluations derived from the more powerful, albeit costlier, GPT-4V model-generated rewards. Employing this reward mechanism as the basis for DPO algorithm, we train LLaVA-Hound-DPO that achieves an 8.1% accuracy improvement over the SFT counterpart. This marks a significant advancement in video LMM alignment and represents the first successful application of a DPO method in this domain.

Our contributions are outlined as follows:

We develop a large-scale, detailed video caption dataset, covering a wide array of content. This dataset serves as a foundational resource for LMM model training and research, facilitating advancements in video understanding tasks.

We introduce a cost-effective method for evaluating video instruction-following tasks, serving as enhanced evaluation of model performance.

We demonstrate the effective application of DPO to improve model performance by leveraging the language model feedback as reward, which substantially improves the alignment of video LMM, establishing a new benchmark for SOTA performance in video QA tasks.

Related Work

LMMs (Liu et al., 2023b; a; Bai et al., 2023; Chen et al., 2023; Li et al., 2023a) have enabled instruction following across modalities by utilizing LLM as backbones. In the context of video understanding, LLMs have been adapted to process video content (Lin et al., 2023a; Zhang et al., 2023a; Maaz et al., 2023; Li et al., 2023b; Luo et al., 2023; Liu et al., 2023c; Jin et al., 2024; Ahn et al., 2024). Our work adots Video-LLaVA backbone, focusing on model enhancement through preference modeling with the DPO technique.

2 Video-text Datasets

Existing video-text datasets typically provide brief sentences or mere keywords as captions, as indicated by Bain et al. (2021); Wang et al. (2023); Yu et al. (2019); Jang et al. (2017); Xu et al. (2016). Shvetsova et al. (2023) uses automatic speech recognition to extract textual content from videos, but it encounters alignment issues when the audio does not match or is absent from the visual content. Video-ChatGPT (Li et al., 2023b) employs human effort to create high-quality video instructions, albeit limited to the ActivityNet domain with only 100k instruction pairs. Our work leverages the GPT-4V model with specifically crafted prompts to produce detailed video captions as community resource for LMM training.

3 Preference Modeling for LMMs

Preference modeling techniques are employed to enhance the utility of LMMs while mitigating the issue of hallucination. Sun et al. (2023a) leveraged Reinforcement Learning with Human Feedback (RLHF) and incorporated caption information into the reward model to improve the assessment of factuality. More recently, Ahn et al. (2024) used RL on AI feedback to improve video LMM performance. For the image understanding, Li et al. (2023d); Gunjal et al. (2023) introduced the application of DPO on the distilled rewards from GPT-4V on a group of model outputs, while Zhao et al. (2023) created preference data using ChatGPT to generate positive and negative pairs informed by detailed image descriptions. Our contribution extends DPO to the video LMM alignment, with the use of detailed captions as factual evidence for reward modeling.

Method

As shown in fig. 1, our methodology enhances video LMM alignment through DPO method using rewards from a language model. We elaborate on constructing a video caption dataset in section 3.1. Subsequently, in section 3.2, we discuss the generation of video instruction data and the fine-tuning process of our model. Lastly, section 3.3 details the incorporation of generated captions as a feedback mechanism for DPO method to refine our model’s factual alignment in video instruction-following tasks.

The selection of dataset includes videos from three sources: the WebVid and VIDAL datasets, which are general domain videos sourced from YouTube with 400k and 450k sampled videos respectively, and the ActivityNet dataset, which adds 50k videos focusing on human activities. The three datasets together result in a comprehensive collection of 900k videos. To accommodate the requirement that GPT-4V only takes images as input, we preprocess videos by uniformly extracting ten frames per video content. These frames are then concatenated into a sequence to serve as a proxy for the video. This sequence is the input into GPT-4V to generate a coherent caption for the represented video based on the frame sequence. The prompt adheres to guidelines covering temporal dynamics, world knowledge, object attributes, spatial relationships, aesthetic assessments, etc., with the goal of comprehensively understanding the video contents.

2 SFT with Generated Video Instruction Data from Detailed Caption

To generate video instruction-following data for SFT, we adopt a similar methodology outlined in Video-ChatGPT (Li et al., 2023b). Specifically, we first randomly sample 20k, 30k, 30k captions in our dataset from ActivityNet, WebVid and VIDAL respective and then employ ChatGPT to generate three question-answer pairs given each detailed video caption, resulting in a total of 240k instruction data for finetuning. This approach ensures that the instructional data remains factually consistent with the content of the detailed captions. The specific prompting strategy used for this instruction generation process is detailed in fig. 13.

3 DPO with Feedback from Language Model as Reward

Acquiring high-quality preference data is both costly and labor-intensive. Although GPT-4V is an effective model for reward distillation, its high cost, slow performance, and limited accessibility hinder scalability, especially for video inputs with multiple frames. We propose a cost-efficient method to generate reward data for DPO using detailed video captions as supporting evidence, as shown in fig. 2.

Initially, we randomly select a subset of 20k instruction pairs from the dataset described in section 3.2. The SFT model uses these sampled questions and their corresponding videos to generate six responses per input pair at a temperature of 1.01.0. This procedure results in 120k question-answer pairs, which will be evaluated. Subsequently, we employ ChatGPT to process inputs including a question, the ground truth answer, the model’s prediction, and a detailed description serving as supportive evidence, with the prompt in fig. 15. This generates an output that includes a natural language explanation as chain-of-thought step, followed by a numerical reward score on a scale from 11 to 55, indicating the level of factual alignment and overall quality.

For each video and question pair, we randomly select an answer with a score ≥\geq 3 as positive example, and an answer with a score below 33 as negative. Cases where all responses are uniformly scored above or below 33 are excluded from the dataset. After the selection process, approximately 17k training instances are compiled for DPO training. Formally, the dataset is denoted as DDPO={(V,x,yw,yl)}\mathcal{D}_{DPO}=\{(\mathcal{V},x,y_{w},y_{l})\}, where V\mathcal{V} is the video, xx is the question, ywy_{w} and yly_{l} are the positive and negative responses. The DPO objective is defined as below:

where πθ\pi_{\theta} is the policy model to be optimized and πref \pi_{\text{ref }} is the base reference model, both models are initialized with SFT weights. σ\sigma is the logistic function and β\beta is set to 0.10.1.

Our approach to reward assignment leverages detailed captions as a proxy for video frames, offering both cost-effectiveness and efficiency. This method incurs costs of less than 20,underapricingmodelof20, under a pricing model of1.5 per million tokens. In comparison, previous methods of preference data collection, such as in Sun et al. (2023a), required an expenditure of 3,000togather10khumanpreferencedatapoints.Additionally,themethodproposedbyLietal.(2023d),whichemploysGPT−4Vforrewarddatalabeling,incursasignificantlyhighercost—3,000 to gather 10k human preference data points. Additionally, the method proposed by Li et al. (2023d), which employs GPT-4V for reward data labeling, incurs a significantly higher cost—30 per million tokens—and demonstrates considerably slower inference speeds.

Assessment of Evaluator with GPT-4V Caption as Evidence

To assess the effectiveness of our proposed reward assignment method, which utilizes detailed captions as a proxy of actual video frames, we conducted a comparative analysis with the GPT-4V, used as a video QA evaluator. The latter reward system employs GPT-4V evaluation directly taking in video frames, a question, and the model prediction as inputs, with detailed prompt in fig. 16. Both reward systems follow the same set of guidelines for scoring reward.

To compare the two methods, we sample 200200 videos from each of the WebVid, VIDAL, and ActivityNet datasets, each associated with one question and two model predictions from our SFT model, with one preferred and one dispreferred by ChatGPT. This results in 1,2001,200 examples, for which we used GPT-4V (with the ”gpt-4-vision-preview” version) version to assign scores. Filtering through the Azure API backend resulted in 196196, 151151, and 143143 videos from each dataset, respectively, having both answers evaluated. The average scores of all examples from ChatGPT and GPT-4V evaluations were 2.92.9 and 3.53.5 respectively, indicating a tendency of GPT-4V to yield slightly positive evaluations. The Pearson Correlation Coefficient (PCC) of 0.470.47 (p<0.01p<0.01) suggests a moderate positive correlation. In fig. 3 (left), the distribution of the difference between ChatGPT and GPT-4V scores reveals that majority (>75%>75\%) of ChatGPT scores fall within one standard deviation (σ=1.31\sigma=1.31) of GPT-4V scores. Additionally, in fig. 3 (right), the agreement on preference between ChatGPT and GPT-4V, excluding ties, exceeded 70%70\%. These findings cautiously support our benchmark’s applicability in video QA evaluation. Further refinements for better alignment—such as incorporating Likert scales Zhou et al. (2023) or GPT-4 evaluation—are areas for future research.

Experimental Results

We adopt Video-LLaVA (Lin et al., 2023a) as the backbone of our video LMM, but our dataset and method can be applied to any other architectures as well. Specifically, Video-LLaVA employs LanguageBind (Zhu et al., 2023) encoder for image and video frame inputs, a MLP projector with 2 fully connected layers to map visual embeddings into text space, and Vicuna Chiang et al. (2023) as large language model. During training, we first initialize the projection MLP layer with the same Video-LLaVA MLP weight. Then we follow the training stages below:

Caption Pre-training Stage (LLaVA-Hound-PT): At pretraining stage, we use captioning data including 650k image caption data from ALLaVA (Chen et al., 2024a) and our distilled 900k video caption. We freeze the LanguageBind visual encoder and fine-tune the MLP projector and LLM, with learning rate 2e-5 and batch size 128.

SFT Stage (LLaVA-Hound-SFT): We use instructional data from both image and video domain to fine-tune the model for instruction-following ability. Our SFT model use 600k image instruction data from ALLaVA and our generated 240k video instruction data, with learning rate 5e-6 and batch size 128.

DPO training Stage (LLaVA-Hound-DPO): We use the 17k preference data introduced in section 3.3 for DPO training. Following Ivison et al. (2023), we train our policy model for 33 epochs with learning rate 5e-7, and a batch size of 128, resulting in roughly 420 training steps. All the experiments are performed on 8 A100 gpus.

2 Existing Benchmark Evaluation

We evaluate model performance on three benchmark datasets: MSVD-QA Chen & Dolan (2011), MSRVTT-QA Xu et al. (2016), and TGIF-QA Jang et al. (2017), using ChatGPT with version gpt-3.5-turbo-0611 to assess model predictions. The evaluation prompts follow Maaz et al. (2023). In our experiment, we found that different ChatGPT versions have high impact on absolute score of metric, but the overall ranking of models is relatively stable. We select gpt-3.5-turbo-0613 due to its closeness to the reported score in Video-LLaVA paper. Further details on the selection rationale and evaluation pitfalls are discussed in Appendix A.

Baseline Selection

Our selection criteria include video LMM models that have demonstrated SOTA performance, specifically including Video-LLaVA, which is also our choice of architecture. We consider other contemporaneous SOTA models with similar reported performance levels to Video-LLaVA, yet have not been directly compared in prior studies. A key consideration in our selection is the availability of models with accessible code and checkpoints, which is crucial for ensuring reproducibility of our findings. To this end, we replicate models including Video-ChatGPT (Maaz et al., 2023), LLaMA-VID (Li et al., 2023e) (7B and 13B), Chat-UniVi (Jin et al., 2023), and Video-LLaVA Lin et al. (2023b). We adopt the results from additional baselines including FrozenBiLM (Yang et al., 2022), VideoChat (Li et al., 2023b) and VideoLLaMA (Zhang et al., 2023a), sourced from their original publication.

Results

In table 1, our analysis shows that within the SFT models, LLaMA-VID-7B and Video-LLaVA exhibit comparable performance, with LLaMA-VID-13B performing the best. Our LLaVA-Hound-SFT model achieves comparable performance to LLaMA-VID-13B. Incorporating preference modeling, LLaVA-Hound-DPO achieves an average accuracy of 70.75%70.75\%, surpassing LLaVA-Hound-SFT, which has an average accuracy of 62.65%62.65\%, by 8.1%8.1\%. Furthermore, LLaVA-Hound-DPO, enhanced by DPO, exhibits superior accuracy compared to VLM-RLAIF’s performance achieved through reinforcement learning.

Error Analysis

Figure 4 illustrates two examples. In the left example, LLaVA-Hound-SFT provides an accurate description of the video’s first half but introduces a hallucination with the phrase “I’m not scared of space,” absent in the video content. LLaVA-Hound-DPO yields a more accurate inference. In the right example, both LLaVA-Hound-SFT and Video-LLaVA models produce incorrect inferences, whereas LLaVA-Hound-DPO successfully correctly identifies the subject in the video. More critically, these examples unveil two significant issues within the current benchmark: (1) the auto-generated questions from existing benchmark may be grammatically incorrect or even nonsensical, and (2) the answers are limited to a single word, which is insufficient for evaluating LMMs with long-form text generation. Such constraints in the ground truth answers hinder the evaluation of crucial aspects like helpfulness and hallucination detection.

3 Proposed Benchmark Evaluation with GPT-4V Caption as Supporting Evidence

As a solution to the above limitations in existing benchmark evaluation, we propose a new set of test questions for same videos in the benchmark datasets with generated QA from detailed captions, illustrated in appendix D. Applying the our reward system in section 4, we report the score from ChatGPT, and a score value ≥3\geq 3 will be considered correct for accuracy calculation. This new long-form QA evaluation potentially support diverse aspects in responses including relevance, accuracy, clarity and completeness in prompt 16.

Table 2 and table 3 shows the in-domain and out-of-domain evaluation. We use ”gpt-3.5-turbo-0301” for evaluation as it is the same version for constructing DPO dataset. The model performance is more distinguishable from our evaluation with Video-LLaVA performing the best among the other baseline models.

Video LMM without Video Instruction: in table 2 is baseline trained with only image instruction fine-tuned on LLaVA-Hound-PT, which achieves an average accuracy of 65.97%65.97\%, comparable to the LLaVA-Hound-SFT model’s 66.06%66.06\% in in-domain QA scenarios. However, its performance significantly drops in out-of-domain QA contexts (49.32%49.32\% vs. 56.50%56.50\%), suggesting that Video QA training could potentially enhance generalization capabilities.

Quality of Generated SFT: substitutes our generated video QA with the Video-ChatGPT dataset for Video-LLaVA fine-tuning. A comparison between the findings of and reveals a marginal performance disparity of 0.2%0.2\% in average accuracy, indicating that the quality of our generated QA closely parallels that of the existing video QA datasets. Given the similar quality in SFT data, the large gain of over can be reasonably concluded from large-scale pre-training on video captions.

Unfreeze MLP: The comparison between and reveals a significant decrease in performance when the MLP is unfrozen during DPO training. Despite this drop, however, the performance remains superior to that of the SFT baseline.

Smaller Learning Rate: The comparison between and reveals that using a smaller learning rate of 3e-7 (vs. 5e-7) results in a decreasing of model performance. This highlights the future improvements by finding better hyperparameters.

Self-Play vs. DPO: Chen et al. (2024b) introduced a self-play methodology for DPO training, which designates ground truth answers as preferred and model-generated responses as dispreferred. When comparing the results of with those in , a notable decrease in accuracy by 3%3\% from the SFT model is observed, suggesting that self-play may be less effective for video LMM alignment, and introducing reward model is helpful.

DPO Accuracy vs. Training Epochs. The left of fig. 5 depicts the generalization performance of the model on out-of-domain video QA tasks with respect to the number of training epochs. We observe a consistent enhancement in model performance among datasets during the initial 0 to 2 epochs, with peak performance materializing at around 2.5 epochs, which corresponds to 350 training steps.

DPO as Ranker vs. Generator. Following Hosseini et al. (2024), we compare the performance of employing the DPO model as a ranker for candidate answers produced by the SFT model, operating at a temperature setting of 1.0. As depicted on the right in fig. 5, we illustrate the test accuracy progression through the selection of the best among NN candidates by the DPO ranker. Initial observations indicate that the SFT model, when set to a temperature of 1.0, demonstrates a reduced accuracy (43.3%) compared to that achieved through greedy decoding (57.8%). A steady enhancement in performance is noted as the number of candidates increases, plateauing at an accuracy of approximately 62% with 64 candidates. This performance, however, falls short when compared with the direct application of the DPO model for answer generation, which yields an accuracy of 68.29%. This difference suggests the stronger generalization of DPO model in answer generation, despite it is trained on a reward classification loss. The contradictory results to Hosseini et al. (2024) may be due to the difference of tasks, i.e. Math vs. Video QA. Refer to appendix E for more results.

4 Analysis on Video Captioning Ability from Pre-training

In Figure 6, we present the video captioning ability of models across various datasets, with a total of 900k distilled data instances. GPT-4V is employed for self-evaluation (fig. 14), serving as the upper-bound performance, while the Video-LLaVA serves for comparative analysis, establishing a baseline. Notably, Video-LLaVA is trained on 54k video QA data instances. However, our first checkpoint, utilizing only 10% of the data, is trained on 90k high-quality caption data instances, likely accounting for the observed performance disparity in the video captioning task. Our results demonstrate that incorporating more distilled data contributes to improved model performance across both in-domain and out-of-domain datasets. Despite these improvements, a performance discrepancy with the GPT-4V model remains. Further, we evaluate the generalization potential in specific data subsets, as shown in fig. 7 in the Appendix. These subsets reveal varying degrees of generalization challenges for different types of dataset. For example, the WebVid subset, which concentrates on relatively static scenes, necessitates less data for effective training compared to the VIDAL subset, which is marked by dynamic scene transitions and a diversity of video themes.

Conclusion

In this study, we propose an cost-effective reward system that utilizes detailed captions as proxies for video content. Our findings demonstrate that the reward scores is well-aligned with the evaluation metrics of GPT-4V, and the incorporation of this reward mechanism enhances DPO training, resulting in SOTA performance on video QA tasks.

Reproducibility Statement

The ensure reproducibility of our work, we plan to release the following items:

Distilled video captions with corresponding frames.

The model weights including the pre-trained, SFT, and DPO models.

Code for training and testing using existing and our proposed benchmark.

References

Appendix A Effect of ChatGPT Version on Official Benchmark Evaluation

In Table 4, we show impact of using different ChatGPT versions on metric scores within zero-shot video question answering benchmarks. Our analysis reveals significant variations in the absolute scores across ChatGPT versions, but based on the average accuracy metric, the relative ranking of models under the same ChatGPT version shows consistency.

This comparison underscores a critical issue: many prior studies neglect to specify the ChatGPT version used, potentially leading to inaccurate conclusions during evaluation. We advocate for the explicit designation of the ChatGPT version in future evaluations. Analysis from Table 4 indicates that the version gpt-3.5-turbo-0613 aligns most closely with the performance of the Video-LLaVA (Lin et al., 2023a) model, serving as the benchmark for model performance comparison in our study.

Appendix B Evaluation of Captioning Ability from pre-training

Appendix C Distilled Caption Demonstration

Appendix D Video QA Dataset Demonstration

Appendix E Additional DPO Results

Appendix F Prompts for GPT-4V and ChatGPT Queries