InstructVideo: Instructing Video Diffusion Models with Human Feedback
Hangjie Yuan, Shiwei Zhang, Xiang Wang, Yujie Wei, Tao Feng, Yining Pan, Yingya Zhang, Ziwei Liu, Samuel Albanie, Dong Ni
Introduction
The emergence of diffusion models has significantly boosted generation quality across a wide range of media content . This generation paradigm has shown promise for video generation , despite the challenges of working with high-dimensional data. While diffusion models are one factor driving progress, the scaling of training datasets has also played a key role . However, despite recent progress, the visual quality of generated videos still leaves room for improvement . A significant contributing factor to this issue is the varying quality of web-scale data employed during pre-training , which can yield models capable of generating content that is visually unappealing, toxic and misaligned with the prompt.
While aligning model outputs with human preferences has proven highly effective for control , text generation and image generation , it remains a notion unexplored in video diffusion models. The most widely-adopted methods for aligning models with human preferences include off-line reinforcement learning (RL) and direct reward back-propagation . Typically, this entails training a reward model on manually annotated datasets that is then subsequently used to fine-tune the pre-trained generative model.
Two major challenges arise when seeking to align video generation models with human preferences: 1) The optimization process for optimizing human preferences is computationally demanding, often requiring video generation from textual inputs. While video generation pre-training using DDPM requires only a single-step inference for every iteration, reward optimization requires a 50-step DDIM inference . 2) The curation of a large annotated dataset to capture human preferences of videos is labor-intensive, while the computation- and memory-intensive demands of utilizing ViT-H or ViT-L -based computational alternatives to evaluate the entire video are high.
To surmount these mentioned challenges, we propose InstructVideo, a model that efficiently instructs text-to-video diffusion models to follow human feedback, as illustrated in Fig. 1. Regarding the first challenge of the demanding reward fine-tuning process caused by generating through the full DDIM sampling chain, we recast the problem of reward fine-tuning as an editing procedure. This reformulation requires only partial inference of the DDIM sampling chain, thereby reducing computational demands while improving fine-tuning efficiency. Drawing inspiration from established editing workflows in diffusion models , where primary visual content is initially corrupted with noise and then reshaped by a target prompt, our method focuses on refining coarse and structural videos into more detailed and nuanced outputs. This contrasts with previous methods that generate results directly from text. Such blurry and structural videos, serving as the starting point for reward fine-tuning, are procured by a simple diffusion process with negligible cost. During generation, the optimized model retains the capability to produce videos directly from textual inputs. In conjunction with back-propagation truncation of the sampling chain, we make reward fine-tuning on text-to-video diffusion models computationally attainable and effective.
Regarding the second challenge (the lack of a reward model tailored for video generation), we postulate that the visual excellence of a video is tied to both the quality of its individual frames and the fluidity of motion across consecutive frames. To this end, we resort to off-the-shelf image reward models, e.g., HPSv2 , to ascertain frame quality. Drawing inspiration from temporal segment networks , we propose Segmental Video Reward (SegVR), which strategically evaluates video quality based on a subset of sparely sampled frames. By providing sparse reward signals, SegVR offers dual benefits: it not only ameliorates computational burden but also mitigates temporal modeling collapse. On the other hand, although LoRA is adopted by default to retain the capability to generate temporally smooth videos, SegVR still leads to videos with visual artifacts, such as structure twitching and color jittering. To mitigate this, we propose Temporally Attenuated Reward (TAR), which operates under the hypothesis that central frames should be assigned paramount importance, with emphasis tapering off towards peripheral frames. This strategic allocation of importance across frames ensures a more stable and visually coherent video generation process.
As part of our pioneering effort to align video diffusion models with human preferences, we conduct extensive experiments to assess the practicality and efficacy of integrating image reward models within InstructVideo. Our findings reveal that InstructVideo markedly enhances the visual quality of generated videos without sacrificing the model’s generalization capabilities, setting a new precedent for future research in video generation.
Related Work
Video generation via diffusion models. Early efforts at video generation focused on GANs and VAEs . However, due to the complexity of jointly modeling spatio-temporal dynamics, generating videos from texts remains an unresolved challenge. Recent methods for video generation aim to mitigate this by utilizing the de facto generation method, i.e., diffusion models , for generating videos with diversity and fidelity and scaling up the pre-training data or model architecture . VDM represents a pioneering work that extended image diffusion models to video generation. Owing to the computation-intensive nature of diffusion models, follow-up research sought to reduce overhead by leveraging the latent space , e.g., ModelScopeT2V , Video LDM , MagicVideo and SimDA , etc.. To enable more controllable generation, further efforts introduce spatio-temporal conditions , e.g., VideoComposer , Gen-1 , DragNUWA , etc.. However, generating videos that adhere to human preferences remains a challenge.
Human preference model. Understanding human preference in visual content generation remains challenging . Some pioneering works target solving this problem by annotating a dataset with human preferences, e.g., Pick-a-pic , ImageReward and HPD . Language-image models, e.g. CLIP and BLIP , are then fine-tuned on the resulting annotated data. As such, the fine-tuned models represent a data-driven approach to modelling human preferences. However, the annotation process for capturing human preferences is highly labor- and cost-intensive. Thus, in this paper, we adopt off-the-shelf image preference models to improve the quality of generated videos.
Learning from human feedback. Learning from human feedback was first studied in the context of reinforcement learning and agent alignment and later in large language models , enabling them to generate helpful, honest and harmless textual outputs. This goal of learning from human feedback is also desirable in visual content generation. In image generation and editing, a key objective is to align generated images with the prompt , preventing surprising and toxic results . Lee et al. utilizes reward-weighted regression on a manually collected dataset, aiming to mitigate misalignment with respect to factors such as count, color and background. DDPO and DPOK propose to use policy gradients on a multi-step diffusion model , demonstrating improved rewards of aesthetic quality, image-text alignment, compressibility, etc.. DRaFT and AlignProp achieve feedback optimization by back-propagating the gradients of a differentiable reward function through the sampling procedure via gradient checkpointing . However, learning from human feedback for video diffusion models remains under-explored owing to its prohibitive cost. InstructVideo aims to fill this gap, providing a solution for more efficient reward fine-tuning.
Methodology
In this section, we commence with preliminaries. Next, we delve into the details of InstructVideo, which encompasses: 1) A reformulation of reward fine-tuning as editing that ensures computational efficiency and efficacy. 2) Segmental Video Reward (SegVR) and Temporally Attenuated Reward (TAR) that enable efficient reward fine-tuning with image reward models.
During inference, we adopt the DDIM sampling method for realistic video generation.
Reward fine-tuning. Reward fine-tuning aims to optimize a pre-trained model to enhance the expected rewards of a reward function . In our case, we target optimizing the parameters of a text-to-video diffusion model to enhance the expected reward of the generated videos given the text distribution:
where is the video generated from the sampled text via the diffusion model through the DDIM sampling chain. The reward function is typically a pre-trained model to assess the quality of the model output.
2 InstructVideo
In Fig. 2, we illustrate InstructVideo’s fine-tuning pipeline and elaborate on the technical contributions below.
Reward fine-tuning with diffusion models is costly due to the iterative refinement process during generation using DDIM . During generation, initial steps are crucial for shaping coarse, structural aspects of videos, with subsequent steps refining the coarse videos. Understanding that the essence of reward fine-tuning is not to drastically alter the model’s output but to subtly adjust it in line with human preferences, we propose to reinterpret reward fine-tuning as a form of editing . This perspective shift allows us to perform partial inference of the DDIM sampling chain, reducing computational demands and easing optimization.
To implement this idea, we first curate a small amount of fine-tuning data from pre-training data for reward fine-tuning. For each video-text pair , we acquire the video’s latent embedding as stated in Sec. 3.1. We aim to smooth out the video to eliminate undesirable artifacts and distortions . To achieve this, we leverage the diffusion process rather than DDIM inversion to enable efficient editing. If we denote the number of DDIM steps as and the number of pre-training DDPM steps as , we define a mapping that maps DDIM step index to the DDPM step indexFor example, if and , then the DDIM step sub-sequence is , i.e., ., formulated as:
Given the noise level , the targeted diffusion step for injecting noise is formulated as:
This allows us to obtain the starting point for reward fine-tuning via the diffusion process:
Based on , we can perform steps of DDIM sampling along the DDIM sub-sequence to obtain the edited result , which consumes of the computation of the full sampling chain. Utilizing the decoder , we decode in the latent space to in the RGB space.
2.2 Reward Fine-tuning with Image Reward Models
where denotes the aggregation function along index . To consider the impact of all frames in , an intuitive implementation of is the function.
However, the simple aggregation function leads to noticeable visual artifacts in the generated videos, such as structural twitching and color jittering. This issue arises because the function places equal weight on all frames, disregarding the inherent dynamic nature of videos where the reward scores of frames can vary throughout the sequence. To address this, we introduce TAR that strategically emphasizes central frames, with the emphasis tapering off towards the peripheral frames, thereby avoiding uniformly optimizing all frames’ reward scores to be equally high. We define the temporally attenuated coefficient as:
where controls the degree of the attenuating rate. We set by default. Incorporating this coefficient, we rewrite the reward score :
The optimization objective in Eq. 2 can be rewritten as:
3 Reward Fine-tuning and Inference
Data preparation and evaluation metric. We follow DDPO to experiment on prompts describing 45 animal species. In contrast to DDPO, since InstructVideo relies on video-text data, we select video-text pairs as the fine-tuning data from the base model’s pre-training dataset, i.e., WebVid10M, ensuring that no extra data is introduced. It is worth noting that we do not apply any quality filtering method to ensure that the selected videos are of high quality. Specifically, we select about 20 video-text pairs for each animal species. To evaluate the model’s ability to optimize the reward scores, we also collect evaluation data comprising about 6 prompts for each animal. We use HPSv2 score to measure the optimization performance of reward fine-tuning on the first frames of all segments.
Reward fine-tuning. We adopt the publicly available text-to-video diffusion model ModelScopeT2V as our base model. ModelScopeT2V is trained on WebVid10M with and is able to generate videos of resolution, which we divide into segments. By default, we adopt the differentiable HPSv2 as the reward model and perform 20-step DDIM inference, i.e., . Classifier-free guidance is adopted by default. Directly back-propagating the reward loss to the diffusion models can be computationally intensive and risks catastrophic forgetting . To circumvent these issues, we incorporate LoRA by default. To further accelerate fine-tuning, we truncate the gradient to only back-propagate the last DDIM sampling step following . Experiments are conducted on 4 NVIDIA A100s, with the batch size set to and the learning rate set to . To strike a cost-performance balance, we fine-tune InstructVideo with default parameters for 20 steps if not otherwise stated.
Inference. After reward fine-tuning, we merge the LoRA weights into the ModelScopeT2V parameters to ensure that InstructVideo’s inference cost is identical to ModelScopeT2V . For text-to-video generation, InstructVideo uses 20-step DDIM inference.
Experiments
Comparison with the base model ModelScopeT2V. To verify the efficacy of InstructVideo, we compare it with ModelScopeT2V utilizing and even DDIM steps in Fig. 3. Examining the examples, we observe that the quality of videos generated by InstructVideo consistently outperforms the base model by a margin. Specifically, notable enhancements include 1) clearer and more coherent structures and scenes even if the animal is moving, exemplified by the walking cat and the swimming fish; 2) more appealing coloration, exemplified by the sunflower, the bee and the mountain goat; 3) an enhanced delineation of scene details, exemplified by the rock and grass on the cliff, and the texture of all the animals; and 4) improved video-text alignment, exemplified by the distinct portrayal of sunflowers and the bird’s reflections on the water. Remarkably, these advancements are achieved without compromising motion fluidity and the resultant videos can often surpass the video quality of the WebVid10M dataset. Notably, InstructVideo even attenuates watermarks present in WebVid10M. These qualitative leaps, consistently favored by human annotators, are attributed to the reward fine-tuning process, which effectively refines the video diffusion model.
Comparison with other reward fine-tuning methods. This aims to validate the efficacy of reward fine-tuning conceptualized as an editing process. We compare with other representative reward fine-tuning methods, including policy gradient algorithm, DDPO , reward-weighted regression, RWR and direct reward back-propagation method, DRaFT . For DRaFT, we adopt DRaFT-1 for efficient fine-tuning. SegVR and TAR are employed for all compared methods to standardize reward signals. The comparative analysis on the evaluation set, presented in Fig. 5(a), leads to two findings: 1) Both RWR and DDPO exhibit a performance plateau after about 11 hours of fine-tuning, with further optimization failing to enhance or even deteriorating performance. 2) Direct reward back-propagation methods, including InstructVideo and DRaFT, initially lag during the first 11 hours but subsequently demonstrate fine-tuning efficiency, especially InstructVideo. To further validate the efficacy of our method, we provide visual comparisons in Fig. 4, where we adopt the optimal fine-tuned checkpoint for each method in Fig. 5(a). The examples reflect InstructVideo ’s superiority, evident in: 1) the clarity and coherence of structural and scenic elements, 2) the vibrancy of colors, 3) the precision in depicting intricacies, and 4) enhanced video-text alignment.
Generalization to unseen text prompts. We assessed the model’s generalization capabilities using two distinct sets of prompts: 1) those describing new animals and 2) those related to non-animals, none of which are present in the fine-tuning data. For this purpose, we curate about 4 prompts each for 6 new animal species following and 46 prompts for non-animals. Our comparative analysis of InstructVideo, the base model ModelScopeT2V and other reward fine-tuning methods is presented in Tab. 1. It is worth noting that during fine-tuning, SegVR and TAR are adopted by default to provide reward signals. Our observations are threefold: 1) Increasing the number of DDIM steps enhances the video quality. 2) Compared to the base model ModelScopeT2V in the second row, all methods improve reward scores for these unseen prompts. 3) Among all methods, InstructVideo outperforms other alternatives, affirming its superior generalization capabilities. To further demonstrate this, we provide visual comparisons in Fig. 6, featuring a set of unseen animal species, various sceneries, and human figures. The presented examples display an enhanced quality and exhibit an improvement in video-text alignment.
User study. To further qualitatively compare the videos generated by InstructVideo and other methods, we conduct a user study comparing our methods with the base model ModelScopeT2V and the reward fine-tuned model DRaFT in Tab. 2. We recruited two participants who have related research experience in generative models to assess the quality of videos in terms of video quality and video-text alignment. To simplify the annotation process, participants were presented with pairs of videos and asked to identify which video was superior or if both were of equal quality. To ensure a comprehensive comparison, we chose 60 prompts from the 45 fine-tuning animal species, 20 prompts from the 6 new animal species and 20 prompts describing non-animals. More details are presented in the Appendix. We observe that our method consistently outperforms other methods. Specifically, improvements in video quality, a noted shortcoming of the base model, are more pronounced than improvements in video-text alignment.
2 Ablation Study
The effect of varying noise level . To determine the optimal choice for noise level , we vary its value and evaluate its impact on reward scores using the evaluation data. We illustrate the results in Fig. 5(b). An increase in from to correlates with a progressive enhancement in the highest reward scores achieved by InstructVideo. However, excessively prolonged fine-tuning precipitates a sharp decline in generative performance. This phenomenon can be attributed to the limited edited space available to the model at lower noise levels, which constricts its ability to find the optimal output space as directed by the reward scores. When we further increase from to , we observe that the reward score enhancement per hour becomes minor, suggesting challenges associated with generating from an extended sampling chain. Optimally, a noise level of strikes a balance, providing a feasible starting point for editing that still allows for a substantial exploration of the edited space. After 20 steps, more optimization leads to over-optimization , meaning that further steps can degrade the visual quality of the output despite potential increases in the reward score. Thus, we finalize on with 20 steps of fine-tuning.
The effect of varying . To determine the optimal choice for , we vary its value and evaluate its impact. We illustrate the results in Fig. 5(c). A relatively high value , such as , results in decaying exponentially faster towards those border frames, thus providing diminished reward signals. A relatively low value , such as , leads to decaying more gently towards those border frames, thus strengthening the reward signals. This equalized weighting across frames can destabilize fine-tuning, leading to a precipitous decline in reward scores. Subsequent increases in scores do not necessarily indicate improved video quality but indicate quality degradation, as revealed by the rising variance. Thus, an appropriate coefficient to ensure stable fine-tuning is imperative and we finalize on .
Ablation on SegVR and TAR. To qualitatively verify the efficacy of SegVR and TAR, we present illustrative results in Fig. 7. Removing either SegVR or TAR results in a noticeable reduction in temporal modeling capabilities. This suggests that overly dense or excessively strong reward signals can lead to generation collapse. The degradation of modeling temporal dynamics often leads to the degraded quality of the individual frames due to the intertwined nature of spatial and temporal parameters. These observations underscore the critical roles of SegVR and TAR in maintaining fine-tuning stability.
3 Further Analysis
The evolution of the generated videos during reward fine-tuning. To elucidate how the reward fine-tuning works, we present a visual progression in Fig. 8. The top row depicts a video generated without fine-tuning. All the frames in this video exhibit a lack of the dog’s fur texture. Moreover, a notable blurriness characterizes the third frame due to sudden and unanticipated motion, while the fourth frame suffers from a loss of facial clarity. As reward fine-tuning proceeds, we observe a noticeable enhancement in terms of all aspects mentioned above. Surprisingly, watermarks, which are consistently present across the dataset, also gradually fade. The resultant video is full of clear details and aesthetically pleasing coloration.
Impact of fine-tuning data quality on the fine-tuning results. To investigate this, we self-collect a dataset comprising an equivalent number of video-caption pairs for 45 animal species, which are employed for fine-tuning. The results are illustrated in Fig. 9. We employ horizontal dashed lines to indicate the quality of different data, inferred from reward scores. While the variance of the generated videos are comparable for two kinds of fine-tuning data, the employment of higher-quality data, i.e., WebVid10M, yields superior average reward scores compared to that obtained using the lower-quality counterpart. This suggests that superior fine-tuning data can facilitate reward fine-tuning.
Constraints of fine-tuning data on resultant video quality. Fig. 9 showcases that InstructVideo is capable of generating videos that achieve reward scores significantly exceeding those of the fine-tuning data itself, as denoted by the horizontal dashed lines. This observation leads us to conclude that the quality of the fine-tuning data does not impose a ceiling on the potential quality of the fine-tuned results. Our fine-tuning pipeline has the propensity to surpass the initial data quality, thus facilitating the generation of videos with substantially enhanced reward scores.
Conslusion
In this paper, we introduce InstructVideo, a method that pioneers instructing video diffusion models with human feedback by reward fine-tuning. We recast reward fine-tuning as an editing process that mitigates computational burden and enhances fine-tuning efficiency. We resort to image reward models to provide human feedback on generated videos and propose SegVR and TAR to ensure effective fine-tuning. Extensive experiments validate that InstructVideo not only elevates visual quality but also maintains robust generalization capabilities.
References
Appendix A Potential Societal Impact
InstructVideo, as the pioneering effort in instructing video diffusion models with human feedback, prioritizes users’ preferences for AI-generated content. We conducted this research, motivated by the varied quality of generated videos induced by the varied quality of the curated web-scale datasets. Pre-training models on such unfiltered data can lead to outputs that deviate from human preferences. In the context of the broader research community, we advocate that video generation systems, akin to other generative models like language models , should prioritize ethical considerations and human values.
Moreover, conventional video generation systems might not always resonate with all users in terms of aesthetic style and often struggle with accurately reflecting textual prompts. InstructVideo steps in as a human-centered technology, efficiently addressing these issues in a data- and computation-efficient way and opening up possibilities for commercial applications, particularly in sectors like education and entertainment.
However, as InstructVideo primarily targets research, aiming at investigating the practicality of aligning video diffusion models with human preferences, its deployment to any circumstance beyond research should be approached with thorough oversight and evaluation to ensure responsible and ethical use.
Appendix B Limitation and Future work
We recognize that InstructVideo, as an initial endeavor in this area, comes with its limitations. Although we validate the efficacy of image reward models, we anticipate that specialized video reward models capturing human preferences might be even more superior since they evaluate one generated video as a whole. Additionally, as a common issue mentioned in previous works , reward fine-tuning carries a risk of over-optimization, meaning that excessive optimization steps will result in the degradation of the video quality despite potential increases in the reward score. Addressing these aspects presents avenues for future research, including the development of a more advanced video reward model and the design of strategic mechanisms to identify and ameliorate over-optimization.
Appendix C More Details about Implementing LoRA
To instantiate LoRA for efficient tuning, we adopt the implementation used in Diffusershttps://github.com/huggingface/diffusers/blob/main/src/diffusers/models/lora.py. Specifically, we configure the intrinsic rank within LoRA to 4 to ensure fast processing. LoRA modifications are applied to every Transformer layer within our model, targeting the linear layers responsible for query, key, value, and output projections. ModelScopeT2V contains 1,347.44M parameters, whereas the additional parameters introduced by adding LoRA amount to only 1.58M – approximately 0.1% of the total ModelScopeT2V parameters.
Appendix D More Details about User Study
In the main paper, we present a user study to demonstrate the effectiveness of InstructVideo. This study involves a comparative analysis of videos generated by InstructVideo and other methods, focusing on two key aspects: video quality and video-text alignment. For video quality, we asked annotators to evaluate: 1) The overall visual quality of the videos, 2) Alignment with general human aesthetic preferences, such as pleasing visuals, texture and details, and 3) The smoothness and consistency in terms of structural and color transitions within the video. Regarding video-text alignment, annotators are tasked with determining the extent to which the generated videos accurately and clearly represent the content of the provided text prompts. This assessment included evaluating the depiction of entities, attributes, relationships, and motions as described in the prompts. To simplify the evaluation process, annotators are asked to perform pairwise comparisons between videos, thereby streamlining their task to direct contrasts rather than isolated assessments.
Appendix E More Visualization Results
We provide more visualization results to exemplify the conclusions we draw in the main paper, including: 1) More results demonstrating how the generated videos evolve as the fine-tuning process proceeds as shown in Fig. A.2; 2) More results showcasing the comparison between InstructVideo and the base model ModelScopeT2V as illustrated in Fig. A.3; 3) More results exemplifying the comparison between InstructVideo and other reward fine-tuning methods as shown in Fig. A.4; 4) More results showing the InstructVideo’s generalization capabilities to unseen text prompts as shown in Fig. A.5.
Appendix F 50-Step Generation with InstructVideo
To showcase the adaptability and effectiveness of InstructVideo, we conduct experiments using a 50-step DDIM inference for generation after initial fine-tuning with 20-step DDIM inference. The results are shown in Tab. A.1. We observe that InstructVideo, despite being fine-tuned with a 20-step protocol, remains effective under a longer 50-step DDIM inference protocol, as demonstrated by the boosted reward scores. We present several cases to further illustrate InstructVideo’s efficacy as shown in Fig. A.6. We observe that both inference schemes can significantly improve over the base model and adopting more inference steps can occasionally lead to better results.
Appendix G 50-Step Reward Fine-tuning
To assess the adaptation of InstructVideo to different DDIM steps, we experiment on reward fine-tuning with the commonly-used 50-step DDIM inference and evaluate its 20-step generation quality for a fair comparison. We present the results in Fig. A.1. The results demonstrate that InstructVideo could be optimized towards higher reward scores in both settings. However, utilizing 50 steps degrades the fine-tuning efficiency, likely due to the increased computation brought by longer sampling chains.
Appendix H Adaptation to Other Reward Functions
In the main paper, we focus on the utilization of HPSv2 . To further validate the generalization of our method to other reward functions, we explore the application of ImageReward as our reward model. ImageReward is a general-purpose text-to-image human preference reward model, fine-tuned on BLIP . We perform reward fine-tuning as HPSv2 and present results in Fig. A.7. We observe that the quality of the videos are generally boosted in terms of structures, color vibrancy and details, despite that the stylistic aspects of the videos differ from those fine-tuned with HPSv2.