T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback
Jiachen Li, Weixi Feng, Tsu-Jui Fu, Xinyi Wang, Sugato Basu, Wenhu Chen, William Yang Wang
Introduction
Diffusion model (DM) (Sohl-Dickstein et al., 2015; Ho et al., 2020) has emerged as a powerful framework for neural image (Betker et al., 2023; Rombach et al., 2022; Esser et al., 2024; Saharia et al., 2022) and video synthesis (Singer et al., 2022; Ho et al., 2022a; He et al., 2022; Wang et al., 2023b; Zhang et al., 2023), leading to the development of cutting-edge text-to-video (T2V) models like Sora (Brooks et al., 2024), Gen-2 (Esser et al., 2023) and Pika (Pika Labs, 2023). Although the iterative sampling process of these diffusion-based models ensures high-quality generation, it significantly slows down inference, hindering their real-time applications. On the other hand, existing open-sourced T2V models including VideoCrafter (Chen et al., 2023, 2024) and ModelScopeT2V (Wang et al., 2023c) are trained on web-scale video datasets, e.g., WebVid-10M (Bain et al., 2021), with varying video qualities. Consequently, the generated videos often appear visually unappealing and fail to align accurately with the text prompts, deviating from human preferences.
Efforts have been made to address the issues listed above. To accelerate the inference process, Wang et al. (2023a) applies the theory of consistency distillation (CD) (Song et al., 2023; Song and Dhariwal, 2023; Luo et al., 2023a) to distill a video consistency model (VCM) from a teacher T2V model, enabling plausible video generations in just 4-8 inference steps. However, the quality of VCM’s generations is naturally bottlenecked by the performance of the teacher model, and the reduced number of inference steps further diminishes its generation quality. On the other hand, to align generated videos with human preferences, InstructVideo (Yuan et al., 2023) draws inspiration from image generation techniques (Dong et al., 2023; Clark et al., 2023; Prabhudesai et al., 2023) and proposes backpropagating the gradients of a differentiable reward model (RM) through the iterative video sampling process. However, calculating the full reward gradient is prohibitively expensive, resulting in substantial memory costs. Consequently, InstructVideo truncates the sampling chain by limiting gradient calculation to only the final DDIM step, compromising optimization accuracy. Additionally, InstructVideo is limited by its reliance on an image-text RM, which fails to fully capture the transition dynamic of a video. Empirically, InstructVideo only conducts experiments on a limited set of user prompts, the majority of which are related to animals. As a result, its generalizability to a broader range of prompts remains unknown.
In this paper, we aim to achieve fast and high-quality video generation by breaking the quality bottleneck of a VCM. We introduce T2V-Turbo, which integrates reward feedback from a mixture of RMs into the process of distilling a VCM from a teacher T2V model. Besides utilizing an image-text RM to align individual video frames with human preference, we further incorporate reward feedback from a video-text RM to comprehensively evaluate the temporal dynamics and transitions in the generated videos. We highlight that our reward optimization avoids tackling the highly memory-intensive issues associated with backpropagating gradients through an iterative sampling process. Instead, we directly optimize rewards of the single-step generations that arise from computing the CD loss, effectively bypassing the memory constraints faced by conventional methods that optimize a DM (Yuan et al., 2023; Xu et al., 2024; Clark et al., 2023; Prabhudesai et al., 2023).
Empirically, we demonstrate the superiority of our T2V-Turbo in generating high-quality videos within 4-8 inference steps. To illustrate the applicability of our methods, we distill T2V-Turbo (VC2) and T2V-Turbo (MS) from VideoCrafter2 (Chen et al., 2024) and ModelScopeT2V (Wang et al., 2023c), respectively. Remarkably, the 4-step generation results from both variants of our T2V-Turbo outperform SOTA models on the video evaluation benchmark VBench (Huang et al., 2024), even surpassing proprietary systems such as Gen-2 (Esser et al., 2023) and Pika (Pika Labs, 2023) that are trained with extensive resources. We further corroborate the results by conducting human evaluation using 700 prompts from the EvalCrafter (Liu et al., 2023) benchmark, validating that the 4-step generations from T2V-Turbo are favored by human over the 50-step DDIM samples from their teacher T2V models, which represents over tenfold inference acceleration and enhanced video generation quality.
Learn a T2V model with feedback from a mixture of RMs, including a video-text model. To the best of our knowledge, we are the first to do so.
Establish a new SOTA on the VBench with only 4 inference steps, outperforming proprietary models trained with substantial resources.
4-step generations from our T2V-Turbo are favored over the 50-step generation from its teacher T2V model as evidenced by human evaluation, representing over 10 times inference acceleration with quality improvement.
Preliminaries
Diffusion models (DMs). In the forward process, DMs progressively inject Gaussian noise into the original data distribution and perturb it into a marginal distribution with the transition kernel at timestep . and correspond to the noise schedule. In the reverse process, DMs sequentially recover the data from a noise sampled from the prior distribution . The reverse-time SDE can be modeled by an ordinary differential equation (ODE), known as the Probability Flow (PF-ODE) (Song et al., 2020a):
where and are the drift and diffusion coefficients, respectively, with the following properties:
The PF-ODE’s solution trajectories, when sampled at any timestep , align with the distribution . Empirically, a denoising model is trained to approximate the score function via score matching. During the sampling phase, one begins with a sample and follows the empirical PF-ODE below to obtain a sample .
Consistency Distillation. Conventional methods (Ho et al., 2020; Song et al., 2020b) generate their samples by solving the PF-ODE sequentially, leading to DM’s slow inference speed. To tackle this problem, consistency models (CM) (Song et al., 2023; Song and Dhariwal, 2023) propose to learn a consistency function to directly map any on the PF-ODE trajectory to its origin, where is a fixed small positive number. And thus, the consistency function has the following self-consistency property
where and are from the same PF-ODE. We can model with a CM . When tackling the PF-ODE of a T2V model that operates on the video latent space , we aim to learn a video consistency model (VCM) (Luo et al., 2023a; Wang et al., 2023a) . To ensure , we parameterize as
where and are differentiable functions with and , and is modeled as a neural network. We can distill a from a pre-trained T2V DM by minimizing the consistency distillation (CD) Song et al. (2023); Luo et al. (2023a) loss as below
where is a distance function. is updated by the exponential moving average (EMA) of , i.e., . is an estimate of obtained by the numerical augmented PF-ODE solver parameterized by and is the skipping interval
We follow the LCM paper (Luo et al., 2023a) to use DDIM Song et al. (2020b) as the ODE solver and defer the formula of the DDIM solver to Appendix A.
Training T2V-Turbo with Mixed Reward Feedback
In this section, we present the training pipeline to derive our T2V-Turbo. To facilitate fast and high-quality video generation, we integrate reward feedback from multiple RMs into the LCD process when distilling from a teacher T2V model. Figure 2 provides an overview of our framework. Notably, we directly leverage the single-step generation arise from computing the CD loss (6) and optimize the video decoded from it towards multiple differentiable RMs. As a result, we avoid the challenges associated with backpropagating gradients through an iterative sampling process, which is often confronted by conventional methods optimizing DMs (Clark et al., 2023; Xu et al., 2024; Yuan et al., 2023).
In particular, we leverage reward feedback from an image-text RM to improve human preference on each individual video frame (Sec. 3.1) and further utilize the feedback from a video-text RM to improve the temporal dynamics and transitions in the generated video (Sec. 3.2).
Chen et al. (2024) achieve high-quality video generation by including high-quality images as single-frame videos when training the T2V model. Inspired by their success, we align each individual video frame with human preference by optimizing towards a differentiable image-text RM . In particular, we randomly sample a batch of frames from the decoded video and maximize their scores evaluated by as below
2 Optimizing Video-Text Feedback Model
Existing image-text RMs (Wu et al., 2023a; Xu et al., 2024; Kirstain et al., 2024) are limited to assessing the alignment between individual video frames and the text prompt and thus cannot evaluate through the temporal dimensions that involve inter-frame dependencies, such as motion dynamic and transitions (Huang et al., 2024; Liu et al., 2023). To address these shortcomings, we further leverage a video-text RM to assess the generated videos. The corresponding objective is given below
3 Summary
To this end, we can define the total learning loss of our training pipeline as a linear combination of the in (6), in (8), and in (9) with weighting parameters and .
To reduce memory and computational cost, we initialize our T2V-Turbo with the teacher model and only optimize the LoRA weights (Hu et al., 2021; Luo et al., 2023b) instead of performing full model training. After completing the training, we merge the LoRA weights so that the per-step inference cost of our T2V-Turbo remains identical to the teacher model. We include pseudo-codes for our training algorithm in Appendix B.
Experimental Results
Our experiments aim to demonstrate our T2V-Turbo’s ability to generate high-quality videos with 4-8 inference steps. We first conduct automatic evaluations on the standard benchmark VBench (Huang et al., 2024) to comprehensively evaluate our methods from various dimensions (Sec. 4.1) against a broad array of baseline methods. We then perform human evaluations with 700 prompts from the EvalCrafter (Liu et al., 2023) to compare the 4-step and 8-step generations from our T2V-Turbo with the 50-step generations from the teacher T2V models as well as the 4-step generations from the baseline VCM (Sec. 4.2). Finally, we perform ablation studies on critical design choices (Sec. 4.3).
Settings. We train T2V-Turbo (VC2) and T2V-Turbo (MS) by distilling from the teacher diffusion-based T2V models VideoCrafter2 (Chen et al., 2024) and ModelScopeT2V (Wang et al., 2023c), respectively. Similar to both teacher models, we conduct our training using the WebVid10M (Bain et al., 2021) datasets. We train our models on 8 NVIDIA A100 GPUs for 10K gradient steps without gradient accumulation. We set the batch size of training videos to 1 for each GPU device. We employ HPSv2.1 (Wu et al., 2023a) as our image-text RM . When distilling from VideoCrafter2, we utilize the 2nd Stage model of InternVideo2 (InternVid2 S2) (Wang et al., 2024) as our video-text RM . When distilling from ModelScopeT2V, we set to be ViCLIP (Wang et al., 2023d). To optimize (8), we randomly sample 6 frames from the video by setting . For the hyperparameters (HP), we set learning rate and guidance scale range . We use DDIM (Song et al., 2020b) as our ODE solver and set the skipping step . For T2V-Turbo (VC2), we set and . For T2V-Turbo (MS), we set and . We include further training details in Appendix A.
We evaluate our T2V-Turbo (VC2) and T2V-Turbo (MS) on the standard video evaluation benchmark VBench (Huang et al., 2024) to compare against a wide array of baseline methods. VBench is designed to comprehensively evaluate T2V models from 16 disentangled dimensions. Each dimension in VBench is tailored with specific prompts and evaluation methods.
Table 1 compares the 4-step generation of our methods with various baselines from the VBench leaderboardhttps://huggingface.co/spaces/Vchitect/VBench_Leaderboard, including Gen-2 (Esser et al., 2023), Pika (Pika Labs, 2023), VideoCrafter1 (Chen et al., 2023), VideoCrafter2 (Chen et al., 2024), Show-1 (Zhang et al., 2023), LaVie (Wang et al., 2023b), and ModelScopeT2V (Wang et al., 2023c). Table 4 in Appendix further compares our methods with VideoCrafter0.9 (He et al., 2022), LaVie-Interpolation (Wang et al., 2023b), Open-Sora (Open-Sora, 2024), and CogVideo (Hong et al., 2022). The performance of each baseline method is directly reported from the VBench leaderboard. To obtain the results of our methods, we carefully follow VBench’s evaluation protocols by generating 5 videos for each prompt to calculate the metrics. We further train VCM (VC2) and VCM (MS) by distilling from VideoCrafter2 and ModelScopeT2V, respectively, without incorporating reward feedback, and then compare their results.
VBench has developed its own rules to calculate the Total Score, Quality Score, and Semantic Score. Quality Score is calculated with the 7 dimensions from the top table. Semantic Score is calculated with the 9 dimensions from the bottom table. And Total Score is a weighted sum of Quality Score and Semantic Score. Appendix C provides further details, including explanations for each dimension of VBench. As shown in Table 1, the 4-step generations of both our T2V-Turbo (MS) and T2V-Turbo (VC2) surpass all baseline methods on VBench in terms of Total Score. These results are particularly remarkable given that we even outperform the proprietary systems Gen-2 and Pika, which are trained with extensive resources. Even when distilling from a less advanced teacher model, ModelScopeT2V, our T2V-Turbo (MS) attains the second-highest Total Score, just below our T2V-Turbo (VC2). Additionally, our T2V-Turbo breaks the quality bottleneck of a VCM by outperforming its teacher T2V model, significantly improving over the baseline VCM.
2 Human Evaluation with 700 EvalCrafter Prompts
To verify the effectiveness of our T2V-Turbo, we compare the 4-step and 8-step generations from our T2V-Turbo with the 50-step DDIM samples from the corresponding teacher T2V models. We further compare the 4-step generations between our T2V-Turbo and their baseline VCMs when distilled from the same teacher T2V model.
We leverage the 700 prompts from the EvalCrafter (Liu et al., 2023) video evaluation benchmark, which are constructed based on real-world user data. We hire human annotators from Amazon Mechanical Turk to compare videos generated from different models given the same prompt. For each comparison, the annotators need to answer three questions: Q1) Which video is more visually appealing? Q2) Which video better fits the text description? Q3) Which video do you prefer given the prompt? Appendix D includes additional details about how we setup the human evaluations.
Figure 3 provides the full human evaluation results. We also qualitatively compare different methods in Figure 4. Appendix F further includes additional qualitative comparison results. Notably, the 4-step generations from our T2V-Turbo are favored by humans over the 50-step generation from their teacher T2V model, representing a 25 times inference acceleration with improving performance. By increasing the inference steps to 8, we can further improve the visual quality and text-video alignment of videos generated from our T2V-Turbo, reflected by the fact that our 8-step generations are more likely to be favored by the human compared to our 4-step generations in terms of all 3 evaluated metrics. Additionally, our T2V-Turbo significantly outperforms its baseline VCM, demonstrating the effectiveness of methods that incorporate a mixture of reward feedback into the model training.
3 Ablation Studies
We are interested in the effectiveness of each RM, and especially in the impact of the video-text RM . Therefore, we ablate and and experiment with different choices of . In Appendix E, we further experiment with different choices of .
Ablating RMs and . Recall that the training of our T2V-Turbo incorporate reward feedback from both and . To demonstrate the effectiveness of each individual RM, we perform ablation study by training VCM (VC2) + and VCM (VC2) + , which only incorporate feedback from and , respectively. Again, we evaluate the 4-step generations from different methods on VBench. Results in Table 2 show that incorporating feedback from either or leads to performance improvement over the baseline VCM. Notably, optimizing alone can already lead to substantial performance gains, while incorporating feedback from can further improve the Semantic Score on VBench, leading to better text-video alignment.
Effect of different choices of . We investigate the impact of different choices of by training T2V-Turbo (VC2) and T2V-Turbo (MS) by setting as ViCLIP (Wang et al., 2023d) and the second stage model of Intervideo2 (InternVid2 S2). In terms of model architecture, ViCLIP employs the CLIP (Radford et al., 2021) text encoder while InternVid2 S2 leverages the BERT-large (Kenton and Toutanova, 2019) text encoder. Additionally, InternVid2 S2 outperforms ViCLIP in several zero-shot video-text retrieval tasks. As shown in Table 3, T2V-Turbo (VC2) can achieve decent performance on VBench when integrating feedback from either ViCLIP or InternVid2 S2. Conversely, T2V-Turbo (MS) performs better with ViCLIP (Wang et al., 2023d). Nevertheless, with InternVid2 S2, our T2V-Turbo (MS) still surpasses VCM (MS) + .
Related Work
Diffusion-based T2V Models. Many diffusion-based T2V models rely on large-scale image datasets for training (Ho et al., 2022a; Wang et al., 2023c; Chen et al., 2023) or inherit weights from pre-trained text-to-image (T2I) models (Zhang et al., 2023; Blattmann et al., 2023; Khachatryan et al., 2023). The scale of text-image datasets (Schuhmann et al., 2022) is usually more than ten times the scale of open-sourced video-text datasets (Bain et al., 2021; Wang et al., 2023d) and with higher spatial resolution and diversity (Wang et al., 2023c). For example, Imagen Video (Ho et al., 2022b) discovers that joint training on a mix of image and video datasets improves the overall visual quality and enables the generation of videos in novel styles. Models trained with WebVid-10M (Bain et al., 2021) like ModelScopeT2V (Wang et al., 2023c) or VideoCrafter (Chen et al., 2023) also treat images as a single-frame video, and use them to improve video qualities. LaVie (Wang et al., 2023b) initialize the training with WebVid-10M and LAION-5B and then continue the training with a curated internal dataset of 23M videos. To overcome the data scarcity of high-quality videos, VideoCrafter2 (Chen et al., 2024) proposes to disentangle motion from appearance at the data level so that it can be trained on high-quality images and low-quality videos. The data limitation of high-quality videos and aligned, accurate video captions has been a longstanding bottleneck of current T2V models. In this paper, we propose to combat this challenge by leveraging reward feedback from a mixture of RMs.
Accelerating inference of Diffusion Models. Various methods have been proposed to accelerate the sampling process of a DM, including advanced numerical ODE solvers Song et al. (2020b); Lu et al. (2022a, b); Zheng et al. (2022); Dockhorn et al. (2022); Jolicoeur-Martineau et al. (2021) and distillation techniques Luhman and Luhman (2021); Salimans and Ho (2021); Meng et al. (2023); Zheng et al. (2023). Recently, Consistency Model (Song et al., 2023; Luo et al., 2023a) is proposed to facilitate fast inference by learning a consistency function to map any point at the ODE trajectory to the origin. Li et al. (2024) proposes to augment consistency distillation with an objective to optimize image-text RM to achieve fast and high-quality image generation. Our work extends it for T2V generation, incorporating reward feedback from both an image-text RM and a video-text RM.
Vision-and-language Reward Models. There have been various open-sourced image-text RMs that are trained to mirror human preferences given a text-image pair, including HPS (Wu et al., 2023b, a), ImageReward Xu et al. (2024), and PickScore Kirstain et al. (2024), which are obtained by finetuning a image-text foundation model such as CLIP Radford et al. (2021) and BLIP Li et al. (2022), on human preference data. However, to the best of our knowledge, no video-text RMs, e.g., T2VScore (Wu et al., 2024), that mirrors human preference on a text-video pair has been released to the public. In this paper, we choose HPSv2.1 as our image-text RM and directly employ the video foundation models ViCLIP (Wang et al., 2023d) and InterVid S2 (Wang et al., 2024) that are trained for general video-text understanding as our video-text RM. Empirically, we show that incorporating feedback from these RMs can improve the performance of our T2V-Turbo.
Learning from Human/AI Feedback has been proven as an effective way to align the output from a generative model with human preference (Leike et al., 2018; Ziegler et al., 2019; Ouyang et al., 2022; Stiennon et al., 2020; Rafailov et al., 2024). In the field of image generation, various methods have been proposed to align a text-to-image model with human preference, including RL based methods Fan et al. (2024); Prabhudesai et al. (2023); Zhang et al. (2024) and backpropagation-based reward finetuning methods Clark et al. (2023); Xu et al. (2024); Prabhudesai et al. (2023). Recently, InstructVideo (Yuan et al., 2023) extends the reward-finetuning methods to optimize a T2V model. However, it still employs an image-text RM to provide reward feedback without considering the transition dynamic of the generated video. In contrast, our work incorporates reward feedback from both an image-text and video-text RM, providing comprehensive feedback to our T2V-Turbo.
Conclusion and Limitations
In this paper, we propose T2V-Turbo, achieving both fast and high-quality T2V generation by breaking the quality bottleneck of a VCM. Specifically, we integrate mixed reward feedback into the VCD process of a teacher T2V model. Empirically, we illustrate the applicability of our methods by distilling T2V-Turbo (VC2) and T2V-Turbo (MS) from VideoCrafter2 (Chen et al., 2024) and ModelScopeT2V (Wang et al., 2023c), respectively. Remarkably, the 4-step generations from both our T2V-Turbo outperform SOTA methods on VBench (Huang et al., 2024), even surpassing their teacher T2V models and proprietary systems including Gen-2 (Esser et al., 2023) and Pika (Pika Labs, 2023). Our human evaluation further corroborates the results, showing the 4-step generations from our T2V-Turbo are favored by humans over the 50-step DDIM samples from their teacher, which represents over ten-fold inference acceleration with quality improvement.
While our T2V-Turbo marks a critical advancement in efficient T2V synthesis, it is important to recognize certain limitations. Our approach utilizes a mixture of RMs, including a video-text RM . Due to the lack of an open-sourced video-text RM trained to reflect human preferences on video-text pairs, we instead use video foundation models such as ViCLIP (Wang et al., 2023d) and InternVid S2 (Wang et al., 2024) as our . Although incorporating feedback from these models has enhanced our T2V-Turbo’s performance, future research should explore the use of a more advanced for training feedback, which could lead to further performance improvements.
References
Appendix A Experiment and Hyperparameter (HP) Details
When performing qualitative comparisons between different methods, we ensure to use the same random seed for head-to-head video comparisons.
As mentioned in Sec. 4, we train T2V-Turbo (VC2) and T2V-Turbo (MS) by distilling from the teacher diffusion-based T2V models VideoCrafter2 [Chen et al., 2024] and the less advanced ModelScopeT2V [Wang et al., 2023c], respectively. Specifically, VideoCrafter2 supports video FPS as an input and generates videos at the resolution of 512x320. For simplicity, we always set FPS to 16 when distilling our T2V-Turbo (VC2) from VideoCrafter2. On the other hand, ModelScopeT2V always generates video at 8 FPS at a resolution of 256x256.
We conduct our training with the WebVid10M [Bain et al., 2021] datasets. Note that both teacher T2V models are also trained with WebVid10M. We train our models for 10K gradient steps with 6 - 8 NVIDIA A100 GPUs without gradient accumulation and set the batch size to 1 for each GPU device. That is, we load 1 video with 16 frames. At each training iteration, we always sample 16 frames from the input video. We employ HPSv2.1 [Wu et al., 2023a] as our image-text RM . When distilling from VideoCrafter2, we utilize the 2nd Stage model of InternVideo2 [Wang et al., 2024] as our video-text RM . When distilling from ModelScopeT2V, we set to be ViCLIP [Wang et al., 2023d]. To optimize (8), we randomly sample 6 frames from the video by setting . For the hyperparameters, we set learning rate and guidance scale range . We use DDIM [Song et al., 2020b] as our ODE solver and set the skipping step . For T2V-Turbo (VC2), we set and . For T2V-Turbo (MS), we set and .
As mentioned in Sec. 2, we employ the DDIM Song et al. [2020b] ODE solver by following the practice of Luo et al. [2023a]. Its formula from to is given below
where denotes the noise prediction model from the teacher T2V model. We refer interested readers to the original LCM paper Luo et al. [2023a] for further details.
Appendix B Psudoe-codes for Training our T2V-Turbo
We include the pseudo-codes for training our T2V-Turbo in Algorithm 1. We use the red color to highlight the difference from the standard (latent) consistency distillation Luo et al. [2023a], Song et al. .
Appendix C Further Details about VBench
We provide a brief introduction of the metrics included in VBench [Huang et al., 2024] followed by introducing the derivation rules for the Quality Score, Semantic Score and Total Score. We refer interested readers to read the VBench paper for further details.
The following metrics are used to construct the Quality Score.
Subject Consistency (Subject Consist.) is calculated by the DINO [Caron et al., 2021] feature similarity across video frames.
Background Consistency (BG Consist.) is calculated by CLIP [Radford et al., 2021] feature similarity across video frames.
Temporal Flickering (Temporal Flicker.) is computed by the mean absolute difference across video frames.
Motion Smoothness (Motion Smooth.) is evaluated by motion priors in the video frame interpolation model [Li et al., 2023a].
Aesthetic Quality is calculated by mean of aesthetic scores evalauted by the LAION aesthetic predictor [Schuhmann et al., 2022].
Dynamic Degree is calculated using RAFT [Teed and Deng, 2020].
Image Quality is evaluated by the MUSIQ [Ke et al., 2021] image quality predictor.
Quality Score is calculated as the weighted sum of the normalized scores of each metric mentioned above. The weight for all metrics is 1, except for Dynamic Degree, which has a weight of 0.5.
The following metrics are used to construct the Semantic Score.
Object Class is calculated by detecting the success rate of generating the object specified by the user using GRiT [Wu et al., 2022].
Multiple Object is calculated by detecting the success rate of generating all objects specified in the prompt using GRiT [Wu et al., 2022].
Human Action is evaluated by the UMT model [Li et al., 2023b].
Color is calculated by comparing the color caption generated by GRiT [Wu et al., 2022] against the expected color.
Spatial Relationship (Spatial Relation.) is calculated by a rule-based method similar to [Huang et al., 2023a].
Scene is calculated by comparing the video captions generated by Tag2Text [Huang et al., 2023b] against the scene descriptions in the prompt.
Appearance Style (Appear Style.) is calculated by using ViCLIP [Wang et al., 2023d] to compare the video feature and the style description in the user prompt.
Temporal Style is calculated based on the similarity between the video feature and the style descrption feature provided by ViCLIP [Wang et al., 2023d].
Overall Consistency (Overall Consist.) is calculated based on the similarity between the video feature and the entire text prompt feature provided by ViCLIP [Wang et al., 2023d]. ViCLIP [Wang et al., 2023d]
Semantic Score is simply calculated as the mean of the normalized scores of each metric mentioned above. And the Total Score is the weighted sum of Quality Score and Semantic Score, which is given by
Appendix D Human Evaluation Details
Figure 5 shows the user interface displayed to the labelers when conducting our human evaluations. Each method generate videos of 16 frames using the 700 prompts from EvalCrafter [Liu et al., 2023]. For our T2V-Turbo (VC2), we collect its 4-step and 8-step generations and compare them with the 50-step DDIM samples from its teacher VideoCrafter2. For our T2V-Turbo (MS), we collect its 2-step and 4-step generations and compare them with the 50-step DDIM samples from its teacher ModelScopeT2V. We also compare the 4-step generations between our T2V-Turbo their baseline VCM, demonstrating the significant quality improvement of our methods.
As mentioned in Sec. 4.2, we hire labelers from Amazon Mechanical Turk platform and form the video comparison task as many batches of HITs. Specifically, we choose lablers from English-speaking countries, including AU, CA, NZ, GB, and US. Each task needs around 30 seconds to complete, and we pay each submitted HIT with 0.2 US dollars. Therefore, the hourly payment is about 24 US dollars.
We note that the data annotation part of our project is classified as exempt by Human Subject Committee via IRB protocols.
Appendix E Additional Ablation Studies
In this section, we provide the full ablation results performed in Table 3, which can be found in Table 5. We further examine the effect of different choices of . In the initial stage of our project, we train VCM (VC2) + with several different image-text RMs, including HPSv2.1 [Wu et al., 2023a], PickScore [Kirstain et al., 2024], and ImageReward [Xu et al., 2024]. We collect the 4-step generations from each method and qualitatively compare them with the 4-step generation from the baseline VCM. As shown in Figure 6, incorporating reward feedback from any of these leads to quality improvment over the baseline VCM (VC2).
Appendix F Qualitative Results
We provide additional qualitative comparisons between our T2V-Turbo, the baseline VCM, and their teacher T2V models in Figures 7, 8, 9, and 10.
Appendix G Broader Impact
The ability to create highly realistic synthetic videos raises concerns about misinformation and deepfakes, which can be used to manipulate public opinion, defame individuals, or perpetrate fraud. Addressing these concerns requires robust regulatory frameworks and ethical guidelines to ensure the technology is used responsibly and for the benefit of society. Therefore, we are committed to installing safeguard when releasing our models. Specifically, we will require users to adhere to usage guidelines.
Despite the challenges, the impact of our T2V-Turbo is profound, offering a scalable solution that significantly enhances the accessibility and practicality of generating high-quality videos at a remarkable speed. This innovation not only broadens the potential applications in fields ranging from digital art to visual content creation but also sets a new benchmark for future research in T2V synthesis, emphasizing the importance of human-centric design in the development of generative AI technologies.