EvalCrafter: Benchmarking and Evaluating Large Video Generation Models

Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, Ying Shan

Introduction

The charm of the large generative models is sweeping the world. e.g., the well-known ChatGPT and GPT4 have shown human-level abilities in several aspects, including coding, solving math problems, and even visual understanding, which can be used to interact with our human beings using any knowledge in a conversational way. As for the generative models for visual content creation, Stable Diffusion and SDXL play very important roles since they are the most powerful publicly available models that can generate high-quality images from any text prompts.

Beyond text-to-image, taming diffusion model for video generation has also progressed rapidly. Early works (Imagen-Viedo, Make-A-Video ) utilize the cascaded models for video generation directly. Powered by the image generation priors in Stable Diffusion, LVDM and MagicVideo have been proposed to train the temporal layers to efficiently generate videos. Apart from the academic papers, several commercial services also can generate videos from text or images. e.g., Gen2 and PikaLabs . Although we can not get the technique details of these services, they are not evaluated and compared with other methods. However, all current large text-to-video (T2V) model only uses previous GAN-based metrics, like FVD , for evaluation, which only concerns the distribution matching between the generated video and the real videos, other than the pairs between the text prompt and the generated video. Differently, we argue that a good evaluation method should consider the metrics in different aspects, e.g., the motion quality and the temporal consistency. Also, similar to the large language models, some models are not publicly available and we can only get access to the generated videos, which further increases the difficulties in evaluation. Although the evaluation has progressed rapidly in the large generative models, including the areas of LLM , MLLM , and text-to-image , it is still hard to directly use these methods for video generation. The main problem here is that different from text-to-image or dialogue evaluation, motion and consistency are very important to video generation which previous works ignore.

We make the very first step to evaluate the large multi-modality generative models for video. In detail, we first build a comprehensive prompt list containing various everyday objects, attributes, and motions. To achieve a balanced distribution of well-known concepts, we start from the well-defined meta-types of the real-world knowledge and utilize the knowledge of large language models, e.g., ChatGPT , to extend our meta-prompt to a wide range. Besides the prompts generated by the model, we also select the prompts from real-world users and text-to-image prompts. After that, we also obtain the metadata (e.g., color, size, etc.) from the prompt for further evaluation usage. Second, we evaluate the performance of these larger T2V models from different aspects, including the video visual qualities, the text-video alignment, and the motion quality and temporal consistency. For each aspect, we use one or more objective metrics as the evaluation metrics. Since these metrics only reflect one of the abilities of the model, we also conduct a multi-aspects user study to judge the model in terms of its qualities. After obtaining these opinions, we train the coefficients of each objective regression model to align the evaluation scores to the user’s choice, so that we can obtain the final scores of the models and also evaluate the new video using the trained coefficients.

Overall, we summarize the contribution of our paper as:

We make the first step of evaluating the large T2V model and build a comprehensive prompt list with detailed annotations for T2V evaluation.

We consider the aspects of the video visual quality, video motion quality, and text-video alignment for the evaluation of video generation. For each aspect, we align the opinions of humans and also verify the effectiveness of the proposed metric by human alignment.

During the evaluation, we also discuss several conclusions and findings, which might be also useful for further training of the T2V generation models.

Related Work

T2V generation aims to generate videos from the given text prompts. Early works generate the videos through Variational AutoEncoders (VAEs ) or generative adversarial network (GAN ). However, the quality of the generated videos is often low quality or can only work on a specific domain, e.g., face or landscape . With the rapid development of the diffusion model , video diffusion model , and large-scale text-image pretraining , current methods utilize the stronger text-to-image pre-trained model prior to generation. e.g., Make-A-Video and Imagen-Video train a cascaded video diffusion model to generate the video in several steps. LVDM , Align Your latent and MagicVideo extend the latent text-to-image model to video domains by adding additional temporal attention or transformer layer. AnimateDiff shows a good visual quality by utilizing the personalized text-to-image model. Similar methods are also been proposed by SHOW-1 and LAVIE T2V generation also raises the enthusiasm of commerce or non-commerce companies. For online model services, e.g., Gen1 and Gen2 , show the abilities of the high-quality generated video in the fully T2V generation or the conditional video generation. For discord-based servers, Pika-Lab , Morph Studio , FullJourney and Floor33 Pictures also show very competitive results. Besides, there are also some popular open-sourced text (or image)-to-video models, e.g., ZeroScope , ModelScope .

However, these methods still lack a fair and detailed benchmark to evaluate the advantages of each method. For example, they only evaluate the performance using FVD (LVDM , MagicVideo , Align Your Latent ), IS (Align Your Latent ), CLIP similarity (Gen1 , Imagen Video , Make-A-Video ), or user studies to show the performance level. These metrics might only perform well on previous in-domain text-to-image generation methods but ignore the alignment of input text, the motion quality, and the temporal consistency, which are also important for T2V generation.

2 Evaluations on Large Generative Models

Evaluating the large generative models is a big challenge for both the NLP and vision tasks. For the large language models, current methods design several metrics in terms of different abilities, question types, and user platform . More details of LLM evaluation and Multi-model LLM evaluation can be found in recent surveys . Similarly, the evaluation of the multi-modal generative model also draws the attention of the researchers . For example, Seed-Bench generates the VQA for multi-modal large language model evaluation.

For the models in visual generation tasks, Imagen only evaluates the model via user studies. DALL-Eval assesses the visual reasoning skills and social basis of the text-to-image model via both user and object detection algorithm . HRS-Bench proposes a holistic and reliable benchmark by generating the prompt with ChatGPT and utilizing 17 metrics to evaluate the 13 skills of the text-to-image model. TIFA proposes a benchmark utilizing the visual question answering (VQA). However, these methods still work for text-to-image evaluation or language model evaluation. For T2V evaluation, we consider the quality of motion and temporal consistency.

Benchmark Construction

Our benchmark aims to create a trustworthy prompt list to evaluate the abilities of various of T2V models fairly. To achieve this goal, we first collect and analyze the T2V prompt from large-scale real-world users. After that, we propose an automatic pipeline to increase the diversity of the generated prompts so that they can be identified and evaluated by pre-trained computer vision models. Since video generation is time-consuming, we collect 500 prompts as our initial version for evaluation with careful annotation. Below, we give the details of each step.

To answer this question, we collect the prompts from the real-world T2V generation discord users, including the FullJourney and PikaLab . In total, we get over 600k prompts with corresponding videos and filter them to 200k by removing repeated and meaningless prompts. Our first curiosity is how long a prompt should be generated, as shown in Fig. 2 2(a), 90% of the prompts contain the words in the range of $$. We also plot the most important words in Fig. 2 2(b) by removing some unclear words like video, camera, high, quality, etc., where the person, the style, human motion, and scene are dominant. Despite the above analysis, we also count the word class to decide the meta class of our prompt list. As shown in Fig. 2 2(c), we use WordNet to identify the meta classes, except for the communication, attribute, and cognition words, the artifacts (human-made objects), human, animal, and the location (landscape) play important roles. We also add the most important word style of Fig. 2 2(b) to the metaclass. Overall, we divide the T2V generation into roughly four meta-subject classes, including the human, animal, object, and landscape. For each type, we also consider the motions and styles of each type and the relationship between the current metaclass and other metaclasses to construct the video. Besides, we include the motion which is relevant to the main object and important for the video. Finally, we consider the camera motion and the style by template.

2 General Recognizable Prompt Generation

Automatically Prompt Generation. After deciding the meta classes of our prompt list, we generate the recognizable prompt by the power of a large language model (LLM) and humans. As shown in Fig 3, for each kind of meta class, we let GPT-4 describe the scenes about this meta class with randomly sampled meta information along with the attributes of the scenes so that we already know the labels. For example, for humans, we can ask GPT-4 to give us the attributes of humankind, age, gender, clothes, and human activity, which are saved as a JSON file as the ground truth computer vision models. However, we also find that the GPT-4 is not fully perfect for this task, the generated attributes are not very consistent with the generated description. Thus, we involve a self-check to the benchmark building, where we also use GPT-4 to identify the similarities of the generated description and the meta data. Finally, we filter the prompts by ourselves to make sure each prompt is correct and meaningful for T2V generation.

Prompts from Real World. Since we have already collected a very large scale of prompts from real-world users and there are also available text-to-image evaluation prompts, e.g., DALL-Eval and Draw-Bench , we also integrate these prompts to our benchmark list. To achieve this, we first filter and generate the metadata using GPT-4. Then, we choose the suitable prompts with the corresponding meta-information as shown in Fig. 3 and check the consistency of the meta-information.

3 Benchmark Analysis

Overall, we get over 500 prompts in the meta classes of human, animal, objects, and landscape. Each class contains the natural scenes, the stylized prompts, and the results with explicit camera motion controls. We give a brief view of the benchmark in Fig. 4. The whole benchmark contains over 500 prompts with careful categories. To increase the diversity of the prompts, our benchmark contains 3 different sub-types as shown in Figure 4, where we have a total of 50 styles and 20 camera motion prompts. We add them randomly in the 50% prompts of the whole benchmark. Our benchmark contains an average length of 12.5 words pre-prompt, which is also similar to the real-world prompts as we find in Figure. 2.

Evaluation Metrics

Different from previous FID based evaluation metrics, we evaluate the T2V models in different aspects, including the visual quality of the generated video, the text-video alignment, the content correctness, the motion quality, and temporal consistency. Below, we give the detailed metrics.

We first consider the visual quality of the generated video, which is the key for visually appealing to the users. Notice that, since the distribution-based method, e.g., FVD still needs the ground truth video for evaluation, we argue these kinds of metrics are not suitable for the general T2V generation cases.

Video Quality Assessment (VQAA, VQAT). We utilize the state-of-the-art video quality assessment method, Dover , to evaluate the quality of the generated video in terms of aesthetics and technicality, where the technical rating measures the quality of the generated video in terms of the common distortion, including noises, artifacts, etc. Dover is trained on a self-collected larger-scale dataset and the labels are ranked by the real users for alignment. We term the aesthetic and technical scores as VQAA and VQAT, respectively.

Inception Score (IS). Following previous metrics in the T2V generation papers, we also use the inception score of the video as one of the video quality assessment indexes. The inception score is proposed to evaluate the performance of GAN , which utilizes a pre-trained Inception Network on the ImageNet dataset as the pre-trained feature extraction method. The inception score reflects the diversity of the generated video, whereas a larger score means the generated content is more diverse.

2 Text-Video Alignment

Another common evaluation direction is the alignment of the input text and the generated video. We not only consider both the global text prompts and the video, and also the content correctness in different aspects. Below, we give the details of each score.

Text-Video Consistency (CLIP-Score). We incorporate the CLIP-Score as one of the evaluation metrics, given its widespread usage and simplicity in quantifying the discrepancy between input text prompts and generated videos. Utilizing the pretrained ViT-B/32 CLIP model as a feature extractor, we obtain frame-wise image embeddings and text embeddings, and compute their cosine similarity. The cosine similarity for the tt-th frame of the ii-th video xtix_{t}^{i} and the corresponding prompt pip^{i} is denoted as C(emb(xti),emb(pi))\mathcal{C}(emb(x_{t}^{i}),emb(p^{i})), emb(⋅)emb(\cdot) means CLIP embedding. The overall CLIP-Score, SCSS_{CS}, is derived by averaging individual scores across all frames and videos, calculated as

where MM is the total number of testing videos and NN is the total number of frames in each video.

Image-Video Consistency (SD-Score). Most current video diffusion models are fine-turned on a base stable diffusion with a larger scale dataset. Also, tuning the new parameters for stable diffusion will cause conceptual forgetting, we thus propose a new metric by comparing the generated quality with the frame-wise stable diffusion . In detail, we use SDXL to generate N1{N_{1}} images {dk}k=1N1\{d_{k}\}_{k=1}^{N_{1}}for every prompt and extract the visual embeddings in both generated images and video frames, and here we set N1N_{1} to 5. We calculate the embedding similarity between the generated videos and the SDXL images, which is helpful to ablate the concept forgotten problems when fine-tuneing the text-to-image diffusion model to video models. The final SD-Score is

Text-Text Consistency (BLIP-BLEU). We also consider the evaluation between the text descriptions of the generated video and the input text prompt. To this purpose, we utilize BLIP2 for caption generation. Similar to text-to-image evaluation methods , we use BLEU for text alignment of the generated and the source prompt across frames:

where B(⋅,⋅)\mathcal{B}(\cdot,\cdot) is the BLEU similarity scoring function, {lki}k=1N2\{l^{i}_{k}\}_{k=1}^{N_{2}} are BLIP generated captions for ii-th video, and N2N_{2} is set to 5 experimentally.

Object and Attributes Consistency (Detection-Score, Count-Score and Color-Score). For general objects, we employ a state-of-the-art segmentation and tracking method, namely SAM-Track , to analyze the correctness of the video content that we are interested in. Leveraging the powerful segmentation model , we can easily obtain the objects and their attributes. In our pipeline, we focus on detecting prompts with COCO classes , which is a widely used dataset for object detection and segmentation tasks. We evaluate T2V models on the existence of objects, as well as the correctness of color and count of objects in text prompts. Specifically, we assess the Detection-Score, Count-Score, and Color-Score as follows:

1. Detection-Score (SDetS_{Det}): Measures average object presence across videos, calculated as:

where M1M_{1} is the number of prompts with objects, and σji\sigma^{i}_{j} is the detection result for frame tt in video ii (1 if an object is detected, 0 otherwise).

2. Count-Score (SCountS_{Count}): Evaluates average object count difference, calculated as:

where M2M_{2} is the number of prompts with object counts, ctic^{i}_{t} is the detected object count frame tt in video ii and c^i\hat{c}^{i} is the ground truth object count for video ii.

3. Color-Score (SColor{S}_{Color}): Assesses average color accuracy, calculated as:

where M3M_{3} is the number of prompts with object colors, stis^{i}_{t} is the color accuracy result for frame ii in video tt (1 if the detected color matches the ground truth color, 0 otherwise).

Human Analysis (Celebrity ID Score). Human is important for the generated videos as shown in our collected real-world prompts. To this end, we also evaluate the correctness of human faces using DeepFace , a popular face analysis toolbox. We do the analysis by calculating the distance between the generated celebrities’ faces with corresponding real images of the celebrities.

where M4M_{4} is the number of prompts that contain celebrities, D(⋅,⋅)\mathcal{D}(\cdot,\cdot) is the Deepface’s distance function, {fki}k=1N3\{f^{i}_{k}\}_{k=1}^{N_{3}} are collected celebrities images for prompt ii, and N3N_{3} is set to 3.

Text Recognition (OCR-Score) Another hard case for visual generation is to generate the text in the description. To examine the abilities of current models for text generation, we utilize the algorithms from Optical Character Recognition (OCR) models similar to previous text-to-image evaluation or multi-model LLM evaluation method . Specifically, we utilize PaddleOCRhttps://github.com/PaddlePaddle/PaddleOCR to detect the English text generated by each model. Then, we calculate Word Error Rate (WER) , Normalized Edit Distance (NED) , Character Error Rate (CER) , and finally we average these three score to get the OCR-Score.

3 Motion Quality

For video, we believe the motion quality is a major difference from other domains, such as image. To this end, we consider the quality of motion as one of the main evaluation metrics in our evaluation system. Here, we consider two different motion qualities introduced below.

Action Recognition (Action-Score). For videos about humans, we can easily recognize the common actions via pre-trained models. In our experiments, we use MMAction2 toolbox , specifically the pre-trained VideoMAE V2 model, to infer the human actions in the generated videos. We then take the classification accuracy (ground truth are actions in the input prompts) as our Action-Score. In this work, we focus on Kinetics 400 action classes , which is widely used and encompasses human-object interactions like playing instruments and human-human interactions, including handshakes and hugs.

Average Flow (Flow-Score). We also consider the general motion information of the video. To this end, we use the pretrained optical flow estimation method, RAFT , to extract the dense flows of the video in every two frames. Then, we calculate the average flow on these frames to obtain the average flow score of every specific generated video clip since some methods are likely to generate still videos which are hard to identify by the temporal consistency metrics.

Amplitude Classification Score (Motion AC-Score). Based on the average flow, we further identify whether the motion amplitude in the generated video is consistent with the amplitude specified by the text prompt. To this end, we set an average flow threshold ρ\rho that if surpasses ρ\rho, one video will be considered large, and here ρ\rho is set to 2 based on our subjective observation. We mark this score to identify the movement of the generated video.

4 Temporal Consistency

Temporal consistency is also a very valuable field in our generated video. To this end, we involve several metrics for calculation. We list them below.

Warping Error. We first consider the warping error, which is widely used in previous blind temporal consistency methods . In detail, we first obtain the optical flow of each two frames using the pre-trained optical flow estimation network , then, we calculate the pixel-wise differences between the warped image and the predicted image. We calculate the warp differences on every two frames and calculate the final score using the average of all the pairs.

Semantic Consistency (CLIP-Temp). Besides pixel-wise error, we also consider the semantic consistency between every two frames, which is also used in previous video editing works . Specifically, we consider the semantic embeddings on each of the two frames of the generated videos and then get the averages on each two frames, which is shown as follows:

Face Consistency. Similar to CLIP-Temp, we evaluate the human identity consistency of the generated videos. Specifically, we select the first frame as the reference and calculate the cosine similarity of the reference frame embedding with other frames’ embeddings. Then, we average the similarities as the final score:

5 User Opinion Alignments

Besides the above objective metrics, we conduct user studies on the main five aspects to get the users’ opinions. These aspects include (1) Video Qualities. It indicates the quality of the generated video where a higher score shows there is no blur, noise, or other visual degradation. (2) Text and Video Alignment. This opinion considers the relationships between the generated video and the input text-prompt, where a generated video has the wrong count, attribute, and relationship will be considered as low-quality samples. (3) Motion Quality. In this metric, the users need to identify the correctness of the generated motions from the video. (4) Temporal Consistency. Temporal consistency is different from motion quality. In motion quality, the user needs to give a rank for high-quality movement. However, in temporal consistency, they only need to consider the frame-wise consistency of each video. (5) Subjective likeness. This metric is similar to the aesthetic index, a higher value indicates the generated video generally achieves human preference, and we leave this metric used directly.

For evaluation, we generate videos using the provided prompts benchmark on five state-of-the-art methods of ModelScope , ZeroScope , Gen2 , Floor33 , and PikaLab , getting 2.5k videos in total. For a fair comparison, we change the aspect ratio of Gen2 and PikaLab to 16:916:9 to suitable other methods. Also, since PikaLab can not generate the content without the visual watermark, we add the watermark of PikaLab to all other methods for a fair comparison. We also consider that some users might not understand the prompt well, for this purpose, we use SDXL to generate three reference images of each prompt to help the users understand better, which also inspires us to design an SD-Score to evaluate the models’ text-video alignments. For each metric, we ask three users to give opinions between 1 to 5, where a large value indicates better alignments. We use the average score as the final labeling and normalize it to range .

Upon collecting user data, we proceed to perform human alignment for our evaluation metrics, with the goal of establishing a more reliable and robust assessment of T2V algorithms. Initially, we conduct alignment on the data using the mentioned individual metrics above to approximate human scores for the user’s opinion in the specific aspects. We employ a linear regression model to fit the parameters in each dimension, inspired by the works of the evaluation of natural language processing . Specifically, we randomly choice 300 samples from four different methods as the fittings samples and left the rest 200 samples to verify the effectiveness of the proposed method (as in Table. 4). The coefficient parameters are obtained by minimizing the residual sum of squares between the human labels and the prediction from the linear regression model. In the subsequent stage, we integrate the aligned results of these four aspects and calculate the average score to obtain a comprehensive final score, which effectively represents the performance of the T2V algorithms. This approach streamlines the evaluation process and provides a clear indication of model performance.

Results

We conduct the evaluation on 500 prompts from our benchmark prompts, where each prompt has a metafile for additional information as the answer of evaluation. We generate the videos using all available high-resolution T2V models, including the ModelScope , Floor33 Pictures , and ZeroScope . We keep all the hyper-parameters, such as classifier-free guidance, as the default value. For the service-based model, we evaluate the performance of the representative works of Gen2 and PikaLab . They generate at least 512p videos with high-quality watermark-free videos. Before our evaluation, we show the differences between each video type in Table 1, including the abilities of these models, the generated resolutions, and fps. As for the comparison on speed, we run all the available models on an NVIDIA A100. For the unavailable model, we run their model online and measure the approximate time. Notice that, PikaLab and Gen2 also have the ability to control the motions and the cameras through additional hyper-parameters. Besides, although there are many parameters that can be adjusted, we keep the default settings for a relatively fair comparison.

We first show the overall human-aligned results in Fig. 5, with also the different aspects of our benchmark in Table 3, which gives us the final and the main metrics of our benchmark. Finally, as in Figure 7, we give the results of each method on four different meta-types (i.e., animal, human, landscape, object) in our benchmark and two different type videos (i.e., general, style) in our benchmark. For comparing the objective and subjective metrics of each method, we give the raw data of each metric in Table. 1 and Fig. 6. We give a detailed analysis in Sec. 5.1.

Finding #1: Evaluating the model using one single metric is unfair. From Table. 3, the rankings of the models vary significantly across these aspects, highlighting the importance of a multi-aspect evaluation approach for a comprehensive understanding of their performance. For instance, while Gen2 outperforms other models in terms of Visual Quality, Motion Quality, and Temporal Consistency, PikaLab demonstrates superior performance in Text-Video Alignment.

Finding #2: Evaluating the models’ abilities by meta-type is necessary. As shown in Fig. 7, most methods show very different values in different meta types. For example, although Gen2 has the best overall T2V alignment in our experiments, the generated videos from this method are hard to recognize by the action recognition models. We subjectively find Gen2 mainly generates the close-up shot from text prompt with a weaker motion amplitude.

Finding #3: Users are more tolerate with the bad T2V alignment than visual quality. As shown in Fig. 7 and Table. 2, even Gen2 can not perform well in all the text-video alignment metrics, the user still likes the results of this model in most cases due to its good temporal consistency, visual quality, and small motion amplitude.

Finding #4: All the methods CAN NOT control their camera motion directly from the text prompt. Although some additional hyper-parameters can be set as additional control handles, the text encoder of the current T2V text encoder still lacks the understanding of the reasoning behind open-world prompts, like camera motion.

Finding #5: Visually appealing has no positive correlation with the generated resolutions. As shown in Tab. 1, gen2 has the smallest resolutions, however, both humans and the objective metrics consider this method to have the best visual qualities and few artifacts as in Tab. 2, Fig. 6.

Finding #6: Larger motion amplitude does not indicate a better model for users. From Fig. 6, both two small motion models, i.e., PikaLab and Gen2 get better scores in the user’s choice than the larger motion model, i.e., Floor33 Pictures . Where users are more likely to see slight movement videos other than a video with bad and unreasonable motions.

Finding #7: Generating text from text descriptions is still hard. Although we report the OCR-Scores of these models, we find it is still too hard to generate realistic fonts from the text prompts, nearly all the methods are fair to generate high-quality and consistent texts from text prompts.

Finding #8: The current video generation model still generates the results in a single shot. All methods show a very high consistency of CLIP-Temp as in Table. 2, which means each frame has a very similar semantic across frames. So the current T2V models are more likely to generate the cinemagraphs, other than the long video with multiple transitions and actions.

Finding #9: Most valuable objective metrics. By aligning the objective metrics to the real users, we also find some valuable metrics from a single aspect. For example, SD-Score and CLIP-Score are both valuable for text-video alignment according to Table. 2 and Table. 3. VQAT and VQAA are also valuable for visual quality assessment.

Finding #10: Gen2 is not perfect also. Although Gen2 achieved the overall top performance in our evaluation, it still has multiple problems. For example, Gen2 is hard to generate video with complex scenes from prompts. Gen2 has a weird identity for both humans and animals, which is also reflected by the IS metric (hard to be identified by the network also) in Table. 1, while other methods do not have such problems.

Finding #11: A significant performance gap exists between open-source and closed-source T2V models. Referring to Table 3, we can observe that open-source models such as ModelScope-XL and ZeroScope have lower scores in almost every aspect compared to closed-source models like PikaLab and Gen2 . This indicates that there is still room for improvement in open-source T2V models to reach the performance levels of their closed-source counterparts.

2 Ablation on Human Preference Alignment

To demonstrate the effectiveness of our model in aligning with human scores, we calculate Spearman’s rank correlation coefficient and Kendall’s rank correlation coefficient , both of which are non-parametric measures of rank correlation. These coefficients provide insights into the strength and direction of the association between our method results and human scores, as listed in Table. 4. From this table, the proposed weighting method shows a better correlation on the unseen 200 samples than directly averaging (we divide all data by 100 to get them to range $$ first). Another interesting finding is that all current Motion Amplitude scores are not related to the users’ choice. We argue that humans care more about the stability of the motion than the amplitude. However, our fitting method shows a higher correlation.

3 Limitation

Although we have already made a step in evaluating the T2V generation, there are still many challenges. (i) Currently, we only collect 500 prompts as the benchmark, where the real-world situation is very complicated. More prompts will show a more detailed benchmark. (ii) Evaluating the motion quality of the general senses is also hard. However, in the era of multi-model LLM and large video foundational models, we believe better and larger video understanding models will be released and we can use them as our metrics. (iii) The labels used for alignment are collected from only 3 human annotators, which may introduce some bias in the results. To address this limitation, we plan to expand the pool of annotators and collect more diverse scores to ensure a more accurate and unbiased evaluation.

Conclusion

Discovering more abilities of the open world large generative models is essential for better model design and usage. In this paper, we make the very first step for the evaluation of the large and high-quality T2V models. To achieve this goal, we first built a detailed prompt benchmark for T2V evaluation. On the other hand, we give several objective evaluation metrics to evaluate the performance of the T2V models in terms of the video quality, the text-video alignment, the object, and the motion quality. Finally, we conduct the user study and propose a new alignment method to match the user score and the objective metrics, where we can get final scores for our evaluation. The experiments show the abilities of the proposed methods can successfully align the users’ opinions, giving the accurate evaluation metrics for the T2V methods.

References