VideoPoet: A Large Language Model for Zero-Shot Video Generation
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Josh Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Xuan Yang, Hartwig Adam, Ming-Hsuan Yang, Irfan Essa, Huisheng Wang, David A. Ross, Bryan Seybold, Lu Jiang
Introduction
Recently, there has been a surge of generative video models capable of a variety of video creation tasks. These include text-to-video , image-to-video , video-to-video stylization , and video editing among other video applications. Most existing models employ diffusion-based methods that are often considered the current top performers in video generation. These video models typically start with a pretrained image model, such as Stable Diffusion , that produces high-fidelity images for individual frames, and then fine-tune the model to improve temporal consistency across video frames.
While Large Language Models (LLMs) are commonly used as foundational models across various modalities including language , code , audio , speech , and robotics , the diffusion model remains the predominant approach for video generation. Although early research has demonstrated the effectiveness of LLMs in text-to-image generation (e.g., DALL-E , Parti and ) and text-to-video (e.g., ), language models have not reached a level of quality on par with video diffusion models in tasks like text-to-video generation as shown in previous studies . In contrast to training exclusively for text-to-video tasks, the generative model of LLMs in the language domain emphasizes a large pretraining stage to learn a foundation by examining pretraining tasks that extend beyond text-to-video generation.
A notable advantage of employing LLMs in video generation lies in the ease of integrating existing LLM frameworks. This integration allows for reusing LLM infrastructure and leverages the optimizations our community has developed over many years for LLMs, including optimizations in learning recipes for model scaling , training and inference infrastructure , hardware, among other innovations. This couples with their flexibility in encoding many diverse tasks in the same model , which stands in contrast to most diffusion models where architectural changes and adapter modules are the dominant approach used to adapt the model to more diverse tasks .
In this paper, we investigate the application of language models in video generation, following the canonical training protocols of LLMs in the language domain. We introduce VideoPoet, a language model for video generation. VideoPoet employs a decoder-only LLM architecture that admits image, video, and audio modalities as discrete tokens, each produced by their respective tokenizer.
The training of VideoPoet consists of two stages: (1) pretraining and (2) task-adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal pretraining objectives within an autoregressive transformer framework. After pretraining, the model functions as a versatile multi-task video generation model such as text-to-video, image-to-video, video editing and video-to-video stylization, as shown in VideoPoet: A Large Language Model for Zero-Shot Video Generation. Unlike , these capabilities are inherently integrated into a single LLM, rather than relying on a separate diffusion model controlled by text prompts. During subsequent task-adaptation, the pretrained model can be further fine-tuned either to enhance its generation quality on the training tasks or to perform new tasks.
Our experimental results demonstrate the VideoPoet’s state-of-the-art capabilities in generating videos with large and high-fidelity motions. In particular, we observe that through the powerful capabilities of the transformer architecture, the VideoPoet can be straightforwardly trained on a multi-task, multimodal generative objective, allowing for generating consistent and realistic motion driven by text, as shown in Figure 2, or other prompts. Additionally, VideoPoet can synthesize coherent long videos of up to 10 seconds by autoregressively extending the content, conditioned on the last second of the generated video.
We found that VideoPoet, a LLM, is capable of zero-shot video generation. We use the term “zero-shot video generation” because VideoPoet exhibits generalization capability in processing new text, image, or video inputs that diverge from the training data distribution. Furthermore, VideoPoet begins to show an ability to handle new tasks that were not included in its training. For example, VideoPoet demonstrates the ability to perform new editing tasks by sequentially chaining training tasks together. See Section 7.
We provide the following contributions in this work:
A simple method for training a Large Language Model (LLM) specifically for video generation tasks, utilizing tokenized video and audio data that seamlessly incorporates both text-paired and unpaired video data.
An approach to super-resolution that increases video resolution within the latent token space using a bidirectional transformer with efficient windowed local attention.
Evaluations and demonstrations that showcase the LLM’s competitive performance, especially in producing realistic and interesting motion.
Related Work
Most video generation works use diffusion-based methods for text-to-video and video-to-video editing . Because video diffusion models are most often derived from text-to-image diffusion models , additional tasks and modalities are added via inference tricks , architectural changes and adapter layers . Although these models are composable after training, they are not trained end-to-end in a unified model. Our multitask pretraining strategy in a single model improves the performance and provides zero-shot video generation capabilities.
In contrast, video language models are derived from the general family of transformer-based language models that easily combine multiple tasks in pretraining and demonstrate powerful zero-shot capabilities. Image generation language models can generate images autoregressively or via masked prediction . Both families have been extended to text-to-video using paired data. Other text-to-video work with transformers only leverages video-text pairs for training, but we also leverage unpaired videos (without text) and the same video for different tasks. Because video language models can flexibly incorporate many tasks , including video-to-video, we extend this family of work to text- and multimodal-conditioned tasks in this work with a synergistic pretraining strategy across many tasks.
Because language models can easily incorporate multiple training tasks, task selection is an important area of research. GPT-3 and PaLM show that training LLMs on diverse tasks leads to positive scaling effects on zero- and few-shot tasks. Other works show that masking approaches are a valuable learning target . And as model sizes grow, training data must grow as well . Our pretraining strategy enables using the same video for multiple training tasks even without paired text. This design facilitates training on a substantial quantity of video-only examples, thereby decreasing the demand for video-text pairs.
Model Overview
We are interested in researching an effective method for leveraging large language models for video generation. Our model consists of three main components: (1) modality-specific tokenizers, (2) a language model backbone (Figure 3), and (3) a super-resolution module (Figure 4).
The tokenizers map input data – i.e. image pixels, video frames, and audio waveforms – into discrete tokens in a unified vocabulary. The visual and audio tokens are flattened into a sequence of integers using raster scan ordering. The LLM accepts image, video and audio tokens as input along with text embeddings, and is responsible for generative multi-task and multimodal modeling. As illustrated in Figure 3, VideoPoet conditions on text embeddings, visual tokens, and audio tokens, and autoregressively predicts visual and audio tokens. Subsequently, the super-resolution module increases the resolution of the video outputs while refining visual details for higher quality. In the following, we discuss design specifics which allow our LLM to generate across video and audio modalities.
We employ the MAGVIT-v2 tokenizer for joint image and video tokenization, and the SoundStream tokenizer for audio. These visual and audio tokens are represented in a unified vocabulary. The unified vocabulary is constructed as follow: the initial 256 codes are reserved for special tokens and task prompts. Subsequently, the next 262,144 codes are allocated for image and video tokenization. This is followed by 4,096 audio codes. The text modality is represented by text embeddings for its better performance compared with training with text tokens from scratch.
Visual tokenizer is key to generating high-quality video content . After exploring various existing tokenizers , we observe superior performance with the MAGVIT-v2 tokenizer. In particular, it represents both images and videos as a sequence of discrete tokens in a unified large vocabulary.
Specifically, a video clip is encoded and quantized into a sequence of integers, with a decoder mapping them back to the pixel space. As the bridge between the token and pixel spaces, the performance of this visual tokenizer sets the upper bound of the video generation quality. Meanwhile, the compression ratio determines the sequence length of the LLM for effective and efficient task setups.
MAGVIT-v2 tokenizes 17-frame 2.125-second 128128 resolution videos sampled at 8 fps to produce a latent shape of , which is then flattened into 1280 tokens, with a vocabulary size of . To facilitate the generation of short-form content for mobile, we also tokenize videos into a portrait aspect ratio at 128224 resolution, producing a latent shape of , or 2240 tokens. When the evaluation protocol is on 16 frames, we discard the generated last frame to make a 16-frame video.
The MAGVIT-v2 tokenizer enforces causal temporal dependency, where the frames are encoded without any information from future frames. This causal property simplifies the setup for frame prediction tasks and supports tokenization and generation of arbitrarily long videos. To jointly represent images and videos, we encode the first frame into (1,16,16) tokens, which can be used to represent a static image as well. And then, every 4-frame chunks are encoded into (1,16,16) tokens. These tokens are concatenated on the first (temporal) dimension. For masked objectives, we adopt the COMMIT encoding as input to the tokenizer to optimally setup other tasks such as inpainting and outpainting. In simple terms, COMMIT encoding processes the input condition video and the target video differently to avoid information leakage during tokenization. The former involves tokenization of the conditional video with pixel masks applied, while the latter uses tokenization on the entire unmasked video.
Since the first frame is tokenized separately, MAGVIT-v2 allows images to be represented in the same vocabulary as video. In addition to being more compact, images provide many learnable characteristics that are not typically represented in videos, such as strong visual styles (e.g., art paintings), objects which are infrequently seen in video, rich captions, and significantly more text-image paired training data. When training on images, we resize the images to 128128 which are then tokenized to a latent shape of , or 256 tokens.
We scale the MAGVIT-v2 model’s size and train it on the datasets discussed in Section 5.1. The training follows two steps: image training, inflation and video training.
We tokenize audio clips with the pretrained SoundStream tokenizer. We embed 2.125 seconds of audio to produce 106 latent frames at a residual vector quantizer (RVQ) of four levels. To improve audio generation performance, we transpose the clip before flattening so that the model predicts the full audio clip at each RVQ granularity level before moving on to the finer grained levels. Finally, each RVQ level has a disjoint vocabulary with each level containing 1,024 codes. This results in a combined audio vocabulary size of 4,096 codes.
We find that a strong text encoding is important for accurate and high quality text-to-video generation. Pretrained text representations in general outperformed training our model with text tokens from scratch. Due to computational constraints, we found it more efficient to leverage off-the-shelf pretrained language embeddings at our model scale so the model can allocate more capacity to generating and understanding vision and audio modalities.
Therefore, instead of inputting text tokens into the model directly, we first input the tokens into a frozen pretrained T5 XL encoder to produce a sequence of text embeddings. For tasks with text guidance, such as text-to-video, T5 XL embeddings are projected into the transformer’s embedding space with a linear layer. We use up to a maximum of 64 text tokens for all of our experiments.
2 Language Model Backbone
Now that we have tokenized all modalities into discrete tokens, we can directly leverage a language model to generate videos and audios in the token space. We use a prefix language model with a decoder-only architecture as the backbone. By constructing different patterns of input tokens to output tokens during training, we can control the types of tasks the model is able to perform as explained in Section 4. As discussed in Section 3.1, we use a shared multimodal vocabulary to represent the generation of all modalities as a language modeling problem. This produces a total vocabulary size of approximately 300,000.
3 Super-Resolution
Generating high-resolution (HR) videos with an autoregressive transformer incurs heavy computational cost due to the increase in sequence length. To illustrate this with an example, the video tokenizer of Section 3.1 operating on a video produces a sequence of tokens, making autoregressive sampling highly impractical.
Aiming at efficient and high-quality generative video upsampling, we develop a custom spatial super-resolution (SR) non-autoregressive video transformer to operate in token space on top of the language model output. To mitigate the computational requirements of the very long sequences involved, and in particular the quadratic memory of the self-attention layers, our design incorporates windowed local attention . More precisely, our SR transformer is composed of blocks of three transformer layers, each of which performs self-attention in a local window aligned with one of three axis : spatial vertical, spatial horizontal and temporal. The cross-attention layers attend to the low-resolution (LR) token sequence and are also divided into local windows, isomorphic to those of the self-attention layers. All blocks also include cross-attention to text embeddings from a frozen T5 XL encoder. See Figure 4 for a schematic representation of the custom transformer architecture.
We train the SR transformer with the MAGVIT objective, and use token factorization to account for the large vocabulary size. For training, the LR token sequences are obtained by tokenizing bicubic-downsampled versions of the ground truth videos and applying noise augmentation in the discrete latent space. Specifically, we randomly resample the value of a random subset of the LR tokens and independently drop the LR condition and text embeddings for 10% of the training samples. During inference, we use non-autoregressive sampling with classifier-free guidance independently on both the LR condition and the text embeddings. We use a cascade of two stages to generate videos of resolution from the samples of VideoPoet. We refer to the appendix for further details on the implementation.
LLM Pretraining for Generation
VideoPoet demonstrates general-purpose generative video modeling by training with a large mixture of multimodal objectives. The objectives work together so that individual tasks can be chained (see Section 7.1), demonstrating a zero-shot capability that goes beyond any individual task.
We design a mixture of tasks used in pretraining to produce a foundation model capable of general purpose video generation. For each task we define a prefix input and output such that the model conditions on the prefix, and we only apply the loss on the output.
Unconditioned video generation: generate video frames without conditioning on an input.
Text-to-video: generate video frames from a text prompt.
Video future prediction: given an input video of variable length, predict future frames.
Image-to-video: given the first frame of a video as an input image, predict the future video frames.
Video inpainting/outpainting: given a masked video, predict the video with the masked contents filled in.
Video stylization: given a text prompt, optical flow, depth, and optionally the first frame from a video, predict the video frames (see Section 4.1).
Audio-to-video: given an input audio waveform, predict the corresponding video.
Video-to-audio: given an input video, predict the corresponding audio waveform.
Below we discuss design decisions within the task prompt design.
Producing a high quality initial frame is crucial for the model to generate good video examples, as motion is anchored on the appearance of the first frame. The causal dependencies of frames in the MAGVIT V2 tokenizer allow us to represent any image as if it were the first frame of a video using the same vocabulary. This design enables us to leverage text-image and text-video paired data in joint training, where the image-text data is orders of magnitude larger.
In the input sequence, we leave out the end-of-sequence token (
For all examples, we apply several variants. We have two resolutions: 128128 and 128224. We also generate on two video lengths: 17 frames (2.125 seconds) and 41 frames (5.125 seconds), both at 8 frames per second. We combine the two resolutions and two video lengths, which leads to a total of 4 combinations. Images are a special case of a 1-frame video, which we tokenize at 128128 resolution. To be able to switch between different resolutions and duration, we use special conditioning tokens that indicate what format of video should be generated.
To perform video stylization, we follow an approach inspired by to predict videos from the combination of text, optical flow, and depth signals. On a subset of steps, we also condition on the first video frame. As described in , the text will generally define the “content” or appearance of the output and the optical flow and depth control the “structure.” In contrast to the diffusion-based approaches that usually use external cross-attention networks or latent blending for stylization, our approach is more closely related to machine translation using large language models in that we only need to provide the structure and text as a prefix to a language model.
To perform the task, we estimate optical flow from RAFT and produce monocular depth maps from MIDAS , and then normalize and concatenate on the channel dimension. This conveniently produces the same number of channels as the RGB ground truth and so can be tokenized in the same fashion as RGB videos with the MAGVIT-v2 tokenizer without retraining the tokenizer. The task of stylization is to reconstruct the ground truth video from the given optical flow, depth, and text information. During inference, we apply optical flow and depth estimation on an input video but then vary the text prompt to generate a new style, e.g. “cartoon”.
In Figure 3 we illustrate a typical input-output sequence layout. For each task, the input sequence may include three types of input tokens:
text tokens (embeddings): the pre-extracted T5 embeddings for any text.
visual tokens: the MAGVIT-v2 tokens representing the images, video subsection, or COMMIT encoded video-to-video task.
audio tokens: the SoundStream tokens representing audio.
Likewise, the model outputs two types of tokens: visual tokens and audio tokens. In addition to video and audio tokens along with text embeddings, we incorporate additional special tokens enumerated as shown in Table 1.
When a modality is not included in a task, such as text and audio for unconditioned video generation, then the corresponding input or output tokens together with the beginning and end special tokens are omitted from the sequence to reduce the sequence length. To indicate the type of task, we condition on the
The video-to-video tasks use the COMMIT encoding to obtain the tokens for the tasks such as inpainting and outpainting. Text is encoded as T5 XL embeddings and are inserted into reserved sequence positions right after the
2 Training Strategy
We train on image-text pairs and video with or without text or audio. Both text and sound are noisy and may not match the visual content. The model is trained on approximately 2 trillion tokens across all modalities.
For multi-task training, we employ accelerated Alternating Gradient Descent (AGD) as formulated in to efficiently train on variable sequence lengths. While packing is an alternative, AGD results in a near 0% padding ratio, providing optimal per-token loss efficiency . We accomplish this by grouping each task by sequence length, and alternately sample one group at each iteration. Because sequence lengths are fixed per task, we can optimally train without any padding. Due to images requiring fewer tokens, we can include roughly 5 more images per batch than videos, i.e. 256 image tokens vs. 1280 video tokens.
We find that sampling from image and video datasets uniformly across time can lead to suboptimal results, as training on images can enhance the model’s understanding of objects, but does not capture any motions that are represented in video data. As a result, we devise a two-stage pretraining strategy, where we augment our sampling weights to sample from the image data 90% of the time and 10% video for the first 25% iterations of training. We then switch to training on video 90% and image 10% for the rest of training iterations.
After pretraining, we can fine-tune the pretrained model to either improve its performance on specific tasks or to enable it to undertake new tasks, i.e., task adaption. For example, we finetune the model on both text-to-video and image-to-video tasks using a high-quality data subset. We observe improved generation quality which aligns with findings from . Furthermore, we note that the fine-tuned model mitigates the issue of decoding collapse, characterized by the degradation of predictions into repetitive tokens. This improvement not only enhances the model’s output diversity but also enables increasing the Classifier-Free Guidance scale, leading to an overall enhancement in quality. In addition, we also finetune the pretrained model to perform video-to-audio generation.
Experiments
As discussed in Section 4, we train the model on a mixture of text-to-image, text-to-video, image-to-video, and video-to-video tasks—including outpainting, inpainting, stylization, and future frame prediction—as well as video-to-audio, audio-to-video, unconditioned image, and unconditioned video generation. We finetune a model on two tasks—text-to-image and text-to-video—for text-to-video evaluations, and video-to-audio for producing some video examples with matching audio. Unless explicitly stated, we do not finetune on specific tasks in the evaluation.
We train on a total of 1B image-text pairs and 270M videos (over 100M of which include paired text) from the public internet and other sources. The data has been filtered to remove egregious content and sampled to improve contextual and demographic diversity.
This paper employs a zero-shot generation evaluation protocol, as the model has not been trained on the training data distributions of target benchmarks. Specifically, the evaluation benchmark includes two text-to-video generation tasks on MSR-VTT and UCF-101 , as well as the frame prediction task on Kinetics 600 (K600) , in which the first 5 frames are provided as condition to predict the next 11 frames. It also includes inpainting and outpainting tasks on Something-Something V2 (SSv2) . Additionally, we assess stylization tasks using a subset of the DAVIS datasethttps://davischallenge.org/ , as further detailed below.
We employ commonly used metrics such as FVD , CLIP similarity score , and Inception Score (IS) for evaluation. It is important to note that the specific metrics and evaluation methods vary across different datasets. Detailed information on these variations can be found in Section A.1.
2 Pretraining Task Analysis
We investigate the learning capabilities of different combinations of pretraining tasks using a model with 300 million parameters. All task combinations are trained using a learning rate of for the same number of steps (300k) with a batch size of 1024.
For the pretraining tasks, we consider text-to-video (T2V), text-to-image (T2I), and four self-supervised learning (SSL) tasks: frame prediction (FP), central inpainting and central outpainting (Painting) and audio-video continuation (AVCont) where the model is provided with the first frame and its corresponding audio to predict the subsequent 16 frames and their matching audio. We use uniform sampling among the selected tasks, so when fewer tasks are selected, they are trained more extensively. For each video task, we randomly select 20% from a training subset of 50 million videos. Regarding the text-to-image task, we randomly sample 50 million text-image pairs from our training dataset. For tasks involving audios, our sampling is exclusive to videos that contain an audio track.
The comparison results are presented in Table 2. We evaluate a model across the four tasks within the zero-shot evaluation benchmark: the text-to-video (T2V) task on MSR-VTT and UCF 101 , frame prediction (FP) on Kinetics 600 (K600) , as well as central inpainting and outpainting on Something-Something V2 (SSv2) . In this experiment, we employ a single model, without task-adaption or finetuning, to perform all the tasks. The evaluation on K600 and SSv2 are is detailed in Section A.2. And the evaluation on the text-to-video task will be discussed in Section 5.4.1. The model is not trained on the data distributions of these evaluation datasets, and thus it is zero-shot evaluation.
The top rows of Table 2 depict each pretraining task configuration of the 300 million parameter model, which are comparable in their setup. Note that these evaluation benchmarks come from distinct visual domains and all 300M models are trained on a smaller training subset. This makes identifying consistent patterns in the results difficult. Nevertheless, we observe that incorporating all pretraining tasks results in the best overall performance, on average, across all evaluated tasks. Additionally, the significant disparity observed in the “SSL” row suggests the limitations of self-supervised training and underscores the necessity for text-paired data during training.
The last row, “Ours (8B)”, represents a model with 8 billion parameters, trained on the pretraining tasks as discussed in Section 3 and utilizing significantly more compute resources.
3 Model Scaling
To study model scaling, this experiment uses a subset of the training set without text-paired data and slightly different task prompt design. We evaluate the video generation quality using Fréchet Video Distance (FVD) and audio generation quality using the Fréchet Audio Distance (FAD), which employs the VGGish model as the embedding function . Both FVD and FAD metrics in these model scaling experiments are computed on a subset of our videos with audio, which were held out from training.
Figure 5 shows that as the model size grows and the amount of training data increases, performance improves across visual and audiovisual tasks.
After obtaining the above results, we retrain our 1B and 8B models using the task design and text-paired training data discussed in Section 3. We include a qualitative comparison of our 1B and 8B pretrained models in Section A.4. Increasing the model size improved temporal consistency, prompt fidelity, and motion dynamics while adding capabilities for limited text rendering, spatial understanding, and counting.
4 Comparison to State-of-the-Art
In Table 3, we conduct zero-shot text-to-video evaluation on the common MSR-VTT and UCF-101 datasets. We measure CLIP similarity scores following an implementation given by Villegas et al. , FVD following for UCF101 and following for MSR-VTT, and Inception Score (IS) . Our model shows highly competitive CLIP similarity score and FVD performance on MSR-VTT and UCF-101. According to Table 3, our pretrained foundation model already achieves competitive performance on all metrics. After finetuned on high-quality subset of text-video pairs, VideoPoet achieves even better CLIPSIM on MSR-VTT. For more details on the evaluation settings, see Section A.1.
4.2 Human Evaluations with Text-to-Video
We analyze VideoPoet using human raters on two tasks: text-to-video and video stylization. For text-to-video we compare VideoPoet with other recently published models, specifically: Show-1 , VideoCrafter and Phenaki . Show-1 and VideoCrafter are state-of-the-art publicly available video diffusion models while Phenaki is a token-based approach using the masked token modeling .
We first developed a unified evaluation prompt bank consisting of 200 selected prompts covering a variety of categories and styles. A large subset of our prompts are sourced from published prompt sets (including, e.g., Show-1, Video LDM ) and the remaining have been manually created for this project. We selected the prompts prior to generating videos and fixed these choices after initial selection. We also selected preferentially for prompts that contain an explicit mention of motion so that the evaluation would not be biased for models that generate high quality videos that are almost still (e.g., “person jumping off of a chair” over “person standing on a chair”). The finetuned model discussed in Section 4.2 was used for the user study.
We then compare each model against VideoPoet in a side-by-side fashion, for each prompt, showing raters videos generated by two models at a time (in randomized order so as to not bias raters). Not all methods generate videos at the same size or aspect ratio, so we resize each video to a fixed area while maintaining its original aspect ratio. Raters are then asked to compare the videos along 5 dimensions and report whether the two videos are similar or one is better than the other. Specifically, we ask raters to consider (1) text fidelity (which video follows the text prompt most faithfully), (2) video quality, (3) motion “interestingness”, (4) motion realism and (5) temporal consistency. Raters are required to undergo a training consisting of going over a collection of “training examples” for each of these five dimensions.
Our findings are summarized in Figure 6, where green, gray, and pink bars represent the proportion of trials where VideoPoet was preferred over an alternative, similar to, or less preferred to an alternative, respectively. Experiments where the green segment is larger than the pink segment mean that VideoPoet is preferred over the alternative on average. Our results show that VideoPoet in general outperforms all baseline models along almost all of the dimensions (text fidelity, quality, motion interestingness and realism) and achieves its most significant wins along the motion categories.
On temporal consistency, VideoPoet shows performance on-par with Phenaki and VideoCrafter but slightly underperforms the Show-1 model. We believe this is due to an inherent trade-off with motion interestingness, i.e., a static scene is more temporally consistent but is less interesting. More interesting larger motions necessitate more possibilities of producing noticable artifacts vs. safer small motions.
4.3 Video Stylization
To evaluate stylization capabilities, we choose 20 videos from the public DAVIS 2016DAVIS license: https://creativecommons.org/licenses/by-nc/4.0/deed.en dataset and provide 2 style prompts for each video. For more details, please refer to Section A.5. Following , we evaluated the CLIP-embedding consistency between each frame and the text prompt to determine if the stylization results matches the text. As shown in Table 4, VideoPoet outperforms Control-A-Video conditioned on depth by a large margin. We also conduct human evaluations as discussed above comparing with Control-A-Video . Human raters consistently prefer our text fidelity and video quality as shown in Figure 7.
Responsible AI and Fairness Analysis
We evaluate whether the generated outputs of our model are fair regarding protected attributes such as (1) Perceived Age (2) Perceived Gender Expression (3) Perceived Skin Tone. We construct 306 prompts with template — “a {profession or people descriptor} looking {adverb} at the camera” with “profession” being crawled from the US Bureau of Labor and Statistics and “people descriptors” including emotion state, socioeconomic class, etc. The “adverb” is used to generate semantically unchanged prompt templates such as “straightly” or “directly”. We generate 8 videos for each prompt and for each generated video we infer an approximation of the expressed attribute regarding the 3 protected attributes. Across 10 prompts that have the same semantic meaning but different “adverbs”, we observe our outputs generally introduced a stronger distribution shift toward “Young Adults” (age 18-35), “Male” and “Light Skin Tone”. However, we observe changing the “adverb” in the prompt template can significantly alter the output distributions. Therefore, our model can be prompted to produce outputs with non-uniform distributions across these groups, but also possess the ability of being prompted to enhance uniformity, though prompts are semantically unchanged. While research has been conducted in the image generation and recognition domain , this finding highlights the importance of continued research to develop strategies to mitigate issues and improve fairness for video generation.
LLM’s Capabilities in Video Generation
In this section we highlight several notable capabilities we discover from the pretrained VideoPoet, shedding light on the Large Language Models (LLMs)’s promising potential in video generation.
A simple example of zero-shot editing is inpainting with text control as in Figure 7, but our model can do even more by chaining multiple capabilities. Because of our multi-task pretraining strategy, our model exhibits task generalization that can be chained together to perform novel tasks. We show an example in Figure 7 that we can apply image-to-video to animate images, and then stylize those images with video-to-video effects. We also show applying video-to-video outpainting followed by video-to-video stylization in Figure 7. On our project websitehttp://sites.research.google/videopoet/, we also show text-to-audiovisual-output by generating video from text followed by video-to-audio tasks. At each stage, the quality of the output seems to be sufficient to remain in-distribution (i.e. teacher forcing) for the next stage without noticeable artifacts.
We hypothesize that these capabilities are attributable to our multimodal task design within a LLM transformer framework that allows for modeling multimodal content using a single transformer architecture over a unified vocabulary. Our approach contrasts with others, such as diffusion models, which typically solve these tasks by adopting multiple individually tuned adapter models to control the diffusion process .
2 Coherent Long Video Generation and Image-to-Video
A benefit of an decoder-based language model is that it pairs well with autoregressively extending generation in time. We present two different variants of this capability: generating longer videos, and converting images to videos.
Because the MAGVIT-v2 tokenizer that we use encodes the first frame independently of the subsequent frames, we can encode an image without any padding as the first frame of a video. We can then predict the remaining tokens for subsequent frames to produce a video from any image as shown in Figure 7.For image-to-video examples we source images from Wikimedia Commons: https://commons.wikimedia.org/wiki/Main_Page
We observe temporally coherent generations of objects in a video scene with dynamic, and meaningful motion (see Figure 7). To predict the future frames, despite the model only being able to view up to a short temporal context, such as the first frame or the first second of video, the model is able to keep the motion, style, and identity of objects consistent across more than a total of 8 seconds of video output.
3 3D Structure, Camera Motion, Visual Styles
Because our training spans videos, images, and text, we can prompt our model to demonstrate many aspects of understanding about the world including 3D structures, camera motions, and visual styles learned from these different sources. Even though we do not specifically add training data or losses to encourage 3D consistency, our model can rotate around objects and predict reasonable visualizations of the backside of objects. Additionally, with only a small proportion of input videos with text describing camera motion, our model can use short text prompts to apply a range of camera motions to image-to-video and text-to-video generations (see Figure 7), which has been noted to be difficult for many state-of-the-art video generation models .
In addition, these controls can be added on top of a wide range of styles, such as watercolor or oil paintings. These stylization training sources are primarily observed in the text-image training data. The ability to generalize across and combine these different types of styles to produce large motions following text prompts underscores the strength of our model’s understanding of objects in a temporal context.
Conclusion
VideoPoet highlights the potential of a large language model that is trained on discrete visual and audio tokens, in generating of videos of compelling, state-of-the-art quality. A particular strength of our model lies in its ability to generate high-fidelity, large, and complex motions. Our large language model formulation benefits from training across a variety of multimodal tasks with a unified architecture and vocabulary. Consequently, the pretrained model is adept at multi-task video creation, and serves as a foundation for a diverse variety of video related capabilities, including multiple forms of editing.
Acknowledgements
We give special thanks to Alex Siegman and Victor Gomes for managing computing resources. We also give thanks to Aren Jansen, Marco Tagliasacchi, Neil Zeghidour, John Hershey for audio tokenization and processing, Angad Singh for storyboarding in “Rookie the Raccoon”, Cordelia Schmid for research discussions, Alonso Martinez for graphic design, David Salesin, Tomas Izo, and Rahul Sukthankar for their support, and Jay Yagnik as architect of the initial concept.
References
Appendix A Appendix
We report the details of our zero-shot text-to-video settings here. We note that some details are missing in previous papers and different papers use different settings. Hence, we provide all the details and hope this evaluation setting can serve as a standard text-to-video generation benchmark. Our results are reported on the 8B model and we adopt classifier-free guidance .
All metrics are evaluated on generated videos containing 16 frames with a resolution of 256 x 256. We first generate videos of 128 x 128 resolution and then resize to 256 x 256 via bicubic upsampling.
For CLIP score, we used all 59,794 captions from the MSR-VTT test set. We use CLIP ViT-B/16 model following Phenaki . We note that some papers use other CLIP models, e.g., VideoLDM uses ViT-B/32. Our CLIP score evaluated on the ViT-B/32 backbone for MSR-VTT is 30.01. For the FVD metric, to evaluate on a wide range of captions as well as to be comparable with previous papers that evaluate on 2,048 videos, we evaluate on the first 40,960 captions in the MSR-VTT test set. More specifically, we report the FVD metrics on 2048 videos with 20 repeats. The FVD real features are extracted from 2,048 videos sampled from the MSR-VTT test set. We sample the central 16 frames of each real video, without any temporal downsampling, i.e., we use the original fps in the MSR-VTT dataset (30 fps as reported in ). The FVD is evaluated with an I3D model trained on Kinetics-400.
Following VDM , we sample 10,000 videos from the UCF-101 test set and use their categories as the text prompts to generate 10,000 videos. We use the class text prompts provided in PYoCo to represent the 101 categories. To compute the FVD real features, we sample 10K videos from the training set, following TGAN2 . We sample the central 16 frames for each real video , without any temporal downsampling, i.e., we use the original fps in the UCF-101 dataset (25 fps as reported in ). The FVD metric is evaluated with an I3D model trained on Kinetics-400 and the IS metric is evaluated with a C3D model trained on UCF-101.
A.2 Self-Supervised Tasks Evaluation Settings
Self-supervised learning tasks include frame prediction on K600 with 5 frames as condition, as well as inpainting and outpainting on SSv2. FVD is used as the primary metric, calculated with 16 frames at 128128 resolution. We follow MAGVIT in evaluating these tasks against the respective real distribution, using 500004 samples for K600 and 50000 samples for SSv2.
A.3 Super-resolution implementation details
We use a 1B model for the first spatial super-resolution stage and a 500M model for the second stage. The first super-resolution stage models videos of pixels with a token sequence of shape . The second stage models videos of pixels with a token sequence of shape . The token sequences are obtained with the same MAGVIT-v2 tokenizer used for the base language model. The custom super-resolution transformer has local self-attention windows for vertical, horizontal and temporal layers of shape in the first stage and in the second stage, respectively (Figure 4). The cross-attention layers attend to local windows in the low-resolution sequence isomorphic to self-attention windows but with half the spatial size.
We train the super-resolution stages on a dataset of 64M high-quality text-video pairs using the masked modeling objective of MAGVIT , with token factorization into groups . During inference, we use the sampling algorithm of MAGVIT-v2 with 24 sampling steps for each stage and classifier-free guidance scale of for the text condition and for the low-resolution condition, in the first/second stage.
A.4 Comparison of 1B and 8B models
In Figure 7, we show outputs of 1B and 8B parameter models on the same prompts. Four frames from the best video output of each model in a batch of four text-to-video samples were selected to represent the model. In the first row, the 1B model is unstable with large changes to the subject over time and misses elements from the complex prompt. This prompt was originally used for scaling comparisons in , and compared to a dedicated image-only model, our model does not preserve text as well given the training data used. In the second row, we use a simpler text task and show that the 8B model can represent a single letter clearly, but the 1B model still produces artifacts. In the third row, we show that the 8B model learns spatial positioning such that the river is in front of the astronaut and horse. In the fourth row, we show that the 8B parameter model learned a stop motion style to have items disappear “one by one” and can follow a complicated layout from a long prompt. In contrast, the 1B model includes all of the nouns, but is unstable over time and does not follow the layout indicated in the prompt. In the bottom row, we show that the 8B model understands counts of objects in that it displays a full bouquet (though 12 roses are not explicitly in frame) and smooth consistent motion as opposed to the 1B model 5 roses and distorting objects produced by the 1B model. Overall, scaling the model improved temporal consistency, prompt fidelity, and motion dynamics while adding capabilities for limited text rendering, spatial understanding, and counting.
A.5 Stylization Evaluation on DAVIS
To evaluate the CLIP similarity score and human preference on video stylization, we use the following set of videos and prompts. We select 20 videos from DAVIS 2016 , and for each video we take 16 frames starting from the initial frame specified below and evaluate stylization on the two text prompts specified below. To be easily reproducible, we use a central square crop at the height of the video and evaluate the output videos at 256x256 resolution. We use CLIP-B/16 for the similarity score. Several prompts below are used in or inspired by previous work .