Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising
Fu-Yun Wang, Wenshuo Chen, Guanglu Song, Han-Jia Ye, Yu Liu, Hongsheng Li
Introduction
Benefitting from pre-training on large-scale text-image datasets and the development and refinement of the diffusion model , we have witnessed a plethora of successful applications, including impressive image generation, editing, and even fine-grained generation control through the injection of additional layout information . A logical progression of this approach is its extension to the video realm for text-driven video generation and editing .
Currently, there are three primary strategies for text-driven video generation and editing:
Pretrained Text-to-Video (pretrained t2v) involves training the diffusion model on a large-scale text-video paired dataset such as WebVid-10M . Typically, a temporal interaction module, like Temporal Attention , is added to the denoising model, fostering inter-frame information interaction to ensure frame consistency.
Tuning-free Text-to-Video (tuning-free t2v) utilizes the pre-trained Text-to-Image model to generate and edit video frame-by-frame, while applying additional controls to maintain consistency across frames (for instance, copying and modifying the attention map , sparse causal attention , etc.).
One-shot tuning Text-to-Video (one-shot-tuning t2v) proposes to fine-tune the pretrained text-to-image generation model on a single video instance to generate videos with similar motions or contents. Despite the extra training cost, one-shot tuning-based methods often offer more editing flexibility compared to tuning-free based methods. As depicted in Fig. 4, both attempt to substitute the rabbit in the source video with a tiger or a puppy. The outcome produced by the tuning-free t2v method reveals elongated ears, losing authenticity. One-shot-tuning-based method, in contrast, effectively circumvents this problem.
Despite significant advances made by previous methods, they are accompanied by some fatal limitations, restricting their practical applications. First of all, the number of video frames generated by these methods is usually limited, generally less than 24 frames . On one hand, the computational complexity of temporal attention scales quadratically with the number of frames, making the direct generation of ultra-long videos infeasible. On the other hand, ensuring consistency becomes more challenging with the increase in the number of frames. Another noteworthy limitation is that these methods typically generate videos controlled by a single text condition and cannot accommodate multiple text prompts. In reality, the content of a video often changes over time, meaning that a comprehensive video often comprises multiple segments each bearing different semantic information. This necessitates the development of video generation methods that can handle multiple text conditions. Though there are already attempts at generating long videos, they typically require additional training on large-scale text-video datasets and follow the autoregressive mechanism (i.e.,the generation of later frames is conditioned on former ones), which suffers from severe content degradation and inference inefficiency (see Sec. 2 for more discussion).
In light of these challenges, we propose a novel framework aimed at generating long videos with consistent, coherent content across multiple semantic segments. Unlike previous methods, we do not construct or train a long-video generator directly. Instead, we view a video as a collection of short video clips, each possessing independent semantic information. Hence, a natural idea is that generation of long videos can be seen as the direct splicing of multiple short videos. However, this simplistic division falls short of generating long videos with consistent content, resulting in noticeable content and detail discrepancies between different video clips. As shown in the third row of Fig. 2, the color of the jeep car changes drastically among different video clips when they are denoised isolatedly. To counter this, we perceive long videos as short video clips with temporal overlapping. We demonstrate that under certain conditions, the denoising path of a long video can be approximated by joint denoising of overlapping short videos in the temporal domain. In particular, as depicted in Fig. 1, the noisy long video is initially mapped into multiple noisy short video clips via a designated function. Subsequently, existing off-the-shelf short video diffusion models can be employed to denoise these video clips under the guidance of various text conditions. These denoised short videos are then merged and inverted back to a less noisy original long video. Essentially, this procedure establishes an abstract long video generator and editor without necessitating any additional training, enabling the generation and editing of videos of any length using established short video generation and editing methodologies.
Our method was tested in three scenarios: pretrained t2v, tuning-free t2v, and one-shot-tuning t2v, all of which yielded favorable results. Furthermore, the incorporation of additional control information and advanced open-set detection and segmentation technologies allows for more impressive results, such as precise layout control and arbitrary object video inpainting. Extensive experimental results validate the broad applicability and effectiveness of our proposed Gen-L-Video.
Related Work
As we mentioned, the current mainstream strategies for video generation and editing can be mainly categorized into three types: pretrained Text-to-Video (pretrained t2v), tuning-free Text-to-Video (tuning-free t2v), and one-shot-tuning Text-to-Video (one-shot-tuning t2v). The breakthroughs in video generation and editing techniques have largely drawn from the existing technologies for image editing and generation. In this section, we first introduce the development of text-to-image technology and then provide a brief overview of the key accomplishments of each of the three strategies. In the end, we discuss recent advances in long video generations and the advantage of Gen-L-Video over them.
Many early works train GANs on image captioning datasets to produce text-conditional image samples. Other works apply vector quantization and then adopt autoregressive transformers to predict image tokens followed by text tokens. Recently, several works adopt diffusion models for Text-to-Image Generation. GLIDE introduces classifier-free guidance in the diffusion model to enhance image quality, while DALLE-2 improves text-image alignment using the CLIP feature space. Imagen employs cascaded diffusion models for high-definition video generation. VQ-diffusion and Latent Diffusion Model (LDM, also known as Stable Diffusion) train diffusion models in an autoencoder’s latent space to boost efficiency. Variants of LDM are fine-tuned to achieve more functionality like inpainting, image variants, etc. ControlNet , T2I-adapter add new modules to accept additional image inputs, achieving precise generative layout control. Many fine-tuning strategies are also developed to force diffusion models to generate new concepts and styles, which shares a similar idea to continual learning .
Pretrained Text-to-Video.
Despite significant advancements in Text-to-Image generation, generating videos from text remains a challenge due to the scarcity of high-quality, large-scale text-video datasets and the inherent complexity of modeling temporal consistency and coherence. Early works primarily focus on generating videos in simple domains, such as moving digits or specific human actions. GODIVA is the first model to utilize VQ-VAE and sparse attention for Text-to-Video generation, enabling more realistic scenes. NÜWA builds upon GODIVA with a unified representation for various generation tasks via multitask learning. CogVideo incorporates additional temporal attention modules on top of the pre-trained Text-to-Image model . Similarly, Video Diffusion Models (VDM) use a space-time factorized U-Net with joint image and video data training. Imagen Video improves VDM with cascaded diffusion models and v-prediction parameterization for high-definition video generation. Make-A-Video , MagicVideo , and LVDM share similar motivations, aiming to transfer progress from t2i to t2v generation. Video Fusion decomposes the denoising process by resolving per-frame noise into base noise and residual noise to reflect the connections among frames. For video editing, Dreamix and Gen-1 utilize the video diffusion model for video editing.
Tuning-free Text-to-Video.
It’s nontrivial to directly apply pretrained Text-to-Image model for video generation or editing without tuning. Recent diffusion-based image editing models , although powerful in processing individual frames in a video, results in inconsistencies between frames due to the models’ lack of temporal awareness. Tune-A-Video finds that extending spatial self-attention to sparse causal attention with pretrained weight produces consistent content across frames. Fate-Zero and Video-P2P apply sparse causal attention and attention control proposed in prompt2prompt, achieving consistent video editing. Pix2Video adds additional regularization to penalize the dramatic frame changes. Text2Video-Zero first proposes to generate videos in zero-shot settings with only pretrained text-to-image diffusion model. It applies the sparse causal attention and object mask to preserve the content consistency among frames and add motion dynamics in different scales to enrich the base latent code to generate consecutive motions.
One-shot tuning Text-to-Video.
Single-video GANs generate new videos with similar appearance and dynamics to the input video, while they suffer from extensive computational burden. SinFusion adapts diffusion models to single-video tasks and enables autoregressive video generation with improved motion generalization capabilities. Tune-A-Video proposes to fine-tune the pretrained text-to-image diffusion model on a video to enable generation of videos with similar motions, demonstrating powerful editing ability.
Long video generation.
The generation of long videos has garnered significant attention in recent years, resulting in various attempts to address the challenges associated with this task . Existing approaches typically rely on autoregressive models, such as NUWA-Infinity , Phenaki , and TATS , or diffusion models, including MCVD , FDM , PVDM , and LVDM . All these methods employ an autoregressive mechanism for long video generation, wherein the generated frames serve as conditioning for subsequent frames. However, this mechanism often leads to significant content degradation after several extrapolations due to error accumulation. Furthermore, the autoregressive mechanism constrains generation to a sequential process, substantially reducing efficiency. Recently, NUWA-XL proposed a novel hierarchical diffusion process that enables parallel long video generation. Despite its advantages, this approach necessitates extensive pretraining on large long video datasets and requires a well-designed global diffusion model to generate key frames at the outset. Our framework, instead of directly training or constructing a long video diffusion model, can approximate the arbitrary length long video denoising path with parallel joint denoising of off-the-shelf short video generation or editing models. In general, Gen-L-Video presents an efficient, convenient, and scalable paradigm for long video generation, addressing the limitations of existing methods and offering new possibilities for future research and applications. We make a direct comparison of our method with existing methods in Table 2, and our approach demonstrates significant superiority.
Method
Diffusion models perturb the data by gradually injecting noise to data , which is formalized by a Markov chain:
where is the noise schedule and . The data can be generated by reversing this process, i.e.,we gradually denoise to restore the original data. The diffusion model parameterized by is trained to approximate the reverse transition , which is formulated as
where is uniformly sampled from and .
DDIM generalize the framework of DDPM and propose a deterministic ODE process, achieving faster sampling speed. The inversion trick of DDIM , based on the assumption that the ODE process can be reversed in the limit of small steps, can be used to approximate the corresponding noise of the given instance:
Latent Diffusion Model (LDM) is a variant of text-to-image diffusion models. An autoencoder is first trained on large image datasets, where the encoder compresses the original image into a latent code , and the decoder reconstructs the original image from the latent code. That is
Then a conditional DDPM is trained to gradually remove noise for data sampling. Classifier-free guidance (GFC) is proposed to improve the text-image alignment by linearly combining the conditional predicted noise and the unconditional one.
where is the guidance scale. Larger typically improves the image-text alignment but overlarge causes the degradation of sample fidelity and diversity.
2 Temporal Co-Denoising
As we discussed above, current Text-to-Video diffusion methods for generation and editing typically view the video as a whole. Given a noisy video , they train a diffusion model with respect to noise prediction model to denoise it as a whole. This greatly limits the video length that they are able to generate and makes it hard for them to accommodate multi-text conditions. Though some works performed long videos generation via the autoregressive mechanism, this manner suffers from severe content degradation and only supports serialization generation, leading to inference inefficiency.
In contrast, we consider the denoising process of the entire video as multiple short videos with temporal overlapping undergoing parallel denoising in the temporal domain. We approximate the denoising trajectory of a long video through the joint denoising model of short videos in the temporal domain. More specifically, we suppose there exists a model (with a corresponding noise prediction network ) capable of denoising the given long video , resulting in the denoising trajectory,
where we use the diffusion model to gradually transform the pure noise to the clean video . can be represented as a single or multiple text prompts.
We define a set of mappings to project all original videos (both noisy and clean) in the trajectory to short video segments , specifically,
where represents the collection of frames with frame id from to , represents the stride among adjacent short video clips, is the fixed length of short videos, and is the total number of clips. Empirically, we find that setting to or yields excellent results and preserves efficiency. When setting , our method degrades into isolated denoising. The total number of frames of the video is . Each short video is guided with an independent text condition . After obtaining these short videos, we are able to denoise these short videos using the off-the-shelf short video diffusion models . For simplicity, we can set the diffusion models for all video clips to a single diffusion model , and we find that it works quite well. Then we have,
The remaining question is how to obtain the denoised after acquiring all short video clips .
Considering that we have assumed for all and , the ideal should satisfy that is as close as as possible. Therefore, the optimal can be obtained by solving the following optimization problem.
where is the pixel-wise weight for the video clip , and means the tensor product. It is not difficult to verify that for an arbitrary frame in the video , namely , it should be equal to the weighted sum of all the corresponding frames in short videos that contain the frame. We provide the proof in Sec. III.
In this way, we are able to approximate the transition function with . As we claimed, our method establishes an abstract long video generator and editor without necessitating any additional training, enabling the generation and editing of videos of any length using established short video generation and editing methodologies.
Integrate Gen-L-Video with Mainstream Paradigms
As previously mentioned, our method can be applied to three mainstream paradigms: pretrained t2v, tuning-free t2v, and one-shot-tuning t2v. In this section, we will introduce our implementation and improvements for these paradigms in long video generation and editing. Furthermore, by utilizing additional control information, we can achieve more accurate layout control. Advances in open-set detection and segmentation allow us to achieve precise editing of arbitrary objects in the video without altering the other contents (e.g.,background), resulting in a more powerful and flexible video editing process. All our implementations are based on the pretrained LDM and its variants.
For pretrained Text-to-Video generation and editing, we choose the open-sourced VideoCrafter . VideoCrafter is a Text-to-Video model fine-tuned from LDM on the large text-video dataset WebVid-10M . For modeling the dynamic relationships among frames, VideoCrafter adds additional temporal attention blocks in the original LDM. The pipeline for long video generation and editing follows our proposed temporal co-denoising as illustrated in Fig. 1.
Tuning-free Text-to-Video.
where and means the operation of sequence dimensional concatenation.
One-shot tuning Text-to-Video.
For one-shot tuning Text-to-Video, we follow the pipeline of Tune-A-Video . Similar to Tuning-free Text-to-Video, we also replace the sparse causal attention mechanism with our proposed bidirectional cross-frame attention. However, applying this pipeline directly to the training and generation of long videos presents non-trivial challenges. Although we can prompt the model to learn denoising for each short video clip, it struggles during generation. This is because many short video clips in a video share the same text description, making it difficult for the model to determine which clip it is denoising based on randomly initialized noise and similar text conditions alone. To address this, we suggest learning clip identifier for each short video clip to guide the model in denoising the corresponding clip.
However, introducing could lead to overfitting of the corresponding clip content, causing the model to overlook the text information and lose its editing ability. Drawing on the idea of CFG , we randomly drop during training. At test time, we base our approach on:
This effectively alleviates the overfitting phenomenon. We believe that video learning consists of learning content and motion information. Although learns both content and motion information for the video clip , we aim to retain only the motion information. By dropping , the model learns across all video clips, gaining a large amount of content information. Shifting the denoising direction away from the scenario when is dropped allows us to reduce overfitting to video content.
Personalized and controllable generation.
Our method can be easily extended to personalized and controllable layout generation. Users can easily combine personalized diffusion models obtained through fine-tuning strategies like DreamBooth and LoRA with our generation pipelines. Besides, we are able to inject additional layout control such as pose and segmentation maps with ControlNet and T2I-Adapter pretrained on image datasets . This allows us for more smooth and more precise video generation and editing.
Edit anything in the video.
Open-set detection and segmentation have demonstrated remarkable ability and inspired plenty of interesting applications. We show that it is possible to combine these with our method to achieve precise arbitrary object editing in long videos. Specifically, given a prompt of a specific object (e.g.,man), we apply the open-set detection model Grouding DINO to detect the corresponding object in each frame. Then, the open-set segmentation model SAM is applied to get the precise mask of the object in each frame of the video. With these masks, we are capable of applying inpainting methodologies for video editing while keeping the other components of the video content unchanged. We find that directly using an pretrained Text-to-Image diffusion model with our proposed bi-directional cross-frame attention can already yield acceptable results. To achieve more precise layout generation, we additional add a controlnet pretrained on Text-To-Image datasets to accept the SAM maps.
Experiments
All our experiments are heavily built upon the pretrained LDM (a.k.a Stable Diffusion), as we mentioned. By default, DDIM sampling strategy is applied for all our experiments and the sampling step and guidance scale is set to , and 13.5, respectively. In most cases, we set the number of frames of short video clips to and stride between adjacent short videos clips to . For the one-shot-tuning Text-to-Video pipeline, we set the basic learning rate as and scale up the learning rate as the batch size, which greatly accelerates the training. The default beach size is set to 5, and the loss typically converges within 100 epochs for videos with around 100 frames.
Benchmarks.
To evaluate our method, we collect a video dataset containing 66 videos whose lengths vary from 32 to hundreds of frames. These videos are mostly drawn from the TGVE competition and the internet. For each video, we label it a source prompt and add four edited prompts for video editing, including object change, background change, style transfer, similar motion changes, and multiple changes. We provide details about the dataset in Sec. I.
Qualitative results.
We provide a visual presentation of representatives of our generated videos in Fig. 5, including results in various lengths generated through pretrained t2v, tuning-free t2v, one-shot tuning t2v, personalized diffusion model, multi-text conditions, pose layout control, and video inpainting through the auto-detected mask, respectively. All of them show favorable results, demonstrating the strong versatility of our framework. More results can be seen in Sec. IV and our project page: https://github.com/G-U-N/Gen-L-Video.
Quantitative results.
For quantitative comparisons, we assess the video frame consistency and the textual alignment following the prior work . We compare Gen-L-Video with Isolated Denoising, where each video clip is denoised isolatedly. Regarding the video frame consistency, we employ the CLIP image encoder to extract the embeddings of individual frames and then compute the average cosine similarity between all pairs of video frames. To evaluate the textual alignment, which refers to the alignment between text and video, we calculate the CLIP score of text and each frame in the video. The average value of the scores is used to measure the alignment degree while the variance of those is used to measure the alignment stability. For human preference, we select several participants to vote on which method yields better frame consistency and textual alignment and get 1040 votes in total. The experimental results are shown in Table. 2.
Ablation study.
Bi-directional cross-frame attention. We compare our proposed Bi-directional cross-frame attention with the sparse causal attention when temporal co-denoising is applied. As illustrated in Fig. 2, our method achieves the most smooth and consistent result while sparse causal attention typically causes the first few frames incwonsistent with subsequent ones. Video clip identifier. We compare the generation results of one-shot tuning t2v with or without the clip identifier in Fig. 4 (Right). When random initial noise is used, one-shot tuning t2v without the clip identifier fails to generate consecutive content in the source video, while the other method succeeds
Conclusion
In this work, we propose Gen-L-Video, a universal methodology that extends short video diffusion models for efficient multi-text conditioned long video generation and editing. We implement mainstream Text-to-Video methods and make additional improvements to integrate them with Gen-L-Video for long video generation and editing. Experiments verify our framework is universal and scalable.
Limitations: In general, our framework should be able to be extended into the co-working of various different video diffusion models with different lengths to obtain more flexibility in generation and editing, but we haven’t experimented with that. This avenue of research is left as future work.
References
I Dataset Details
To evaluate our method, we have selected several videos (totaling 66) from TGVE competition and Internet. For internet videos, we have designed 4 distinct prompts that introduce dynamic changes in the areas of object recognition, stylistic elements, background variations, similar motion patterns, or a combination thereof, based on the original prompts.
II User Study Details
We utilized the aforementioned dataset as the benchmark to contrast our approach, Gen-L-Video, with the Isolated method. Beyond quantitative indices (frame consistency and textual alignment), we also enlisted several participants to vote on which method was superior. To assess frame consistency, we generated 264 videos from 4 prompts and an additional 65 videos using both methods, respectively. Then we replicate and shuffle them to 1040 videos. We distributed pairs of these videos to several individuals, posing the question, "Which video exhibits superior frame consistency between these two?" In order to gauge textual alignment, we created 279 videos from varying prompts using both methods separately. We subsequently asked the participants, "Which video aligns more accurately with the text description among these two videos? A?" No B nr diffeerenc?? Most participants reported no significant difference in the degree of textual alignment, but a clear improvement in alignment stability, indicating that our method maintains good text-based editing capabilities.The more detailed results of these comparisons are depicted in Table 2.
III Proof for the Optimal Approximation
As we claimed in Sec. 3.2, the ideal should satisfy that is as close as as possible. The optimal can be obtained by solving the following quadratic optimization problem:
where is the pixel-wise weight for the video clip , and means the tensor product. Here we show that, for an arbitrary frame in the video , namely , it should be equal to the weighted sum of all the corresponding frames in short videos that contain the frame.
Considering that there are video clips in total, we denote the set of all video clips and the set of their indices as and , respectively. Further, we assume that there exists a set of video clips consisting of short video clips containing the corresponding frame in the original long video. Similarly, the set of indices corresponding to is denoted as . Here we use the to represent that it could be suitable for all the time in the denoising path.
Given a short video clip , we denote the corresponding frame of in as for simplicity. Note that the in different video clips represents different values.
Therefore, the original optimization objective can be written as:
Where, is the pixel-wise weight for the frame in (i.e.,), and is the frame in . It is not difficult to observe that the last two terms in the formula have nothing to do with . We denote them as constant . Then we have,
Denote the above objective as and take the gradient of with respect to , and then we have
Note that , therefore and we can replace the in the above formula with . Then, we have
Therefore, set the gradient to be zero, and then we get the optimal ,
which is the weighted sum of all the corresponding frames in short video clips that contain the frame.
IV Additional Results
Multi-text long video. Fig. 6 illustrates an example of our multi-text conditioned long video. Specifically, we first split the original video into several short video clips with obvious content changes and of various lengths, and then we label them with different text prompts. The different colors in Fig. 6 indicate short videos with different text prompts. Then, we split the original long video into short video clips with fixed lengths and strides. For video clips only containing frames conditioned on the same prompt, we can directly set it as the condition. In contrast, for video clips containing frames conditioned on different prompts, we apply our proposed condition interpolation to get the new condition. After all of these, our paradigm Gen-L-Video can be applied to approximate the denoising path of the long video.
Pretrained Text-to-Video. Gen-L-Video can also be applied to the pretraiend short video generation model for longer video generation. We compare the results generated through our method and isolated denoising in Fig. 7, and the result reveals that our method significantly enhances the relevance between different video clips.
Controllabel video generation. We showcase the results generated by injecting additional control information in Fig. 8. The results show that our method can be easily combined with additional information to achieve precise layout control.
Edit anything. Our approach demonstrates significant compatibility with inpainting tasks. As depicted in Fig.9 and Fig.10, our method can reliably edit very long videos and maintain consistent content. The examples given in Fig. 9 are longer than 12 and 20 seconds, respectively.
Long video with smooth semantic changes. Our paradigm also allows us a pleasant application where we are able to edit the source video to generate videos with smooth semantic changes. For example, we are able to generate cars running on the road from day to night to reflect the time flies. The generated results are represented in Fig. 11.