A Recipe for Scaling up Text-to-Video Generation with Text-free Videos

Xiang Wang, Shiwei Zhang, Hangjie Yuan, Zhiwu Qing, Biao Gong, Yingya Zhang, Yujun Shen, Changxin Gao, Nong Sang

Introduction

Video generation aims to synthesize realistic videos that possess visually appealing spatial contents and temporally coherent motions. It has witnessed unprecedented progress in recent years with the advent of deep generative techniques , especially with the emergence of video diffusion models . Pioneering approaches utilize pure image diffusion models or fine-tuning on a small amount of video-text data to synthesize videos, leading to temporally discontinuous results due to insufficient motion perception . To achieve plausible results, current text-to-video methods like VideoLDM and ModelScopeT2V usually insert temporal blocks into latent 2D-UNet and train the model on expansive video-text datasets, e.g., WebVid10M . To enable more controllable generation, VideoComposer proposes a compositional paradigm that incorporates additional conditions (e.g., depth, sketch, motion vectors, etc.) to guide synthesis, allowing customizable creation.

Despite this, the progress in text-to-video generation still falls behind text-to-image generation . One of the key reasons is the limited scale of publicly available video-text data, considering the high cost of video captioning . Instead, it could be far easier to collect text-free video clips from media platforms like YouTube. There are some works sharing similar inspiration, Make-A-Video and Gen-1 employ a two-step strategy that first leverages a large (∼\sim1B parameters) diffusion prior model to convert text embedding into image embedding of CLIP and then enters it into an image-conditioned generator to synthesize videos. However, the separate two-step manner may cause issues such as error accumulation , increased model size and latency , and does not support text-conditional optimization if extra video-text data is available, leading to sub-optimal results. Moreover, the characteristics of scaling potential on video generation are still under-explored.

In this work, we aim to train a single unified video diffusion model that allows text-guided video generation by exploiting the widely accessible text-free videos and explore its scaling trend. To achieve this, we present a novel two-branch framework named TF-T2V, where a content branch is designed for spatial appearance generation, and a motion branch specializes in temporal dynamics synthesis. More specifically, we utilize the publicly available image-text datasets such as LAION-5B to learn text-guided and image-guided spatial appearance generation. In the motion branch, we harness the video-only data to conduct image-conditioned video synthesis, allowing the temporal modules to learn intricate motion patterns without relying on textual annotations. Paired video-text data, if available, can also be incorporated into co-optimization. Furthermore, unlike previous methods that impose training loss on each frame individually, we introduce a temporal coherence loss to explicitly enforce the learning of correlations between adjacent frames, enhancing the continuity of generated videos. In this way, the proposed TF-T2V achieves text-to-video generation by assembling contents and motions with a unified model, overcoming the high cost of video captioning and eliminating the need for complex cascading steps.

Notably, TF-T2V is a plug-and-play paradigm, which can be integrated into existing text-to-video generation and compositional video synthesis frameworks as shown in A Recipe for Scaling up Text-to-Video Generation with Text-free Videos. Different from most prior works that rely heavily on video-text data and train models on the widely-used watermarked and low-resolution (around 360P) WebVid10M , TF-T2V opens up new possibilities for optimizing with text-free videos or partially paired video-text data, making it more scalable and versatile in widespread scenarios, such as high-definition video generation. To study the scaling trend, we double the scale of the training set with some randomly collected text-free videos and are encouraged to observe the performance improvement, with FID from 9.67 to 8.19 and FVD from 484 to 441. Extensive quantitative and qualitative experiments collectively demonstrate the effectiveness and scaling potential of the proposed TF-T2V in terms of synthetic continuity, fidelity, and controllability.

Related Work

In this section, we provide a brief review of relevant literature on text-to-image generation, text-to-video generation, and compositional video synthesis.

Text-to-image generation. Recently, text-to-image generation has made significant strides with the development of large-scale image-text datasets such as LAION-5B , allowing users to create high-resolution and photorealistic images that accurately depict the given natural language descriptions. Previous methods primarily focus on synthesizing images by adopting generative adversarial networks (GANs) to estimate training sample distributions. Distinguished by the promising stability and scalability, diffusion-based generation models have attracted increasing attention . Diffusion models utilize iterative steps to gradually refine the generated image, resulting in improved quality and realism. Typically, Imagen and GLIDE explore text-conditional diffusion models and boost sample quality by applying classifier-free guidance . DALL⋅\cdotE 2 first leverages an image prior to bridge multi-modal embedding spaces and then learns a diffusion decoder to synthesize images in the pixel space. Stable Diffusion introduces latent diffusion models that conduct iterative denoising processes at the latent level to save computational costs. There are also some works that generate customized and desirable images by incorporating additional spatial control signals .

Text-to-video generation. This task poses additional challenges compared to text-to-image generation due to the temporal dynamics involved in videos. Various early techniques have been proposed to tackle this problem, such as recurrent neural networks combined with GANs or transformer-based autoregressive models . With the subsequent advent of video diffusion models pretrained on large-scale video-text datasets , video content creation has demonstrated remarkable advances . Imagen Video learns cascaded pixel-level diffusion models to produce high-resolution videos. Following , Make-A-Video introduces a two-step strategy that first maps the input text to image embedding by a large (∼\sim1B parameters) diffusion prior model and then embeds the resulting embedding into an image-conditional video diffusion model to synthesize videos in pixel space. VideoLDM and ModelScopeT2V extend 2D-UNet into 3D-UNet by injecting temporal layers and operate a latent denoising process to save computational resources. In this paper, we present a single unified framework for text-to-video generation and study the scaling trend by harnessing widely accessible text-free videos.

Compositional video synthesis. Traditional text-to-video methods solely rely on textual descriptions to control the video generation process, limiting desired fine-grained customization such as texture, object position, motion patterns, etc. To tackle this constraint and pursue higher controllability, several controllable video synthesis methods have been proposed. These methods utilize additional control signals, such as depth or sketch, to guide the generation of videos. By incorporating extra structured guidance, the generated content can be precisely controlled and customized. Among these approaches, VideoComposer stands out as a pioneering and versatile compositional technique. It integrates multiple conditioning signals including textual, spatial and temporal conditions within a unified framework, offering enhanced controllability, compositionality, and realism in the generated videos. Despite the remarkable quality, these methods still rely on high-quality video-text data to unleash powerful and customizable synthesis. In contrast, our method can be directly merged into existing controllable frameworks to customize videos by exploiting text-free videos.

Method

We first provide a brief introduction to the preliminaries of the video diffusion model. Then, we will elaborate on the mechanisms of TF-T2V in detail. The overall framework of the proposed TF-T2V is displayed in Fig. 2.

Diffusion models involve a forward diffusion process and a reverse iterative denoising stage. The forward process of diffusion models is gradually imposing random noise to clean data x0x_{0} in a Markovian chain:

where βt∈(0,1)\beta_{t}\in(0,1) is a noise schedule and TT is the total time step. When TT is sufficiently large, e.g. T=1000T=1000, the resulting xTx_{T} is nearly a random Gaussian distribution N(0,I)\mathcal{N}(0,I). The role of diffusion model is to denoise xTx_{T} and learn to iteratively estimate the reversed process:

We usually train a denoising model x^θ\hat{x}_{\theta} parameterized by θ\theta to approximate the original data x0x_{0} and optimize the following v-prediction problem:

where cc is conditional information such as textual prompt, and vv is the parameterized prediction objective. In representative video diffusion models , the denoising model x^θ\hat{x}_{\theta} is a latent 3D-UNet modified from its 2D version by inserting additional temporal blocks, which is optimized in the latent feature space by applying a variational autoencoder , and Eq. 3 is applied on each frame of the input video to train the whole model.

2 TF-T2V

The objective of TF-T2V is to learn a text-conditioned video diffusion model to create visually appealing and temporally coherent videos with text-free videos or partially paired video-text data. Without loss of generality, we first describe the workflow of our TF-T2V in the scenario where only text-free video is used. With merely text-free videos available for training, it is challenging to guide content creation by textual information since there lacks text-visual correspondence. To tackle this issue, we propose to resort to web-scale and high-quality image-text datasets , which are publicly accessible on the Internet. However, this raises another question: how can we leverage the image-text data and text-free videos in a unified framework?

Recalling the network architecture in 3D-UNet, the spatial modules mainly focus on appearance modeling, and the temporal modules primarily aim to operate motion coherence. The intuition is that we can utilize image-text data to learn text-conditioned spatial appearance generation and adopt high-quality text-free videos to guide consistent motion dynamic synthesis. In this way, we can perform text-to-video generation in a single model to synthesize high-quality and consistent videos during the inference stage. Based on this, the proposed TF-T2V consists of two branches: a content branch for spatial appearance generation and a motion branch for motion dynamic synthesis.

Like previous text-to-image works , the content branch of TF-T2V takes a noised image Iimage∈H×W×CI_{image}\in H\times W\times C as input, where HH, WW, CC are the height, width, and channel dimensions respectively, and employs conditional signals (i.e., text and image embeddings) to offer semantic guidance for content generation. This branch primarily concentrates on optimizing the spatial modules in the video diffusion model and plays a crucial role in determining appealing visual quality. In order to ensure that each condition can also control the created content separately, we randomly drop text or image embeddings with a certain probability during training. The text and image encoders from CLIP are adopted to encode embeddings.

2.2 Motion dynamic synthesis

The pursuit of producing highly temporally consistent videos is a unique hallmark of video creation. Recent advancements in the realm of video synthesis usually utilize large-scale video-text datasets such as WebVid10M to achieve coherent video generation. However, acquiring large-scale video-text pairs consumes extensive manpower and time, hindering the scaling up of video diffusion models. To make matters worse, the widely used WebVid10M is a watermarked and low-resolution (around 360P) dataset, resulting in unsatisfactory video creation that cannot meet the high-quality video synthesis requirements.

To mitigate the above issues, we propose to leverage high-quality text-free videos that are easily accessible on video media platforms, e.g., YouTube and TikTok. To fully excavate the abundant motion dynamics within the text-free videos, we train a image-conditioned model. By optimizing this image-to-video generation task, the temporal modules in the video diffusion model can learn to perceive and model diverse motion dynamics. Specifically, given a noised video Ivideo∈F×H×W×CI_{video}\in F\times H\times W\times C, where FF is the temporal length, the motion branch of TF-T2V learns to recover the undisturbed video guided by the image embedding. The image embedding is extracted from the center frame of the original video by applying CLIP’s image encoder .

Since large-scale image-text data used for training contains abundant movement intentions , TF-T2V can achieve text-to-video generation by assembling spatial appearances involving motion trends and predicted motion dynamics. When extra paired video-text data is available, we conduct both text-to-video and image-to-video generation based on video-text pairs to train TF-T2V and further enhance the perception ability for desirable textual control.

In addition, we notice that previous works apply the training loss (i.e., Eq. 3) on each frame of the input video individually without considering temporal correlations between frames, suffering from incoherent appearances and motions. Inspired by the early study finding that the difference between two adjacent frames usually contains motion patterns, e.g., dynamic trajectory, we thus propose a temporal coherence loss that utilizes the frame difference as an additional supervisory signal:

where ojo_{j} and vjv_{j} are the predicted frame and corresponding ground truth. This loss term measures the discrepancy between the predicted frame differences and the ground truth frame differences of the input parameterized video. By minimizing Eq. 4, TF-T2V helps to alleviate frame flickering and ensures that the generated videos exhibit seamless transitions and promising temporal dynamics.

2.3 Training and inference

In order to mine the complementary advantages of spatial appearance generation and motion dynamic synthesis, we jointly optimize the entire model in an end-to-end manner. The total loss can be formulated as:

where Lbase\mathcal{L}_{base} is imposed on video and image together by treating the image as a “single frame” video, and λ\lambda is a balance coefficient that is set empirically to 0.1.

After training, we can perform text-guided video generation to synthesize temporally consistent video content that aligns well with the given text prompt. Moreover, TF-T2V is a general framework and can also be inserted into existing compositional video synthesis paradigm by incorporating additional spatial and temporal structural conditions, allowing for customized video creation.

Experiments

In this section, we present a comprehensive quantitative and qualitative evaluation of the proposed TF-T2V on text-to-video generation and composition video synthesis.

Implementation details. TF-T2V is built on two typical open-source baselines, i.e., ModelScopeT2V and VideoComposer . DDPM sampler with T=1000T=1000 steps is adopted for training, and we employ DDIM with 50 steps for inference. We optimize TF-T2V using AdamW optimizer with a learning rate of 5e-5. For input videos, we sample 16 frames from each video at 4FPS and crop a 448×256448\times 256 region at the center as the basic setting. Note that we can also easily train high-definition video diffusion models by collecting high-quality text-free videos (see examples in the Appendix). LAION-5B is utilized to provide image-text pairs. Unless otherwise stated, we treat WebVid10M, which includes about 10.7M video-text pairs, as a text-free dataset to train TF-T2V and do not use any textual annotations. To study scaling trends, we gathered about 10M high-quality videos without text labels from internal data, termed the Internal10M dataset.

Metrics. (i) To evaluate text-to-video generation, following previous works , we leverage the standard Fréchet Inception Distance (FID), Fréchet Video Distance (FVD), and CLIP Similarity (CLIPSIM) as quantitative evaluation metrics and report results on MSR-VTT dataset . (ii) For controllability evaluation, we leverage depth error, sketch error, and end-point-error (EPE) to verify whether the generated videos obey the control of input conditions. Depth error measures the divergence between the input depth conditions and the eliminated depth of the synthesized video. Similarly, sketch error examines the sketch control. EPE evaluates the flow consistency between the reference video and the generated video. In addition, human evaluation is also introduced to validate our method.

2 Evaluation on text-to-video generation

Tab. 1 displays the comparative quantitative results with existing state-of-the-art methods. We observe that TF-T2V achieves remarkable performance under various metrics. Notably, TF-T2V trained on WebVid10M and Internal10M obtains higher performance than the counterpart on WebVid10M, revealing promising scalable capability. We show the qualitative visualizations in Fig. 3. From the results, we can find that compared with previous methods, TF-T2V obtains impressive video creation in terms of both temporal continuity and visual quality. The human assessment in Tab. 2 also reveals the above observations. The user study is performed on 100 randomly synthesized videos.

3 Evaluation on compositional video synthesis

We compare the controllability of TF-T2V and VideoComposer on 1,000 generated videos in terms of depth control (Tab. 3), sketch control (Tab. 4) and motion control (Tab. 5). The above experimental evaluations highlight the effectiveness of TF-T2V by leveraging text-free videos. In Fig. 4 and 5, we show the comparison of TF-T2V and existing methods on compositional video generation. We notice that TF-T2V exhibits high-fidelity and consistent video generation. In addition, we conduct a human evaluation on 100 randomly sampled videos and report the results in Tab. 6. The preference assessment provides further evidence of the superiority of the proposed TF-T2V.

4 Ablation study

Effect of temporal coherence loss. To enhance temporal consistency, we propose a temporal coherence loss. In Tab. 7, we show the effectiveness of the proposed temporal coherence loss in terms of frame consistency. The metric results are obtained by calculating the average CLIP similarity of two consecutive frames in 1,000 videos. We further display the qualitative comparative results in Fig. 6 and observe that temporal coherence loss helps to alleviate temporal discontinuity such as color shift.

5 Evaluation on semi-supervised setting

Through the above experiments and observations, we verify that text-free video can help improve the continuity and quality of generated video. As previously stated, TF-T2V also supports the combination of annotated video-text data and text-free videos to train the model, i.e., the semi-supervised manner. The annotated text can provide additional fine-grained motion signals, enhancing the alignment of generated videos and the provided prompts involving desired motion evolution. We show the comparison results in Tab. 8 and find that the semi-supervised manner reaches the best performance, indicating the effectiveness of harnessing text-free videos. Notably, TF-T2V-Semi outperforms ModelScopeT2V trained on labeled WebVid10M, possessing good scalability. Moreover, the qualitative evaluations in Fig. 7 show that existing methods may struggle to synthesize text-aligned consistent videos when textual prompts involve desired temporal evolution. In contrast, TF-T2V in the semi-supervised setting exhibits excellent text-video alignment and temporally smooth generation.

Conclusion

In this paper, we present a novel and versatile video generation framework named TF-T2V to exploit text-free videos and explore its scaling trend. TF-T2V effectively decomposes video generation into spatial appearance generation and motion dynamic synthesis. A temporal coherence loss is introduced to explicitly constrain the learning of correlations between adjacent frames. Experimental results demonstrate the effectiveness and potential of TF-T2V in terms of fidelity, controllability, and scalability.

Acknowledgements. This work is supported by the National Natural Science Foundation of China under grant U22B2053 and Alibaba Group through Alibaba Research Intern Program.

References

More experimental details

In A Recipe for Scaling up Text-to-Video Generation with Text-free Videos, we show the comparison on compositional motion-to-video synthesis. TF-T2V achieves more appealing results than the baseline VideoComposer. Following prior works, we use an off-the-shelf pre-trained variational autoencoder (VAE) model from Stable Diffusion 2.1 to encode the latent features. The VAE encoder has a downsample factor of 8. In the experiment, the network structure of TF-T2V is basically consistent with the open source ModelScopeT2V and VideoComposer to facilitate fair comparison. Note that TF-T2V is a plug-and-play framework that can also be applied to other text-to-video generation and controllable video synthesis methods. For human evaluation, we randomly generate 100 videos and ask users to rate and evaluate them. The highest score for each evaluation content is 100%, the lowest score is 0%, and the final statistical average is reported.

Additional ablation study

Effect of joint training. In our default setting, we jointly train the spatial and temporal blocks in the video diffusion model to fully exploit the complementarity between image and video modalities. An alternative strategy is to separate the spatial and temporal modeling into sequential two stages. We conduct a comparative experiment on these two strategies in Tab. 9. The results demonstrate the rationality of joint optimization in TF-T2V.

Scaling trend under semi-supervised settings. In Fig. 9, we vary the number of text-free videos and explore the scaling trend of TF-T2V under the semi-supervised settings. From the results, we can observe that FVD (↓\downarrow) gradually decreases as the number of text-free videos increases, revealing the strong scaling potential of our TF-T2V.

More experimental results

To further verify that TF-T2V can be extended to high-definition video generation, we leverage text-free videos to train a high-resolution text-to-video model, such as 896×512896\times 512. As shown in Fig. 10, we can notice that in addition to generating 448×256448\times 256 videos, our method can be easily applied to the field of high-definition video synthesis. For high-definition compositional video synthesis, we additionally synthesize 1280×6401280\times 640 and 1280×7681280\times 768 videos to demonstrate the excellent application potential of our method. The results are displayed in Fig. 11 and Fig. 12.

Limitations and future work

In this paper, we only doubled the training set to explore scaling trends due to computational resource constraints, leaving the scalability of larger scales (such as 10×10\times or 100×100\times) unexplored. We hope that our approach can shed light on subsequent research to explore the scalability of harnessing text-free videos. The second limitation of this work is the lack of exploration of processing long videos. In the experiment, we follow mainstream techniques and sample 16 frames from each video clip to train our TF-T2V for fair comparisons. Investigating long video generation with text-free videos is a promising direction. In addition, we find that if the input textual prompts contain some temporal evolution descriptions, such as “from right to left”, “rotation”, etc., the text-free TF-T2V may fail and struggle to accurately synthesize the desired video. In the experiment, even though we noticed that semi-supervised TF-T2V helped alleviate this problem, it is still worth studying and valuable to precisely generate satisfactory videos with high standards that conform to the motion description in the given text.