ModelScope Text-to-Video Technical Report
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, Shiwei Zhang
Introduction
Artificial intelligence has expanded the boundaries of content generation in diverse modalities following simple and intuitive instructions. This encompasses textual content , visual content and auditory content . In the realm of visual content generation, research efforts have been put into image generation and editing , leveraging diffusion models .
While video generation continues to pose challenges. A primary hurdle lies in the training difficulty, which often leads to generated videos exhibiting sub-optimal fidelity and motion discontinuity. This presents ample opportunities for further advancements. The open-source image generation methods (e.g., Stable Diffusion ) have significantly advanced research in the text-to-image synthesis. Nonetheless, the field of video generation has yet to benefit from a publicly available codebase, which could potentially catalyze further research efforts and progress.
To this end, we propose a simple yet easily trainable baseline for video generation, termed ModelScope Text-to-Video (ModelScopeT2V). This model, which has been publicly available, presents two technical contributions to the field. Firstly, regarding the architecture of ModelScopeT2V, we explore LDM in the field of text-to-video generation by introducing the spatio-temporal block that models temporal dependencies. Secondly, regarding the pre-training technique, we propose a multi-frame training strategy that utilizes both the image-text and video-text paired datasets, thereby enhancing the model’s semantic richness. Experiments have shown that videos generated by ModelScopeT2V perform quantitatively and qualitatively similar or superior to other state-of-the-art methods. We anticipate that ModelScopeT2V can serve as a powerful and effective baseline for future research related to the video synthesis, and propel innovative advancements and exploration.
Related work
Diffusion probabilistic models. Diffusion Probabilistic Models (DPM) was originally proposed in . The early efforts of utilizing DPM into the image synthesis task at scale has been proven effective , surpassing dominant generative models, e.g., generative adversarial networks and variational autoencoders , in terms of diversity and fidelity. The original DPM suffers the problem of low-efficiency when adopted for image/video generation due to the iterative denoising process and the high-resolution pixel space. To solve the first obstacle, research efforts focused on improving the sampling efficiency by learning-free sampling and learning based sampling . To address the second obstacle, methods like LDM , LSGM and RDM resorted to manifolds with lower intrinsic dimensionality . Our modelScopeT2V follows LDM but modifies it to the video generation task.
Text-to-image synthesis via diffusion models. By receiving knowledge from natural language instructions (e.g., CLIP and T5 ), diffusion models can be utilized for text-to-image synthesis. LDM designed language-conditioned image generator by augmenting the UNet backbone with cross-attention layers . DALL-E 2 generated image embeddings for a diffusion decoder with CLIP text encoder. The concurrent work, Imagen , found the scalibility of T5, which means increasing the size of T5 could boost image fidelity and language-image alignment. Building on existing image generation framework , Imagic achieves text-based semantic image edits by leveraging intermediate text embeddings that align with the input image and the target text. Composer reformulates images as various compositions, thus enabling image generation from not just texts, but also sketches, masks, depthmaps and more. The ModelScopeT2V initialize the spatial part from Stable Diffusion model , and proposes the spatio-temporal block that empowers the capacity of temporal dependencies.
Text-to-video synthesis via diffusion models. Generating realistic videos remains challenging due to the difficulty in generating videos with high fidelity and motion continuity . Recent works have utilized diffusion models to generate authentic videos . Text, as a highly intuitive and informative instruction, has been employed to guide video generation. Various approaches have been proposed, such as Video Diffusion , which introduces a spatio-temporal factorized 3D Unet with a novel conditional sampling mechanism. Imagen Video synthesizes high definition videos given a text prompt by designing a video generator and a video super-resolution model. Make-A-Video employs off-the-shelf text-to-image generation model combined with spatio-temporal factorized diffusion models to generate high-quality videos without relying on paired video-text data. Instead of modeling the video distribbution in the visual space (e.g., RGB format), MagicVideo designed a generator in the latent space with a distribution adapter, which did not require temporal convolution. In oder to improve the content and motion performance, VideoFusion decouples per-frame noise into base noise and residual noise, which benefits from a well-pretrained DALL-E 2 and provides better control over content and motion. Targetting decoupling, Gen-1 defines the content latents as CLIP embeddings and the structure latents as the monocular depth estimates, attaining superior decoupled controllability. ModelScopeT2V proposes a simple yet effective training pipeline that benefits from semantic diversity inherent in image-text and video-text paired datasets, further enhancing the learning process and performance in video generation.
Methodology
In this section, we introduce the overall architecture of ModelScopeT2V (Sec. 3.1), the key spatio-temporal block used in UNet (Sec. 3.2) and a multi-frame training mechanism that stabilizes training (Sec. 3.3).
Structure overview. Given a text prompt , ModelScopeT2V outputs a video through a latent video diffusion model that conforms to the semantic meaning of the prompt. The architecture of the latent video diffusion model is shown in Fig. 2. As illustrated in the figure, the training video and generated video are in the visual space. The diffusion process and denoising UNet are in the latent space. VQGAN converts the data between visual space and latent space through its encoder and decoder . The latent space in this paper is proposed by VQGAN . Given a training video with frames, we could encode the video with VQGAN encoder as,
The UNet includes different types of blocks such as the initial block, downsampling block, spatio-temporal block and upsampling block, represented by the dark blue squares in Figure 2. Most parameters of the model are concentrated in the denoising UNet. So the denoising UNet is considered as the core of the latent video diffusion model, which performs the diffusion process in the latent space. The UNet aims to denoise from to by predicting the noise of each step. Given a specific step index , the predicted noise can be formulated as:
where denotes the prompt, represent the text embedding, and is the latent variable in -th step. In this case, denotes during training, and denotes the denoised latent video representation during inference. The model’s objective is to minimize the discrepancy between the predicted noise and ground-truth noise . Consequently, the training loss of the UNet can be formulated as:
2 Spatio-temporal block
Our video diffusion model is built upon a UNet architecture, which consists of four key building blocks: the initial block, the downsampling block, the spatio-temporal block and the upsampling block. The initial block projects the input into the embedding space, while the downsampling and upsampling blocks spatially downsample and upsample the feature maps, respectively. The spatio-temporal block plays a crucial role in capturing complex spatial and temporal dependencies in the latent space, thereby enhancing the quality of video synthesis. To this end, we leverage the power of spatio-temporal convolutions and attentions to comprehensively obtain such complex dependencies.
Figure 3 illustrates the architecture of the spatio-temporal block in our video diffusion model. To ensure effectively synthesising videos, we factorise the convolution and the attention mechanism over space and time . Therefore, the spatio-temporal block is composed of four sub-components, namely spatial convolution, temporal convolution, spatial attention and temporal attention. The spatio-temporal convolutions capture correlations across frames by convolving over both the spatial and temporal dimensions of the video, while the spatio-temporal attentions capture correlations across frames by selectively attending to different regions and time steps within the video. By utilizing these spatio-temporal building blocks, our model can effectively learn spatio-temporal representations and generate high-quality videos. Specifically, one spatio-temporal block consists of spatial convolutions, temporal convolutions, spatial attention, and temporal attention operations. In our experiments, we set by default to achieve a balance between performance and computational efficiency. Regarding the architecture of the two spatial attentions, we instantiate it in two different ways. The first attention is a cross-attention module that conditions the visual features on the textual features, allowing cross-modal interactions. The other attention is a self-attention module that operates solely on visual features, responsible for spatial modeling. While the two temporal attentions are both self-attention module.
Figure 4(a) illustrates the spatio-temporal convolution, which is composed of both spatial and temporal convolutions. The spatial convolution employs a convolution kernel of size to extract features from the latent features within each frame, where and denote the height and width of video frames in visual space. Meanwhile, the temporal convolution adopts a convolution kernel of size to extract features from frames, where represents the number of frames for each video.
Figure 4(b) displays the spatio-temporal attention, which consists of spatial and temporal attention modules. In detail, the spatial attention operates on the latent features in the spatial dimension of size , while the temporal attention operates on the temporal dimension of size . We adopt the popular Transformer architecture to instantiate both attention mechanisms.
3 Multi-frame training
ModelScopeT2V is designed to be trained on large-scale video-text paired datasets, such as WebVid , which is domain-aligned with video generation. Nonetheless, the scale of such datasets is orders of magnitude smaller compared to image-text paired datasets, such as LAION . Despite initializing the spatial part of ModelScopeT2V with Stable Diffusion , training solely on video-text paired datasets can hinder semantic diversity and lead to catastrophic forgetting of image-domain expertise during training . To overcome this limitation and leverage the strengths of both datasets, we propose a multi-frame training approach. Specifically, one eighth of GPUs for training are applied to image-text paired datasets, while the remaining GPUs handle video-text paired datasets. Since the model structure could adapt to any frame length, one image could be considered as a video with frame length 1 for those GPUs training on image-text paired datasets.
Experiments
We utilize the LAION-5B dataset as image-text pairs, specifically the LAION2B-en subset, as the model focuses on English input. The LAION dataset encompasses objects, people, scenes, and other real-world elements.
The WebVid dataset comprises almost 10 million video-text pairs, with a majority of the videos having a resolution of . Each video clip has a duration of approximately 30 seconds. During the model training, we selected the middle square portion and randomly picked 16 frames with 3 frames per second as training data.
The MSR-VTT dataset is used to validate our model performance and is not utilized for training. This dataset includes 10k video clips, each of which is annotated with 20 sentences. To obtain FID-vid and FVD metric results, 2,048 video clips were randomly selected from the test set, and one sentence was randomly chosen from each clip to generate videos. When evaluating on CLIPSIM metric, we followed previous works and use nearly 60k sentences from the whole test split as prompts to generate videos.
1.2 Model instantiation and hyper-parameters
We employ DDPM with steps for training and use DDIM sampler in classifier-free guidance with steps for inference by default. ModelScopeT2V primarily consists of three modules: the Text encoder , VQGAN, and Denoising UNet. The pretrained checkpoint for initializing VQGAN and Denoising UNet are obtained from Stable Diffusion version 2.1https://github.com/Stability-AI/stablediffusion. The parameters in VQGAN remains frozen during training and inference. The outputs of temporal convolution and temporal attention are initialized as zeros, enabling ModelScopeT2V to generate meaningful yet temporally discontinuous frames at the beginning of training. As the training progresses, the temporal structures will be optimized to learn the temporal correspondence between frames, thereby synthesising continuous videos.
1.3 Training details
We train ModelScopeT2V using the AdamW optimizer with a learning rate of . Our model is trained on 80G NVIDIA A100 GPUs. We perform multi-frame training as detailed in Section 3.3, specifically using a batch size of 1,400 for images and a batch size of 3,200 for videos, and training 267 thousand iterations. The compression factor of VQGAN is 8, meaning that it converts RGB images of size into latent representations of size . For the text encoder, we set the maximum text length to , and embedding dim , which are consistent with the pre-trained OpenCLIPhttps://github.com/mlfoundations/open_clip.
We empirically observe that employing either temporal convolution or temporal attention can augment ModelScopeT2V’s ability to capture temporal dependency. This observation is partly supported by VideoCrafthttps://github.com/VideoCrafter/VideoCrafter which only contains temporal attention for temporal modeling. We take a step further by employing both the temporal convolution and the temporal attention, which facilitates the ModelScopeT2V to achieve superior temporal modeling. In detail, we use temporal convolution blocks and temporal attention block for each spatio-temporal block. These temporal blocks account for 552 million parameters out of the total 1,345 million parameters in our UNet, indicating that 39% parameters of the UNet parameters are dedicated to capturing temporal information. As a result, the entire ModelScopeT2V model (including VQGAN and the text encoder) comprises approximately 1.7 billion parameters.
We observe that use more layers of temporal convolution would lead to better temporal ability. Since the kernel size of 1D CNN in temporal convolution is 3, temporal convolution layers could lead the local receptive field as 81 in each spatio-temporal block, which is enough for 16 output frames per video. For multi-frame training, the temporal convolution and temporal attention mechanisms are still active. Our experiments show it is unnecessary to change the range of parameters for different frame settings.
2 Main results
ModelScopeT2V is evalutated on MSR-VTT dataset. We conduct the evaluation under a zero-shot setting since ModelScopeT2V is not trained on MSR-VTT. We compare ModelScopeT2V with several state-of-the-art models using FID-vid , FVD , and CLIPSIM metrics. The FID-vid and FVD are assessed based on 2,048 randomly selected videos from MSR-VTT test split, where we compute the metrics using the middle 16 frames of each video with an FPS of 3. CLIPSIM are evaluated based on all captions from MSR-VTT test split following . The resolution of the generated videos is consistently .
As shown in Table 1, ModelScopeT2V achieves the best performance on both FID-vid (i.e., 11.09) and FVD (i.e., 550), indicating that our generated videos are visually similar to the ground truth videos. Our model also obtains a competitive score of 0.2930 on CLIPSIM, suggesting that our generated videos are semantically similar to the text prompts. The CLIPSIM score of our model is only marginally lower than that of Make-A-Video , while they utilize additional data from HD-VILA-100M for training.
3 Qualitative results
In this subsection, we compare the qualitative results of ModelScopeT2V with other state-of-the-art methods. To facilitate comparison with Make-A-Video and Imagen Video, generated video frames with the same frame index are presented in the same column with ModelScopeT2V. The videos generated by Make-A-Video https://makeavideo.studio and Imagen Video https://imagen.research.google/video were downloaded from their official webpages. Six frames are uniformly sampled from each video for comparison. This comparison is fair in terms of video duration, as all three methods (i.e., Make-A-Video, Imagen Video, and ModelScopeT2V) generate 16-frame videos aligned with given texts. One difference is that Imagen Video generates videos with a aspect ratio of , while the other two generate videos with a aspect ratio of .
The quanlitative comparison between ModelScopeT2V and Make-A-Video is displayed in Figure 5. We can observe that both methods generate videos of high quality, which is consistent with the quantitative results. However, the “robot” in the first example and the “dog” in the second example generated by ModelScopeT2V exhibit a superior degree of realism. We attribute this advantage to our model’s joint training with image-text pairs, which enhances its comprehension of the correspondence between textual and visual data. In the third example, while the “industrial site” generated by Make-A-Video is more closely aligned with the prompt, depicting the overall scene of a “storm”, ModelScopeT2V produces a distinctive interpretation showcasing two “abandoned” factories and a gray sky, captured from various camera angles. This difference stems from Make-A-Video’s use of image CLIP embedding to generate videos, which can result in less dynamic motion information. In general, ModelScopeT2V demonstrates a wider range of motion in its generated videos, distinguishing it from Make-A-Video.
The comparison of our method, ModelScopeT2V, with Imagen Video is illustrated in Figure 6. While Imagen Video generate more vivid and contextually relevant video content, ModelScopeT2V effectively depicts the content of the prompt, albeit with some roughness in the details. For instance, in the first example of Figure 6, it’s noteworthy that Imagen Video generates a video whose second frame illustrates a significantly distorted dog’s tongue, exposing the model’s limitations in accurately rendering the real world. On the other hand, ModelScopeT2V demonstrates its potential in robustly representing the content described in the prompt. It’s worth noting that the superior performance of Imagen Video can be attributed to the employment of the T5 text encoder, a base model with a larger number of parameters, and a larger-scale training dataset, which is not utilized in ModelScopeT2V. Considering the performance, our ModelScopeT2V lays a strong foundation for future improvements and shows considerable promise in the domain of text-to-video generation.
4 Community development
We have made the code for ModelScopeT2V publicly available on the GitHub repositories of ModelScopehttps://github.com/modelscope/modelscope/blob/master/modelscope/models/multi_modal/video_synthesis and Diffuserhttps://huggingface.co/spaces/damo-vilab/modelscope-text-to-video-synthesis/blob/main/app.py. Additionally, we have provided online demos of ModelScopeT2V on ModelScopehttps://modelscope.cn/studios/damo/text-to-video-synthesis/summary and HuggingFace https://huggingface.co/spaces/damo-vilab/modelscope-text-to-video-synthesis. The open-source community has actively engaged with our model and uncovered several applications of ModelScopeT2V. Notably, projects such as sd-webui-text2videohttps://github.com/deforum-art/sd-webui-text2video and Text-To-Video-Finetuninghttps://github.com/ExponentialML/Text-To-Video-Finetuninghave extended the model’s usage and broadened its applicability. Additionally, the video generation feature of ModelScopeT2V has already been successfully utilized for the creation of short videoshttps://youtu.be/Ank49I99EI8.
Conclusion
This paper proposes ModelScopeT2V, the first open-source diffusion-based text-to-video generation model. To enhance the ModelScopeT2V’s ability to modeling temporal dynamics, we design the spatio-temporal block that incorporates spatio-temporal convolution and spatio-temporal attention. Furthermore, to leverage semantics from comprehensive visual content-text pairs, we perform multi-frame training on both text-image pairs and text-video pairs. Comparative analysis of videos generated by ModelScopeT2V and those produced by other state-of-the-art methods demonstrate similar or superior performance quantitatively and qualitatively.
As for future research directions, we expect to adopt additional conditions to enhance video generation quality. Potential strategies include using multi-condition approaches or the LoRA technique . Additionally, an interesting topic to explore could be the generation of longer videos that encapsulate more semantic information.