MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation
Yanhui Wang, Jianmin Bao, Wenming Weng, Ruoyu Feng, Dacheng Yin, Tao Yang, Jingxu Zhang, Qi Dai Zhiyuan Zhao, Chunyu Wang, Kai Qiu, Yuhui Yuan, Chuanxin Tang, Xiaoyan Sun, Chong Luo, Baining Guo
Introduction
Diffusion models have achieved remarkable success in text-to-image generation, such as DALL-E , Stable Diffusion , Imagen , among others. They can generate unseen image content based on novel text concepts, showcasing impressive capabilities for image content generation and manipulation. Consequently, researchers have sought to extend the success of diffusion models to text-to-video generation.
One prevalent strategy involves training large-scale text-to-video diffusion models directly . These models employ cascade spatiotemporal diffusion models to learn from text and video pairs. While capable of producing high-quality videos, they pose challenges due to substantial GPU resource requirements and the need for extensive training data. Recently, some works have presented a cost-effective strategy. These methods entail the insertion of temporal layers into a text-to-image model, followed by fine-tuning on paired text and video data to create a text-to-video model. However, videos generated using this approach may encounter issues related to appearance and temporal coherence. We argue that maintaining appearance and temporal coherence is crucial for effective video generation.
In this paper, we present a novel approach, named MicroCinema, which employs a divide-and-conquer strategy to address appearance and temporal coherence challenges in video generation. The model features a two-stage generation pipeline. In the first stage, we generate a center frame, which serves as the foundation for subsequent video clip generation based on the input text. This design offers the flexibility to utilize any existing text-to-image generator for the initial stage, allowing users to incorporate their own images to establish the desired scene.
The second stage, known as image&text-to-video, concentrates on motion modeling. To achieve this, we leverage the open-source text-to-image generation model called Stable Diffusion (SD) and inject temporal layers into it to obtain a three-dimensional (3D) network structure. The SD model has been trained on the filtered large-scale LAION dataset . Its strong performance in generating high-quality images demonstrates its ability to capture spatial information within visual signals. To further enhance the model’s ability to capture motion, we propose two core designs for the image&text-to-video model.
First, we introduce an Appearance Injection Network to inject the given image as a condition to guide the video generation. Concretely, it shares the structure of the encoder and middle part of the 3D U-Net and feeds the learned feature into the main branch via dense injection in a multi-layer manner. The dense injection operation better injects the appearance into the main branch, thus releasing the model from appearance modeling and encouraging the model dedicated to motion modeling. Second, we propose an appearance-aware noise strategy to preserve the pre-trained capability of the SD model by modifying the i.i.d. noise in the diffusion process. Specifically, we add an appropriate amount of center frame to the i.i.d. noise without altering the overall diffusion training and inference process. This appearance-aware noise provides an intuitive cue to the model to generate a video whose appearance is similar to the given center frame, thereby unleashing its motion modeling capabilities.
Equipped with these designs, our framework can generate appearance-preserving and coherent videos with a given image and text. Extensive experiments demonstrate the superiority of MicroCinema. We achieve a state-of-the-art zero-shot FVD of 342.86 on UCF101 and 377.40 on MSR-VTT when training on the public WebVid-10M dataset.
In summary, our contributions are presented as follows:
We introduce an innovative two-stage text-to-video generation pipeline that capitalizes on a key-frame image generated by any off-the-shelf text-to-image generator in the initial stage. Subsequently, both the generated key-frame image and text serve as inputs for the video generation process in the second stage.
We propose an Appearance Injection Network structure to encourage the 3D model to focus on motion modeling during the image&text-to-video generation process.
We introduce an effective and distinctive Appearance Noise Prior tailored for fine-tuning text-to-image diffusion models. This modification significantly elevates the quality of video generation.
In-depth quantitative and qualitative results are presented to validate the video generation capability of our proposed MicroCinema.
Related Work
The task of video generation involves addressing two fundamental challenges: image generation and motion modeling. Various approaches have been employed for image generation, including Generative Adversarial Networks (GANs) , Variational autoencoder (VAE) and flow-based methods . Recently, the state-of-the-art methods are built on top of diffusion models such as DALLE-2 , Stable Diffusion , GLIDE and Imagen , which achieved impressive results. Extending these models for video generation is a natural progression, though it necessitates non-trivial modifications.
Text-to-Video Models. Image diffusion models adopt 2D U-Net with few exceptions . To generate temporally smooth videos, temporal convolution (conv) or attention layers are also introduced. Notably, in Align-your-latents , 3D conv layers are interleaved with the existing spatial layers to align individual frames in a temporally consistent manner. This factorized space-time design has become the de facto standard and has been used in VDM , Imagen Video , and CogVideo . Besides, it creates a concrete partition between the pre-trained two-dimensional (2D) conv layers and the newly initialized temporal conv layers, allowing us to train the temporal convolutions from scratch while retaining the previously learned knowledge in the spatial convolutions’ weights. More recent work Latent-Shift introduces no additional parameters but shifts channels of spatial feature maps along the temporal dimension, enabling the model to learn temporal coherence. Many approaches rely on temporal layers to implicitly learn motions from paired text and videos . The generated motions, however, still lack satisfactory global coherence and fail to faithfully capture the essential movement patterns of the target subjects.
Leveraging Prior for Text-to-Video Diffusion Models. Generating natural motions poses a significant challenge in video generation. Many attempts are focused on leveraging prior into the text-to-video generation process. ControlVideo directly utilizes ground truth motions, represented as depth maps or edge maps, as conditions for video diffusion models, demonstrating the importance of motion in video generation. GD-VDM involves a two-phase generation process leveraging generating depth videos followed by a novel diffusion Vid2Vid model that generates a coherent real-world video in the autonomous driving scenario. However, it is not clear whether it can be applied to general scenes due to the lack of depth training data. Make-Your-Video utilizes a standalone depth estimator to extract depth from a driving video, bypassing the need for depth generation, to generate new videos. In Leo , a motion diffusion model is trained to generate a sequence of motion latents, fed to a decoder network to recover the optical flows to animate the input image. Meanwhile, other methods involve linear displacement of codes in latent space , noise correlation , and generating textual descriptions for motion , serving as conditions for video generation models. More recent work PYoCo proposes the video diffusion noise prior for a diffusion model and cost-effectively fine-tuning the text-to-image model.
Our proposed framework differs significantly from existing methods by employing a Divide-and-Conquer strategy. In our approach, we first generate images and subsequently capture motion dynamics along the temporal dimension. We also notice that a previous method Make-A-Video has adopted a similar approach. However, our method introduces a novel model network design and incorporates an appearance-noise prior. This innovation ensures the generated video not only maintains the appearance established in the initial stage but also demonstrates superior motion modeling capabilities, a feature notably absent in Make-A-Video and concurrently related methods .
MicroCinema
Our approach decomposes the text-to-video generation process into two distinct stages. Initially, we employ prevalent off-the-shelf text-to-image generation techniques to produce a key frame. Subsequently, both the key frame, acting as the center frame, and the text prompts are used as input to the image&text-to-video model to generate videos. We argue that the image&text-to-video model in a two-stage framework exhibits the potential for yielding more natural videos compared to the single-stage text-to-video model. This argument rests on the premise that by incorporating the center frame as a condition, our approach mitigates the model’s burden in learning complicated appearance.
In the image&text-to-video generation stage, we adopt a cascaded approach to produce high-quality videos. First, we use a base image&text-to-video model to generate low frame rate videos from given image and text. Then, an adapted temporal interpolation model, derived from the base model, is employed to augment the frame rate. Finally, an off-the-shelf spatial super-resolution model is incorporated to render high-definition videos. This paper focuses on explaining the base model design and detailing its adaptation into the temporal interpolation model.
Base image&text-to-video model. Fig. 2 illustrates the overall architecture of the base image&text-to-video model in MicroCinema. This model is extended from the widely recognized Stable Diffusion (SD) model . Following previous attempts , we first extend the 2D U-Net into a 3D structure. We first enhance the original model by adding a 1D temporal convolution (conv) layer following each 2D spatial conv layer, enhancing its ability to handle temporal alignments. Additionally, we introduce a 1D temporal attention layer after every 2D spatial attention layer. These attention layers effectively capture long-range temporal correspondence, complementing the functionality of the 1D conv layers. To protect the strong capability of SD, we zero-initialize all the convolution and attention temporal layers and add a skip connection to it. Based on these modifications, we obtain a 3D model that can handle text-to-video generation. The base image&text-to-video model showcases two crucial innovations: the AppearNet and the appearance noise prior. Both are designed to incorporate appearance information from the key frame. A detailed explanation of these technical advancements will be provided in Sec. 3.2 and Sec. 3.3.
Temporal interpolation model. Our base model generates videos at a resolution of pixels with a frame rate of 2 frames per second (fps). To enhance temporal quality, we train a temporal interpolation model designed for four-fold temporal super-resolution (TSR). This TSR model mirrors the architecture of the base model with slight modifications. The base model employs only one conditional image (the center frame) while there are two conditional images (the start and end frames) in the TSR model. Accordingly, we alter the input of the AppearNet, shifting from duplicating the center frame to utilizing the interpolated latent representations of the given first and last frames. Leveraging this model consecutively on adjacent frames from previous steps boosts the frame rate from 2 fps to 32 fps.
2 Appearance Injection Network
To enhance the model’s capability in handling reference center frame, we introduce the Appearance Injection Network, abbreviated as AppearNet, to the 3D network as depicted in Fig. 2. Inspired by ControlNet , we let AppearNet inherit the encoder and the middle part of the backbone network. Let be the frame length of the output video. Then the center frame is replicated for times to create an image sequence, denoted as . It is used as input to the AppearNet to offer a robust appearance cue for generating output video frames.
We apply a multi-scale and dense fusion mechanism to seamlessly integrate the outputs of the AppearNet into the main branch. The multi-scale output of AppearNet is injected into both the encoder and the decoder of the main branch at the corresponding scales. In addition to the commonly used additive operation, we introduce an effective strategy of de-normalization to inject the feature into the corresponding normalization layer of the main branch. As shown in Fig. 3, at the -th feature level, let denote the activation map in the main branch. Before integration, we perform 3D Group Normalization on :
Here and are the means and standard deviations of ’s group-wise activations. For AppearNet feature integration, let be the AppearNet embedding on this feature level, we compute the output activation by denormalizing the normalized according to , formulated as
where and are obtained by convolving from the feature map . The computed and are multiplied and added to in an element-wise manner. Equipped with this design, our entire structure could better maintain the appearance from a given center frame while possessing the ability to generate videos based on text and image conditions.
3 Appearance Noise Prior
Fine-tuning from a text-to-image model proves to be a cost-effective approach for acquiring a video generation model. However, this process presents challenges due to the transition of the output space from images to videos. In the context of a typical T2I diffusion model, it tends to generate appearance-irrelevant images from a sequence of independent noise (sampled from ). In video generation, a sequence of independent noise should ideally yield a video with a coherent appearance. Therefore, the fine-tuning process may potentially compromise the capability of the original 2D T2I model. Our focus lies in preserving the effectiveness of the original 2D T2I model during the fine-tuning process for the image&text-to-video model.
For our proposed image&text-to-video model, the model should expand the given center image to a sequence of frames, which have a similar appearance to the center frame. Consequently, the output video is predominantly determined by the center frame rather than the sampled noise in the original diffusion process. To address this, we modify the noise distribution to align with the appearance of the given center frame. Leveraging the denoising property of the diffusion model, we introduce Appearance Noise Prior by adding an appropriate amount of the center frame into the noise, in order to generate appearance-conditioned frames.
Let denote the noise corresponding to a video clip with frames, represents the noise added to the frame. is the latent tensor of center frame, is the randomly sampled noise from . The training noise for our model is defined as:
where is the coefficient that controls the amount of the center frame.
Consequently, the diffusion process of our model can be expressed in the following form, the t-step noisy input of the diffusion model is:
where is the latent tensors of an input video and is the same as defined in DDPM .
For training, we adhere to the stable diffusion training setting and use noise prediction with the following loss function:
where is the time step, is the reference image input, is the text input, , are the ground-truth video and noisy input, represents the output of the model., respectively. Our appearance noise prior employs the same inference strategy as previous methods, differing only in the initiation of noise, which aligns with our formulation. This consistency allows for the direct application of existing ODE sample algorithms. For a thorough understanding of the proofs, please refer to the supplementary materials.
Experiments
Datasets. MicroCinema is trained using the public WebVid-10M dataset , comprising ten million video-text pairs. This dataset exhibits a wide spectrum of video motions, ranging from near-static sequences to those with frequent and abrupt scene changes. Text captions are automatically sourced from alt text, resulting in some noise. Therefore, we perform a filtering process which excludes video-text pairs with a low CLIP score or with excessively high or low motions.
Evaluation metrics. The quantitative evaluations are conducted on UCF-101 and MSR-VTT benchmark datasets under the zero-shot setting. On UCF-101, Frechet Video Distance (FVD) and Inception Score (IS) are reported to validate the temporal consistency, where 10K or 2K video clips are generated using a sentence template of the category names. On MSR-VTT, Frechet Inception Distance (FID) and CLIPSIM are provided to assess the quality of generated frames and the semantic correspondence, where CLIPSIM is computed by averaging the cosine similarity of CLIP embeddings between generated frames and captions. We utilized captions from the MSR-VTT validation set, comprising 2.9K entries, to generate the video clips. The condition images are generated with SDXL model on all evaluations unless otherwise specified.
Implementation details. MicroCinema generates video from text in a two-stage process. In the first stage, we employ a SOTA T2I model SDXL to generate an image according to the text. Then in the second stage, the image&text-to-video generation model is built upon the pre-trained weights of Stable Diffusion 2.1. Temporal layer is zero initialized. During training, the learning rate for the temporal modules is set to 2e-5, while the learning rate for the spatial model is 10 times smaller than that of the temporal modules. The output of the image&text-to-video model yields a video clip with a spatial resolution of 320x320, consisting of 9 frames at a rate of 2fps. The model is trained on the filtered WebVid dataset for one epoch, employing the same diffusion noise schedule as SD2.1.
Quantitative evaluation. We evaluate zero-shot text-to-video generation performance on both UCF101 and MSR-VTT. In the case of UCF101, we produce 10K samples using simple clip captions. For MSR-VTT, we generate 2.9K samples using the captions provided within the MSR-VTT dataset. Tab. 1 presents a quantitative comparison between MicroCinema and alternative text-to-video models. These models are categorized into two groups based on whether they leverage additional data beyond WebVid-10M. As data is of paramount importance to the training of video generation model, we can observe that the methods in the first group (with additional data) achieve superior overall performance compared to those in the second group. Remarkably, despite being exclusively trained on the WebVid-10M dataset, our proposed MicroCinema, with its innovative design, achieves the most outstanding performance among all methods on both datasets. It achieves the lowest FVD values of 342.86 on UCF101 and 377.40 on MSR-VTT. Notably, MicroCinema surpasses methods employing additional data and notably outperforms those relying solely on the WebVid-10M dataset by a considerable margin.
Qualitative evaluation Fig. 4 compares the video clips generated by MicroCinema and two other methods, known as Make-A-Video and Video LDM. Compared to the other two methods, our approach can generate noticeable and accurate motion.
2 Ablation Studies
We conduct ablation studies to validate our design choices concerning appearance injection and shifted noise training. For efficiency purposes, we adopt several different settings from the experiments used for system comparison. First, models employing different options are trained using a 1M subset of the filtered WebVid-10M dataset. Each model undergoes training for 64K steps (equivalent to one epoch) with a batch size set at 16. Second, during inference, we directly generate 17 frames without using the TSR module. Third, for the zero-shot FVD and IS evaluation on UCF101, we uniformly select 2K samples instead of using the entire 10K test set. It’s notable that while using this smaller 2K-sample test set, the absolute FVD values are higher compared to those derived from the larger 10K-sample test set for the same model.
In an image&text-to-video model, the most important design choice is how to inject the appearance information into the primary U-Net of the generation model.
Concatenation (Concat). A common approach in related work is to direct concatenation of the latent features from the reference image to the noise input of the U-Net.
Addition to Decoder (Add-to-Dec). Our approach, however, adopts an AppearNet, akin to ControlNet for structure control. In the vanilla ControlNet, embeddings from the ControlNet are added to the decoder of the U-Net. We employ a similar operation in this setting.
Addition to Encoder and Decoder (Add-to-EncDec). Considering that the reference image contains more appearance details than the structural information in ControlNet, we propose injecting appearance into both the encoder and the decoder of the U-Net. This improvement is expected to elevate generation quality through a more comprehensive integration of appearance features.
Addition to Encoder and Decoder with SPADE (Add-to-EncDec-SPADE). Expanding further, we integrate the SPADE technique, commonly used in image generation models, by infusing information into the GroupNorm layers of the U-Net. This final design constitutes the core of our method, MicroCinema.
Tab. 2 presents a comparative analysis of the zero-shot FVD performance among these four design choices. The results clearly demonstrate that our final model achieves the most superior performance.
2.2 Appearance Noise Prior
Another key mechanism we propose for injecting appearance information into the image&text-to-video generation network is the Appearance Noise Prior. One crucial and intricate parameter within this mechanism is the proportion, denoted by , determining the addition of the reference image to the noise input of the diffusion model. Selecting an optimal value for involves balancing potential harm to the pre-trained image generation model and the advantages gained from additional information.
This set of ablation studies aims to empirically identify the most effective parameter for use with Appearance Noise Prior. Alongside , which we test at values of 0 (no Appearance Noise Prior), 0.03, and 0.1. Besides, according to our formulation, an appropriate amount of appearance may also help during the inference stage. Therefore, we also explore the impact of adding extra to during the inference stage. Therefore, the sampling noise during the inference stage is , where is sampled from .
Fig. 5 shows the FVD scores across various combinations of . We find that the lowest FVD score occurs when and . Notably, this configuration leads to a substantial reduction in FVD compared to the baseline (, ), dropping from 692 to 508, alongside a notable increase in IS from 18.5 to 29.6.
3 Control in Image&Text-to-Video Model
Our image&text-to-video model relies on both a reference image and a text prompt for conditioning. Our findings emphasize that the reference image’s quality profoundly influences the resultant video quality. Consequently, both the text caption and the text-to-image model used to generate the reference image significantly impact the system’s performance. We simplify our experiments by using the base image&text-to-video model without using the temporal super-resolution component. In this setup, we adopt the resulting model to generate 17 frames with 10K samples on UCF101 for evaluating IS and FVD.
Tab. 3 illustrates the influence of various prompts on the model’s generated outputs. We utilize the state-of-the-art SDXL model for text-to-image generation. Within the table, “simple” denotes a straightforward prompt created by connecting “a video of” with the motion tag, while “LLaVA-1.5” signifies a generated caption via the LLaVA-1.5 model using the key frame as input. Results indicate that a well-crafted prompt correlates with higher-quality videos generated by the model.
Moreover, we assess the impact of employing different Text-to-Image (T2I) models. Tab. 4 underscores the substantial influence of T2I models on the FVD and IS of the generated videos. Notably, the design of MicroCinema affords us the flexibility to integrate various T2I models for generating the first-stage reference image, with potential performance enhancements stemming from advancements in text-to-image models.
Lastly, we demonstrate the controllability of text prompts on the model’s output in Fig. 6. Conditioning the image&text-to-video model on the same image but varying text prompts results in significantly different videos aligned with their respective text prompts. This exemplifies the high level of control embedded within our model.
Conclusion
We presented MicroCinema, an innovative text-to-video generation approach that employs the Divide-and-Conquer paradigm to tackle two key challenges in video synthesis: appearance generation and motion modeling. Our strategy employs a two-stage pipeline, utilizing any existing text-to-image generator for initial image generation and subsequently introducing a dedicated image&text-to-video framework designed to focus on motion modeling. To improve motion capture, we propose an Appearance Injection Network structure, complemented by an appearance-aware noise prior. Experimental results showcase MicroCinema’s superiority, achieving a state-of-the-art zero-shot Frechet Video Distance (FVD) of 342.86 on UCF101 and 377.40 on MSR-VTT. We anticipate our research will inspire future advancements in this direction.
References
Appendix A More Qualitative Results
we fist examine the influence of the Appearance Noise Prior on the quality of generated video results by varying the parameters and . As illustrated in Fig. 9, videos generated with the integration of the Appearance Noise Prior display heightened coherence and superior quality in comparison to those generated without this prior. The introduced prior proves beneficial by endowing the model with enhanced capabilities to preserve distinctive characteristics of input images, even when they deviate from the training data in WebVid-10M.
Besides, our empirical findings indicate that adjusting the ratio of the Appearance Noise Prior contributes to the production of high-resolution videos by our model. As illustrated in Fig. 10, the model demonstrates effective generation of 512x512 resolution videos, surpassing its original training resolution of 320x320, thanks to the integration of the Appearance Noise Prior.
Additionally, we discover that the Appearance Noise Prior plays a crucial role in enhancing the efficiency of the diffusion process. As illustrated in Fig. 11, in situations involving simpler motion patterns, the integration of the Appearance Noise Prior empowers the network to produce satisfactory results even with a reduced sampling step count, set at 5. This decrease in steps significantly improves the efficiency of video production. For instance, employing our base image&text-to-video model to generate a 9-frame video at 2 fps now requires only 1.3 seconds.
A.2 More qualitative Results of MicroCinema
In this section, we present additional video generation results Fig. 12 and Fig. 13. We utilize Midjourney as the initial stage text-to-video model. It is evident that the videos generated through our method not only maintain aesthetic quality in imagery but also exhibit clear and coherent motion.
A.3 Qualitative Comparison with Previous Work.
We provide additional examples for comparison with previous works in Fig. 14, Fig. 15, and Fig. 16. Our approach demonstrates the ability to generate visually stunning videos, akin to cinematic quality. In comparison to prior work, it showcases superior image quality, enhanced temporal consistency, greater stylistic diversity, and improved textual coherence.
Appendix B Proof of Appearance Noise Prior
In this section, we present a proof of the compatibility of the Appearance Noise Prior with all ODE samplers. We demonstrate that incorporating the Appearance Noise Prior and employing new noise as supervision does not necessitate alterations to the sampler process itself. Instead, it only requires modifications to the initial noise during sampling.
Firstly, we introduce the standard framework of Denoising Diffusion Probabilistic Models (DDPM). The forward process in DDPM, when articulated in discrete form, is as follows:
The corresponding Stochastic Differential Equation (SDE) process of the DDPM can be represented by a unified expression, given by the following equation:
where is a standard Wiener process. To derive the expressions for and , as approaches infinity, two continuous functions, and , can be defined:
where and are the coefficients corresponding to those in equations Eq. 6 and Eq. 7, respectively. By substituting Eq. 10 into Eq. 7, utilizing , and considering the limit as , and subsequently applying a Taylor series expansion for approximation, equation Eq. 7 can be reformulated as follows:
By setting , and incorporating into the equation, we obtain:
Upon simplification, we obtain the Stochastic Differential Equation (SDE) formulation of DDPM:
For the SDE process described in Eq. 8, the corresponding reverse Ordinary Differential Equation (ODE) process is represented by the following equation:
Given that follows a Gaussian distribution , its score function can be related to the noise as follows:
where is estimated using the following loss function:
where denotes the noisy data distribution of and . By utilizing equation Eq. 16, for DDPM models that implement the -prediction, the reverse ODE process is articulated as follows:
By incorporating Eq. 14, the final form can be derived as follows:
B.2 Appearance Noise Prior
To simplify notation, let . Then the forward process of Appearance Noise Prior is change to:
Applying a transformation to the coefficient preceding in Eq. 21 yields:
Similar to equation (6), considering the limit conditions , equation Eq. 22 can be reformulated as follows:
By setting , and incorporating into the equation, we obtain:
Reverse ODE process can be represented by the following equation:
As we employ the following form of the loss function:
Therefore, the relationship between the score function and the network’s estimated value becomes:
By substituting Eq. 28 and Eq. 25 into Eq. 26, and noting that the coefficient preceding is eliminated, we obtain the final form of the Reverse ODE:
In the context of the Appearance Noise Prior, functions as the network’s output, paralleled by in the DDPM framework. Notably, Eq. 29 and Eq. 19 exhibit identical forms. This similarity enables the straightforward integration of existing ODE sampling algorithms, with the only requisite modification being the adjustment of the initial sampling noise.
B.3 Implementation of Appearance Noise Prior
The implementation of Appearance Noise Prior in noise prediction models is straightforward. Traditionally, noise is added and trained using samples from a standard Gaussian distribution. With the Appearance Noise Prior, we modify this approach by superimposing an image prior onto the original noise, creating a new noise term for noise addition and supervision. During inference with ODE samplers, the initial sampling noise should be changed from to , where . To achieve more consistent results, the strength of the prior can be appropriately enhanced by adjusting to , thereby improving the consistency of the generated videos.
Appendix C Implementation Details
For text-to-image stage, we use SD2.1-Base and SDXL for Quantitative Experiments. Specific details of the samples are provided in the Tab. 5. And we use Midjourney and DALL-E 2 for Qualitative Results. In the Tab. 7, we present the specific details of our image&text-to-video model. For the spatial layer, we utilized the SD2.1-Base model architecture and initial parameters. Additionally, we incorporated the VAE provided by SD2.1-Base, along with the CLIP text encoder, both of which were frozen during the training process. The Tab. 6 displays the parameter count for each component; the image&text-to-video model possesses 2.0 billion parameters, which were actively trained, while the spatial learning rate was set to one-tenth of the temporal learning rate.
Appendix D Limitations
Our method is based on the latent diffusion approach of SD2.1, utilizing an SD-pretrained VAE to encode images into the latent space. Currently, the VAE exhibits limited reconstruction capabilities for small objects, particularly small faces, leading to sub-optimal performance in these cases, as illustrated in the Fig. 7. Conversely, the model performs significantly better with larger faces, also demonstrated in the Fig. 8. To address this issue, it is necessary to re-train the VAE with increased channel size.
Another limitation of our approach is that we focused solely on temporal super-resolution (TSR) without incorporating spatial super-resolution (SSR). Ideally, a joint spatial-temporal super-resolution process could potentially achieve further improvements in the quality of the generated videos. This will be one of our future work.