Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

David Junhao Zhang, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao, Lingmin Ran, Yuchao Gu, Difei Gao, Mike Zheng Shou

Introduction

Remarkable progress has been made in developing large-scale pre-trained text-to-Video Diffusion Models (VDMs), including closed-source ones (e.g., Make-A-Video (Singer et al., 2022), Imagen Video (Ho et al., 2022a), Video LDM (Blattmann et al., 2023a), Gen-2 (Esser et al., 2023)) and open-sourced ones (e.g., VideoCrafter (He et al., 2022), ModelScopeT2V (Wang et al., 2023a). These VDMs can be classified into two types: (1) Pixel-based VDMs that directly denoise pixel values, including Make-A-Video (Singer et al., 2022), Imagen Video (Ho et al., 2022a), PYoCo (Ge et al., 2023), and (2) Latent-based VDMs that manipulate the compacted latent space within a variational autoencoder (VAE), like Video LDM (Blattmann et al., 2023a) and MagicVideo (Zhou et al., 2022).

However, both of them have pros and cons. Pixel-based VDMs can generate motion accurately aligned with the textual prompt but typically demand expensive computational costs in terms of time and GPU memory, especially when generating high-resolution videos. Latent-based VDMs are more resource-efficient because they work in a reduced-dimension latent space. But it is challenging for such small latent space (e.g., 8×58\times 5 for 64×4064\times 40 videos) to cover rich yet necessary visual semantic details as described by the textual prompt. Therefore, as shown in Fig. 2, the generated videos often are not well-aligned with the textual prompts. On the other hand, if the generated videos are of relatively high resolution (e.g., 256×160256\times 160 videos), the latent model will focus more on spatial appearance but may also ignore the text-video alignment.

To marry the strength and alleviate the weakness of pixel-based and latent-based VDMs, we introduce Show-1, an efficient text-to-video model that generates videos of not only decent video-text alignment but also high visual quality. Further, Show-1 can be trained on large-scale datasets with manageable computational costs. Specifically, we follow the conventional coarse-to-fine video generation pipeline (Ho et al., 2022a; Blattmann et al., 2023a) which starts with a module to produce keyframes at a low resolution and a low frame rate. Then we employs a temporal interpolation module and super-resolution module to increase temporal and spatial quality respectively.

In these modules, prior studies typically employ either pixel-based or latent-based VDMs across all modules. While purely pixel-based VDMs tend to have heavy computational costs, exclusively latent-based VDMs can result in poor text-video alignment and motion inconsistencies. In contrast, we combine them into Show-1 as shown in Fig. 3. To accomplish this, we employ pixel-based VDMs for the keyframe module and the temporal interpolation module at a low resolution, producing key frames of precise text-video alignment and natural motion with low computational cost. Regarding super-resolution, we find that latent-based VDMs, despite their inaccurate text-video alignment, can be re-purposed to translate low-resolution video to high-resolution video, while maintaining the original appearance and the accurate text-video alignment of low-resolution video. Inspired by this finding, for the first time, we propose a novel two-stage super-resolution module that first employs pixel-based VDMs to upsample the video from 64×4064\times 40 to 256×160256\times 160 and then design a novel expert translation module based on latent-based VDMs to further upsample it to 572×320572\times 320 with low computation cost.

In summary, our paper makes the following key contributions:

Upon examining pixel and latent VDMs, we discovered that: 1) pixel VDMs excel in generating low-resolution videos with more natural motion and superior text-video synchronization compared to latent VDMs; 2) when using the low-resolution video as an initial guide, conventional latent VDMs can effectively function as super-resolution tools by simple expert translation, refining spatial clarity and creating high-quality videos with greater efficiency than pixel VDMs.

We are the first to integrate the strengths of both pixel and latent VDMs, resulting into a novel video generation model that can produce high-resolution videos of precise text-video alignment at low computational cost (15G GPU memory during inference).

Our approach achieves state-of-the-art performance on standard benchmarks including UCF-101 and MSR-VTT.

Previous Work

Text-to-image generation. (Reed et al., 2016) stands as one of the initial methods that adapts the unconditional Generative Adversarial Network (GAN) introduced by (Goodfellow et al., 2014) for text-to-image (T2I) generation. Later versions of GANs delve into progressive generation, as seen in (Zhang et al., 2017) and (Hong et al., 2018). Meanwhile, works like (Xu et al., 2018) and (Zhang et al., 2021) seek to improve text-image alignment. Recently, diffusion models have contributed prominently to advancements in text-driven photorealistic and compositional image synthesis (Ramesh et al., 2022; Saharia et al., 2022). For attaining high-resolution imagery, two prevalent strategies emerge. One integrates cascaded super-resolution mechanisms within the RGB domain (Nichol et al., 2021; Ho et al., 2022b; Saharia et al., 2022; Ramesh et al., 2022). In contrast, the other harnesses decoders to delve into latent spaces (Rombach et al., 2022; Gu et al., 2022). Owing to the emergence of robust text-to-image diffusion models, we are able to utilize them as solid initialization of text to video models.

Text-to-video generation. Past research has utilized a range of generative models, including GANs (Vondrick et al., 2016; Saito et al., 2017; Tulyakov et al., 2018; Tian et al., 2021; Shen et al., 2023), Autoregressive models (Srivastava et al., 2015; Yan et al., 2021; Le Moing et al., 2021; Ge et al., 2022; Hong et al., 2022), and implicit neural representations (Skorokhodov et al., 2021; Yu et al., 2021). Inspired by the notable success of the diffusion model in image synthesis, several recent studies have ventured into applying diffusion models for both conditional and unconditional video synthesis (Voleti et al., 2022; Harvey et al., 2022; Zhou et al., 2022; Wu et al., 2022b; Blattmann et al., 2023b; Khachatryan et al., 2023; Höppe et al., 2022; Voleti et al., 2022; Yang et al., 2022; Nikankin et al., 2022; Luo et al., 2023; An et al., 2023; Wang et al., 2023b). Several studies have investigated the hierarchical structure, encompassing separate keyframes, interpolation, and super-resolution modules for high-fidelity video generation. Magicvideo (Zhou et al., 2022) and Video LDM (Blattmann et al., 2023a) ground their models on latent-based VDMs. On the other hand, PYoCo (Ge et al., 2023), Make-A-Video (Singer et al., 2022), and Imagen Video (Ho et al., 2022a) anchor their models on pixel-based VDMs. Contrary to these approaches, our method seamlessly integrates both pixel-based and latent-based VDMs.

Show-1

Denoising Diffusion Probabilistic Models (DDPMs). DDPMs, as detailed in (Ho et al., 2020), represent generative frameworks designed to reproduce a consistent forward Markov chain x1,…,xTx_{1},\ldots,x_{T}. Considering a data distribution x0∼q(x0)x_{0}\sim q(x_{0}), the Markov transition q(xt∣xt−1)q(x_{t}|x_{t-1}) is conceptualized as a Gaussian distribution, characterized by a variance βt∈(0,1)\beta_{t}\in(0,1). Formally, this is defined as:

Applying the principles of Bayes and the Markov characteristic, it’s feasible to derive the conditional probabilities q(xt∣x0)q(x_{t}|x_{0}) and q(xt−1∣xt,x0)q(x_{t-1}|x_{t},x_{0}), represented as:

The model’s adaptable parameters θ\theta are optimized to ensure the synthesized reverse sequence aligns with the forward sequence.

UNet architecture for text to image model. The UNet model is introduced by (Spr, 2015) for biomedical image segmentation. Popular UNet for text-to-image diffusion model usually contains multiple down, middle, and up blocks. Each block consists of a resent2D layer, a self-attention layer, and a cross-attention layer. Text condition cc is inserted into to cross-attention layer as keys and values. For a text-guided Diffusion Model, with the text embedding cc the objective is given by:

2 Turn Image UNet to Video

We incorporate the spatial weights from a robust text-to-image model. To endow the model with temporal understanding and produce coherent frames, we integrate temporal layers within each UNet block. Specifically, after every Resnet2D block, we introduce a temporal convolution layer consisting of four 1D convolutions across the temporal dimension. Additionally, following each self and cross-attention layer, we implement a temporal attention layer to facilitate dynamic temporal data assimilation. Specifically, a frame-wise input video x∈RT×C×H×Wx\in\mathcal{R}^{T\times C\times H\times W}, where CC is the number of channels and HH and WW are the spatial latent dimensions. The spatial layers regard the video as a batch of independent images (by transposing the temporal axis into the batch dimension), and for each temporal layer, the video is reshaped back to temporal dimensions.

3 Pixel-based Keyframe Generation Model

Given a text input, we initially produce a sequence of keyframes using a pixel-based Video UNet at a very low spatial and temporal resolution. This approach results in improved text-to-video alignment. The reason for this enhancement is that we do not require the keyframe modules to prioritize appearance clarity or temporal consistency. As a result, the keyframe modules pays more attention to the text guidance. The training objective for the keyframe modules is following Eq. 5.

Why we choose pixel diffusion over latent diffusion here? Latent diffusion employs an encoder to transform the original input xx into a latent space. This results in a reduced spatial dimension, for example, H/8,W/8H/8,W/8, while concentrating the semantics and appearance into this latent domain. For generating keyframes, our objective is to have a smaller spatial dimension, like 64×4064\times 40. If we opt for latent diffusion, this spatial dimension would shrink further, perhaps to around 8×58\times 5, which might not be sufficient to retain ample spatial semantics and appearance within the compacted latent space. On the other hand, pixel diffusion operates directly in the pixel domain, keeping the original spatial dimension intact. This ensures that necessary semantics and appearance information are preserved. For the following low resolution stages, we all utilize pixel-based VDMs for the same reason.

4 Temporal Interpolation Model

To enhance the temporal resolution of videos we produce, we suggest a pixel-based temporal interpolation diffusion module. This method iteratively interpolates between the frames produced by our keyframe modules. The pixel interpolation approach is built upon our keyframe modules, with all parameters fine-tuned during the training process. We employ the masking-conditioning mechanism, as highlighted in (Blattmann et al., 2023a), where the target frames for interpolation are masked. In addition to the original pixel channels C, as shown in Fig. 4 we integrate 4 supplementary channels into the U-Net’s input: 3 channels are dedicated to the RGB masked video input, while a binary channel identifies the masked frames. As depicted in the accompanying figure, during a specific noise timestep, we interpolate three frames between two consecutive keyframes, denoted as xtix^{i}_{t} and xti+1x^{i+1}_{t}. For the added 3 channels, values of ziz^{i} and zi+1z^{i+1} remain true to the original pixel values, while the interpolated frames are set to zero. For the final mask channel, the mask values mim^{i} and mi+1m^{i+1} are set to 1, signifying that both the initial and concluding frames are available, with all others set to 0. In conclusion, we merge these components based on the channel dimension and input them into the U-Net. For ziz^{i} and zi+1z^{i+1}, we implement noise conditioning augmentation. Such augmentation is pivotal in cascaded diffusion models for class-conditional generation, as observed by (Ho et al., 2022a), and also in text-to-image models as noted by (He et al., 2022). Specifically, this method aids in the simultaneous training of diverse models in the cascade. It minimizes the vulnerability to domain disparities between the output from one cascade phase and the training inputs of the following phase. Let the interpolated video frames be represented by x,∈R4T×C×H×Wx^{,}\in\mathcal{R}^{4T\times C\times H\times W}. Based on Eq. 5, we can formulate the updated objective as:

5 Super-resolution at Low Spatial Resolution

To improve the spatial quality of the videos, we introduce a pixel super-resolution approach utilizing the video UNet. For this enhanced spatial resolution, we also incorporate three additional channels, which are populated using a bilinear upscaled low-resolution video clip, denoted as xu,,∈R4T×C×4H×4Wx^{,,}_{u}\in\mathcal{R}^{4T\times C\times 4H\times 4W} through bilinear upsampling. In line with the approach of(Ho et al., 2022c), we employ Gaussian noise augmentation to the upscaled low resolution video condition during its training phase, introducing a random signal-to-noise ratio. The model is also provided with this sampled ratio. During the sampling process, we opt for a consistent signal-to-noise ratio, like 1 or 2. This ensures minimal augmentation, assisting in the elimination of artifacts from the prior phase, yet retaining a significant portion of the structure.

Given that the spatial resolution remains at an upscaled version throughout the diffusion process, it’s challenging to upscale all the interpolated frames, denoted as x′∈R4T×C×H×Wx^{{}^{\prime}}\in\mathcal{R}^{4T\times C\times H\times W}, to x′′∈R4T×C×4H×4Wx^{{}^{\prime\prime}}\in\mathcal{R}^{4T\times C\times 4H\times 4W} simultaneously on a standard GPU with 24G memory. Consequently, we must divide the frames into four smaller segments and upscale each one individually.

However, the continuity between various segments is compromised. To rectify this, as depicted in the Fig. 4, we take the upscaled last frame of one segment to complete the three supplementary channels of the initial frame in the following segment.

6 Super-resolution at High Spatial Resolution

Through our empirical observations, we discern that a latent-based VDM can be effectively utilized for enhanced super-resolution with high fidelity. Specifically, we design a distinct latent-based VDM that is tailored for high-caliber, high-resolution data. We then apply a noising-denoising procedure, as outlined by SDEdit (Meng et al., 2021), to the samples from the preliminary phase. As pointed out by (Balaji et al., 2022), various diffusion steps assume distinct roles during the generation process. For instance, the initial diffusion steps, such as from 1000 to 900, primarily concentrate on recovering the overall spatial structure, while subsequent steps delve into finer details. Given our success in securing well-structured low-resolution videos, we suggest adapting the latent VDM to specialize in high-resolution detail refinement. More precisely, we train a UNet for only the 0 to 900 timesteps (with 1000 being the maximum) instead of the typical full range of 0 to 1000, directing the model to be a expert emphasizing high-resolution nuances. This strategic adjustment significantly enhances the end video quality, namely expert translation. During the inference process, we use bilinear upsampling on the videos from the prior stage and then encode these videos into the latent space. Subsequently, we carry out diffusion and denoising directly in this latent space using the latent-based VDM model, while maintaining the same text input. This results in the final video, denoted as x′′′∈R4T×C×16H×16Wx^{{}^{\prime\prime\prime}}\in\mathcal{R}^{4T\times C\times 16H\times 16W}.

Why we choose latent-based VDM over pixel-based VDM here? Pixel-based VDMs work directly within the pixel domain, preserving the original spatial dimensions. Handling high-resolution videos this way can be computationally expensive. In contrast, latent-based VDMs compress videos into a latent space (for example, downscaled by a factor of 8), which results in a reduced computational burden. Thus, we opt for the latent-based VDMs in this context.

Experiments

For the generation of pixel-based keyframes, we utilized DeepFloydhttps://github.com/deep-floyd/IF as our pre-trained Text-to-Image model for initialization, producing videos of dimensions 8×64×40×3(T×H×W×3)8\times 64\times 40\times 3(T\times H\times W\times 3). In our interpolation model, we initialize the weights using the keyframes generation model and produce videos with dimensions of 29×64×40×329\times 64\times 40\times 3. For our initial model, we employ DeepFloyd’s SR model for spatial weight initialization, yielding videos of size 29×256×16029\times 256\times 160. In the subsequent super-resolution model, we modify the ModelScope text-to-video model and use our proposed expert translation to generate videos of 29×576×32029\times 576\times 320.

The dataset we used for training is WebVid-10M (Bain et al., 2021). Training and hyperparameterdetails can be found in appendix Table 5.

2 Quantitative Results

UCF-101 Experiment. For our preliminary evaluations, we employ IS and FVD metrics. UCF-101 stands out as a categorized video dataset curated for action recognition tasks. When extracting samples from the text-to-video model, following PYoCo (Ge et al., 2023), we formulate a series of prompts corresponding to each class name, serving as the conditional input. This step becomes essential for class names like jump rope, which aren’t intrinsically descriptive. We generate 20 video samples per prompt to determine the IS metric. For FVD evaluation, we adhere to methodologies presented in prior studies (Le Moing et al., 2021; Tian et al., 2021) and produce 2,048 videos.

From the data presented in Table 1, it’s evident that Show-1’s zero-shot capabilities outperform or are on par with other methods. This underscores Show-1’s superior ability to generalize effectively, even in specialized domains. It’s noteworthy that our keyframes, interpolation, and initial super-resolution models are solely trained on the publicly available WebVid-10M dataset, in contrast to the Make-A-Video models, which are trained on other data.

MSR-VTT Experiment. The MSR-VTT dataset (Xu et al., 2016) test subset comprises 2,9902,990 videos, accompanied by 59,79459,794 captions. Every video in this set maintains a uniform resolution of 320×240320\times 240. We carry out our evaluations under a zero-shot setting, given that Show-1 has not been trained on the MSR-VTT collection. In this analysis, Show-1 is compared with state-of-the-art models, on performance metrics including FID-vid (Heusel et al., 2017), FVD (Unterthiner et al., 2018), and CLIPSIM (Wu et al., 2021). For FID-vid and FVD assessments, we randomly select 2,048 videos from the MSR-VTT testing division. CLIPSIM evaluations utilize all the captions from this test subset, following the approach (Singer et al., 2022). All generated videos consistently uphold a resolution of 256×256256\times 256.

Table 2 shows that, Show-1 achieves the best performance in both FID-vid (a score of 13.08) and FVD (with a score of 538). This suggests a remarkable visual congruence between our generated videos and the original content. Moreover, our model secures a notable CLIPSIM score of 0.3076, emphasizing the semantic coherence between the generated videos and their corresponding prompts. It is noteworthy that our CLIPSIM score surpasses that of Make-A-Video (Singer et al., 2022), despite the latter having the benefit of using additional training data beyond WebVid-10M.

Human evaluation. We gather an evaluation set comprising 120 prompts that encompass camera control, natural scenery, food, animals, people, and imaginative content. The survey is conducted on Amazon Mechanical Turk. Following Make a Video (Singer et al., 2022), we assess video quality, the accuracy of text-video alignment and motion fidelity. In evaluating video quality, we present two videos in a random sequence and inquire from annotators which one possesses superior quality. When considering text-video alignment, we display the accompanying text and prompt annotators to determine which video aligns better with the given text, advising them to overlook quality concerns. For motion fidelity, we let annotators to determine which video has the most natural notion. As shown in Table 3, our method achieves the best human preferences on all evaluation parts.

3 Qualitative Results

As depicted in Fig. 5, our approach exhibits superior text-video alignment and visual fidelity compared to the recently open-sourced ModelScope (Wang et al., 2023a) and ZeroScopehttps://huggingface.co/cerspense/zeroscope-v2-576w. Additionally, our method matches or even surpasses the visual quality of the current state-of-the-art methods, including Imagen Video and Make-A-Video.

4 Ablation studies.

Impact of different combinations of pixel-based and latent-based VDMs. To assess the integration method of pixel and latent-based VDMs, we conduct several ablations. For fair comparison, we employe the T5 encoder (Raffel et al., 2020) for text embedding in all low-resolution stages and the CLIP text encoder (Radford et al., 2021) for high-resolution stages. As indicated in Tab. 4, utilizing pixel-based VDMs in the low-resolution stage and latent diffusion for high-resolution upscaling results in the highest CLIP score with reduced computational expenses. On the other hand, implementing pixel-based VDMs during the high-resolution upscaling stage demands significant computational resources. These findings reinforce our proposition that combining pixel-based VDMs in the low-resolution phase and latent-based VDMs in the high-resolution phase can enhance text-video alignment and visual quality while minimizing computational costs.

Impact of expert translation of latent-based VDM as super-resolution model. We provide visual comparison between models with and without expert translation. As elaborated in Section 3.6, “with expert translation” refers to training the latent-based VDMs using timesteps 0-900 (with a maximum timestep of 1000), while “w/o expert translation” involves standard training with timesteps 0-1000. As evident in Fig. 6, the model with expert translation produces videos of superior visual quality, exhibiting fewer artifacts and capturing more intricate details.

Conclusion

We introduce Show-1, an innovative model that marries the strengths of pixel and latent based VDMS. Our approach employs pixel-based VDMs for initial video generation, ensuring precise text-video alignment and motion portrayal, and then uses latent-based VDMs for super-resolution, transitioning from a lower to a higher resolution efficiently. This combined strategy offers high-quality text-to-video outputs while optimizing computational costs.

Ethics Statement

Our pretrained T2I model, Deep-IF, is trained using web data, and our models utilize WebVid-10M. Given this, there’s a potential for our method to not only learn but also amplify societal biases, which could include inappropriate or NSFW content. To address this, we can integrate the CLIP model to detect NSFW content and filter out such instances.

Reproducibility Statement

We take the following steps to guarantee reproducibility: (1) Our codes, along with model weights, will be public available. (2) The training and hyperparameter details can be found in appendix Table 5.

References

Appendix A Appendix

We list details of our models in Table 5.