GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, Nan Duan

Introduction

“Creativity is a fundamental feature of human intelligence, and a challenge for AI.” . Recent advances in image and text generation have shown great creativity of machines, including GANs , VAEs , RNNs and Self-Attentions . However, it is still a challenge for the AI agent to create videos, especially for real-world diversity ones. Generating videos requires the machine to not only create a large number of pixels but also ensure semantic coherence among them.

We take up these challenges of generating videos from the text, namely the text-to-video generation (T2V) task. Given a natural description, T2V requires the machine to understand it and create a semantically consistent video. Although not much, there are still some works studying this topic using GANs. Firstly, and use GANs with 3D convolutions to generate fixed-length low-resolution videos. Then, uses a conditional filter to generate videos of varying lengths. integrates LSTM cells with 2D convolutional networks to model both frame quality and temporal coherence. However, these works conduct experiments on simple or small datasets, where generalization ability is limited.

Besides GAN-based methods, VQ-VAE is another promising research direction and has been shown great progress in generating images and videos, especially DALL-E for text-to-image generation. It successfully generates high-quality images from text. In this paper, we turn to the more challenging text-to-video generation task, where both spatial and temporal coherence of the visual information must be taken into account. Some other recent works apply VQ-VAE for the task of video prediction—forecasting future video frames given the past. Concurrently, we are the first to design a VQ-VAE pretrained model for the T2V task.

In this paper, we propose GODIVA to generate open-domain videos from text using VQ-VAE and three-dimensional sparse attention. Firstly, a VQ-VAE auto-encoder is trained to represent continuous video pixels with discrete video tokens. Then, a three-dimensional sparse attention model is trained using language as input and the discrete video tokens as labels to generate videos, considering temporal, column, and row information, as shown in Fig. 1.

Our contributions are three-fold: (1) We proposed an open-domain text-to-video pretrained model with a three-dimensional sparse attention mechanism, which can significantly reduce the computation cost; (2) We proposed a new Relative Matching (RM) Metric, which can evaluate both visual quality and semantic match for video generation; (3)We pretrained our proposed model on the HowTo100M dataset and demonstrated its video generation capabilities on both fine-tuning and zero-shot settings.

Related Works

In this section, we briefly review related works for video generation. We first review the video-to-video generation task, which has been widely studied in recent years. Then we review text-to-image and text-to-video generation. We also highlight the differences between previous models and ours.

Most video generation studies focus on video prediction tasks. Input the first few frames of a video, the video prediction task predicts the following frames of a video. We call it video-to-video (V2V) generation for comparison with text-to-video (T2V) generation.

Existing video-to-video generation can be divided into three categories. Firstly, deterministic methods directly model the tractable density using RNNs and CNNs and exploit both spatial and temporal information of a video. used ConvLSTM as the basic block to predict pixel motions instead of values. proposed PredNet, which predicts future frames by incorporating previous predictions. Further, proposed stacked ConvLSTM, which shares the hidden state among the layers in the stack. Recently, proposed ContextVP, which aggregates contextual information for each pixel in all possible directions. Secondly, the GAN-based methods avoid explicit density function and use a generator to generate videos and a discriminator to judge if the video is generated. proposed VGAN, which is the first model to generate videos using GANs. After that, proposed TGAN, which separates a spatiotemporal generator into time-series and space models to generate videos. Then, proposed MoCoGAN, which produces videos more efficiently by decomposing the latent space into the motion and the content subspaces. Recently, proposed TGAN2, which trains each sub-generator with its specific discriminator. Thirdly, VAE methods model the approximate density by capturing a low-dimensional representation z and optimize a lower bound on the likelihood. proposed SV2P to capture sequence uncertainty in a single set of latent variables kept fixed for each predicted sequence. Then, proposed SVG. They used a per-step latent variable (SVG-FP) and a variant with a learned prior (SVG-LP), which makes the prior at a certain timestep a function of previous frames. Recently, proposed a Latent Video Transformer, which encodes each frame of a video and predicts the discrete video features. GAN-based models.

Our model can be categorized into the VAE-based models. Different from recent VQ-VAE based works such as Latent Video Transformer , our work focus on text-to-video generation task instead of video-to-video generation task. We also incorporate a three-dimensional sparse attention to model the sparse relations between visual tokens.

2 Text-to-image generation

Text-to-image generation has been widely researched in recent years . The most similar work is DALL-E which successfully generates high-quality images from text. In this paper, we turn to a more challenging text-to-video generation task, which considers both spatial and temporal coherence of the visual information.

3 Text-to-video generation

Different from video-to-video generation, text-to-video generation has been few studied. Firstly, and use GANs with 3D convolutions to generate fixed-length low-resolution videos. Then, uses a conditional filter to generate videos of varying lengths. integrates LSTM cells with 2D convolutional networks to model both frame quality and temporal coherence.

Most text-to-video generation methods use GAN-based methods while our model incorporates a VQ-VAE for this task. As far as we know, this is the first paper that uses VQ-VAE for this task.

The GODIVA Method

Let xx be an observable video, and we use a discrete latent code zz to represent it, which has a lower dimension. In the following, we show how to represent xx using zz with VQ-VAE in section 3.1, and generate videos from the text by modeling P(z∣t)P(z|t) in section 3.2, where tt denotes the given text.

where the three items are reconstruction loss, codebook loss and commitment loss resepectively. β\beta is the weighting factor.sgsg denotes the stop gradient operator.

2 GODIVA video generator

Experiments

We pretrain GODIVA on Howto100M dataset , which consists of more than 136 million text-video pairs. We then evaluate our model on the MSR-VTT dataset , which consists of 10000 video clips with 20 human-annotated captions for each of them. We also train GODIVA from scratch on the Moving Mnist dataset and Double Moving Mnist dataset, both were automatically generated from the Mnist dataset . The original Moving Mnist dataset has two motions: up-down and left-right. In this paper, we follow and add four more directions: move left then right, move right then left, move up then down and move down then up.

2 Evaluation Metrics

It is challenging to quantitatively evaluate the performance of text-to-video generation models. This is mainly due to two reasons: Firstly, given a piece of text, there are countless corresponding videos. It is hard to objectively judge which one is better. Secondly, an evaluation metric should consider both visual quality and semantic matching of the generated video. To handle these challenges, we introduce two kind of metrics: A CLIP Similarity (SIM) metric and a Relative Matching (RM) metric for automatic evaluation in Sec. 4.2.1, A Visual Realisticity (VR) and Semantic Consistency (SC) metric for human evaluation in Sec. 4.2.2.

The key factor for judging the quality of the generated video is whether it matches the text. Using a pretrained visual-language matching model will inevitably introduce the bias of its domain data. Thanks to recent zero-shot work CLIP , which provides a strong zero-shot ability for visual-text matching and thus reduced those data biases. Since CLIP is pretrained between image and text, we calculate the similarities between text and each frame of the video and then take the average value as the semantic matching in Eq. (14).

where tt denotes the input text. v^\hat{v} is the predicted video with LL frames. Note that SIMSIM only provides the absolute score of the semantic match. To further reduce the influence of the CLIP model, we divide SIMSIM by the similarity between text and the ground-truth video to get a relative matching score, which we call Relative Matching (RM) metric, as denoted in Eq. (15).

where vv is the ground-truth video with LL frames. The RM metric reveals the domain-independent generation quality since if the generated video is more relevant to the text, it will obviously have a higher RM value. If the generated video is not relevant to the text or has a low quality, the RM value will be lower.

2.2 Human Evaluation Metrics

To conduct a human evaluation, we invite 200 evaluators as testees and conduct a human evaluation. Let {M1,M2,..,MN}\{M_{1},M_{2},..,M_{N}\} be a set of models to evaluate, TT be the number of samples in the test set. To reduce the subjective biases, we ask the testees to compare the Visual Realisticity (VRVR) and Semantic Consistency (SCSC) of two videos (vi,vjv_{i},v_{j}) generated from two models (Mi,MjM_{i},M_{j}) with the same query qq respectively, as denoted in Eq. (16)∼\sim(17).

3 Implementation details

In Sec. 3.1, the size of the input video is L=10,H=64,W=64,C=3L=10,H=64,W=64,C=3. Both the encoder EE in Eq. (1) and the decoder DD in Eq. (4) are implemented with two CNN layers. The kernel size is 4 and stride is 2. Thus the latent variable has the size of h×w=16×16h\times w=16\times 16. The latent variable dimension dB=128d_{B}=128. The VQ-VAE codebook has a total of K=10000K=10000 tokens. The VQ-VAE model is pretrained on ImageNet with a learning rate of 1e-3 and batch size 32. Note that when we conduct experiments on Moving Mnist Dataset, we train another VQ-VAE on this dataset. We found this will lead to better generation performance.

In Sec. 3.2, the input text has a maximum length of N=35N=35. The dimension D=1024D=1024. The maximum size of the visual tokens is M=2560M=2560. The Self-Att in Eq. (10) uses 16 attention heads. The GODIVA has a total of R=12R=12 layers in Eq. (11). The GODIVA model is pretrained on the Howto100M dataset with 64 V100 GPUs. It is finetuned on the MSR-VTT dataset with 8 V100 GPUs. Both settings have the same batch size of 32 and a learning rate of 5e-4. More details, including the source code, will come soon in Github.

4 Qualitative Results

We qualitatively evaluate our model from two aspects. Firstly, we evaluate the zero-shot ability of our model by comparing GODIVA to two prior approaches: T2V and TFGAN . Both approaches are trained on the real-world dataset created from a clean-up of Kinetics and Youtube videos. As shown in Fig. (3), for the same query "Play golf on grass", T2V successfully generates the grass and action of "playing golf" in a resolution of 64×\times64, but the result looks blurry (see the first row). TFGAN was successfully trained on a resolution of 128×\times128 and generates a higher quality result (see the second row). Both T2V and TFGAN generate text-related videos, but the generated frames are in a single scene and the difference between frames is not significant. This limits the creativity of neural models. Interestingly, GODIVA not only generates text-related videos but also changing scenes (see the third and fourth row). For example, GODIVA(64×\times64) first shows the grass field, then it gives the athlete a close-up shot, and finally the action of hitting the golf ball. Note that GODIVA(64×\times64) and GODIVA(128×\times128) are different models, thus they generate totally different videos. The last row gives another (128×\times128) resolution results generated by GODIVA. In total, GODIVA is able to generate videos with clear frames and coherent semantics.

Secondly, we evaluate the unseen video generation ability by comparing GODIVA to several GAN-based approaches. The models in Fig. (4) are trained on the Moving MNIST dataset and Double Moving MNIST dataset respectively. Note that there is no video in the training set that shows "Digit 9 is moving down than up", but there are some variant samples such as "Digit 9 is moving left and right" or "Digit 3 is moving down than up". GODIVA successfully generates semantic correct results(see the last row of the left part). This shows GODIVA learns to capture the semantic alignment between text and video, rather than just search videos in the training set to find the one most similar to the input sentence. Besides, GODIVA generates high-quality videos, even compared with the state-of-the-art IRC-GAN approach. The digit "9" is both spatially clear and temporally consistent. Another example in Double Moving MNIST on the right shows a similar phenomenon.

5 Quantitative Results

We quantitatively evaluate our model through both automatic and human metrics. To validate the effectiveness of RM metric, we first draw the SIM scores between text and ground-truth videos in Fig. 5(a). It can be seen from the diagonal that SIM is basically able to distinguish semantically similar videos from other videos. Tab. 1 shows the ablation experiments of different settings of GODIVA. We pretrain GODIVA on Howto100M dataset and finetune it on MSR-VTT dataset.We find that SIM and RM have the same trend with the human evaluation metrics. The first row shows the results between input text and ground-truth videos. The second row shows that sufficient scale for GODIVA is crucial. GODIVA (6 layer) shows worse performance than the default GODIVA setting (12 layer). The next three rows show the effectiveness of the three dimensional attentions. We found that the Row Attention is the most important. Following DALL-E , we randomly sample 32 times in the top 10 probabilities in Eq. (12) during inference and using CLIP ranking to find the best generated video. The performance is then significantly improved to 98.34 in RM metric.

Conclusions

In this paper, we propose a three-dimensional sparse attention to generate open-domain videos from natural descriptions using VQ-VAE discrete visual tokens. We also propose a new Relative Matching metric to automatically evaluate generation quality. Experiments show that our model not only can be fine-tuned on downstream video generation tasks, but also has a good zero-shot capability on unseen texts. However, there are still several challenges: Firstly, it is still a great challenge to generate long videos with high resolution. When generating only 64×6464\times 64 resolution videos with 10 frames, the total of visual tokens MM already becomes 2560. Secondly, automatically evaluating text-to-video generation task remains a challenge. In the future, video-based CLIP metric may give more accurate results for semantic consistency for text and videos. Thirdly, GAN-based methods show a great potential for text-to-video generation (see Fig. (4)), their generative abilities for open-domain dataset remain a good research direction.

References

Appendix