To Create What You Tell: Generating Videos from Captions
Yingwei Pan, Zhaofan Qiu, Ting Yao, Houqiang Li, Tao Mei
Introduction
Characterizing and modeling natural images and videos remains an open problem in computer vision and multimedia community. One fundamental issue that underlies this challenge is the difficulty to quantify the complex variations and statistical structures in images and videos. This motivates the recent studies to explore Generative Adversarial Nets (GANs) (Goodfellow et al., 2014) in generating plausible images (Denton et al., 2015; Radford et al., 2016). Nevertheless, a video is a sequence of frames which additionally contains temporal dependency, making it extremely hard to extend GANs to video domain. Moreover, as videos are often accompanied by text descriptors, e.g., tags or captions, learning video generative models conditioning on text then reduces sampling uncertainties and has a great potential real-world applications. Particularly, we are interested in producing videos from captions in this work, which is a brave new and timely problem. It aims to generate a video which is semantically aligned with the given descriptive sentence as illustrated in Figure 1.
In general, there are two critical issues in video generation employing caption conditioning: temporal coherence across video frames and semantic match between caption and the generated video. The former yields insights into the learning of generative model that the adjacent video frames are often visually and semantically coherent, and thus should be smoothly connected over time. This can be regarded as an intrinsic and generic property to produce a video. The later pursues a model with the capability to create realistic videos which are relevant to the given caption descriptions. As such, the conditioned treatment is taken into account, on one hand to create videos resembling the training data, and on the other, to regularize the generative capacity by holistically harnessing the relationship between caption semantics and video content.
By jointly consolidating the idea of temporal coherence and semantic match in translating text in the form of sentence into videos, this paper extends the recipe of GANs and presents a novel Temporal GANs conditioning on Caption (TGANs-C) framework for video generation, as shown in Figure 2. Specifically, sentence embedding encoded by the Long-Short Term Memory (LSTM) networks is concatenated to the noise vector as an input of the generator network, which produces a sequence of video frames by utilizing 3D convolutions. As such, temporal connections across frames are explicitly strengthened throughout the progress of video generation. In the discriminator network, in addition to determining whether videos are real or fake, the network must be capable of learning to align videos with the conditioning information. In particular, three discriminators are devised, including video discriminator, frame discriminator and motion discriminator. The former two classify realistic videos and frames from the generated ones, respectively, and also attempt to recognize the semantically matched video/frame-caption pairs from mismatched ones. The latter one is to distinguish the displacement between consecutive real or generated frames to further enhance temporal coherence. As a result, the whole architecture of TGANs-C is trained end-to-end by optimizing three losses, i.e., video-level and frame-level matching-aware loss to correct label of real or synthetic video/frames and align video/frames with correct caption, respectively, and temporal coherence loss to emphasize temporal consistency.
The main contribution of this work is the proposal of a new architecture, namely TGANs-C, which is one of the first effort towards generating videos conditioning on captions. This also leads to the elegant views of how to guarantee temporal coherence across generated video frames and how to align video/frame content with the given caption, which are the problems not yet fully understood in the literature. Through an extensive set of quantitative and qualitative experiments, we validate the effectiveness of our TGANs-C model on three different benchmarks.
Related Work
We briefly group the related work into two categories: natural image synthesis and video generation. The former draws upon research in synthesizing realistic images by utilizing deep generative models, while the latter investigates generating image sequence/video from scratch.
Image Synthesis. Synthesizing realistic images has been studied and analyzed widely in AI systems for characterizing the pixel level structure of natural images. There are two main directions on automatically image synthesis: Variational Auto-Encoders (VAEs) (Kingma and Welling, 2013) and Generative Adversarial Networks (GANs) (Goodfellow et al., 2014). VAEs is a directed graphical model which firstly constrains the latent distribution of the data to come from prior normal distribution and then generates new samples through sampling from this distribution. This direction is straightforward to train but introduce potentially restrictive assumptions about approximate posterior distribution, always resulting in overly smoothed samples. Deep Recurrent Attentive Writer (DRAW) (Karol et al., 2015) is one of the early works which utilizes VAEs to generate images with a spatial attention mechanism. Furthermore, Mansimov et al. extend this model to generate images conditioning on captions by iteratively drawing patches on a canvas and meanwhile attending to relevant words in the description (Mansimov et al., 2016).
GANs can be regarded as the generator network modules learnt with a two-player minimax game mechanism and has shown the distinct ability of producing plausible images (Denton et al., 2015; Radford et al., 2016). Goodfellow et al. propose the theoretical framework of GANs and utilize GANs to generate images without any supervised information in (Goodfellow et al., 2014). Although the earlier GANs offer a distinct and promising direction for image synthesis, the results are somewhat noisy and blurry. Hence, Laplacian pyramid is further incorporated into GANs in (Denton et al., 2015) to produce high quality images. Later in (Odena et al., 2016), GANs is expended with a specialized cost function for classification, named auxiliary classifier GANs (AC-GANs), for generating synthetic images with global coherence and high diversity conditioning on class labels. Recently, Reed et al. utilize GANs for image synthesis based on given text descriptions in (Reed et al., 2016), enabling translation from character level to pixel level.
Video Generation. When extending the existing generative models (e.g., VAEs and GANs) to video domain, very few works exploit such video generation from scratch task as both the spatial and temporal complex variations need to be characterized, making the problem very challenging. In the direction of VAEs, Mittal et al. employ Recurrent VAEs and an attention mechanism in a hierarchical manner to create a temporally dependent image sequence conditioning on captions (Mittal et al., 2016). For video generation with GANs, a spatio-temporal 3D deconvolutions based GANs is firstly proposed in (Vondrick et al., 2016) by untangling the scene’s foreground from the background. Most recently, the 3D deconvolutions based GANs is further decomposed into temporal generator consisting of 1D deconvolutional layers and image generator with 2D deconvolutional layers for video generation in (Saito and Matsumoto, 2016).
In short, our work in this paper belongs to video generation models capitalizing on adversarial learning. Unlike the aforementioned GANs-based approaches which mainly focus on video synthesis in an unconditioned manner, our research is fundamentally different in the way that we aim at generating videos conditioning on captions. In addition, we further improve video generation from the aspects of involving frame-level discriminator and strengthening temporal connections across frames.
Video Generation from Captions
The main goal of our Temporal GANs conditioning on Captions (TGANs-C) is to design a generative model with the ability of synthesizing a temporal coherent frame sequence semantically aligned with the given caption. The training of TGANs-C is performed by optimizing the generator network and discriminator network (video and frame discriminators which simultaneously judge synthetic or real and semantically mismatched or matched with the caption for video and frame) in a two-player minimax game mechanism. Moreover, the temporal coherence prior is additionally incorporated into TGANs-C to produce temporally coherent frame sequence in two different schemes. Therefore, the overall objective function of TGANs-C is composed of three components, i.e., video-level matching-aware loss to correct the label of real or synthetic video and align video with matched caption, frame-level matching-aware loss to further enhance the image reality and semantic alignment with the conditioning caption for each frame, and temporal coherence loss (i.e., temporal coherence constraint loss/temporal coherence adversarial loss) to exploit the temporal coherence between consecutive frames in unconditional/conditional scheme. The whole architecture of TGANs-C is illustrated in Figure 2.
The basic generative adversarial networks (GANs) consists of two networks: a generator network that captures the data distribution for synthesizing image and a discriminator network that distinguishes real images from synthetic ones. In particular, the generator network takes a latent variable randomly sampled from a normal distribution as input and produces a synthetic image . The discriminator network takes an image as input stochastically chosen (with equal probability) from real images or synthetic ones through and produces a probability distribution over the two image sources (i.e., synthetic or real). As proposed in (Goodfellow et al., 2014), the whole GANs can be trained in a two-player minimax game. Concretely, given an image example , the discriminator network is trained to minimize the adversarial loss, i.e., maximizing the log-likelihood of assigning correct source to this example:
where the indicator function if condition is true; otherwise . Meanwhile, the generator network is trained to maximize the adversarial loss in Eq.(1), targeting for maximally fooling the discriminator network with its generated synthetic images .
2. Temporal GANs Conditioning on Captions (TGANs-C)
In this section, we elaborate the architecture of our TGANs-C, the GANs based generative model consisting of two networks: a generator network for synthesizing videos conditioning on captions, and a discriminator network that simultaneously distinguishes real videos/frames from synthetic ones and aligns the input videos/frames with semantically matching captions. Moreover, two different schemes for modeling temporal coherence across frames are incorporated into TGANs-C for video generation.
2.2. Discriminator Network
The discriminator network is designed to enable three main abilities: (1) distinguishing real video from synthetic one and aligning video with the correct caption, (2) determining whether each frame is real/fake and semantically matched/mismatched with the conditioning caption, (3) exploiting the temporal coherence across consecutive real frames. To address the three crucial points, three basic discriminators are particularly devised:
Specifically, in the training epoch, we can easily obtain a set of real-synthetic video triplets according to the prior given captions, where each tuple consists of one synthetic video conditioning on given caption , one real video described by the same caption , and one real video described by different caption from . Therefore, three video-caption pairs are generated based on the caption and its corresponding video tuple: the synthetic and semantically matched pair , real and semantically matched pair , and another real but semantically mismatched pair . Each video-caption pair is then set as the input to the discriminator network , followed by three kinds of losses to be optimized and each for one discriminator accordingly.
By minimizing this loss over positive video-caption pair (i.e., ) and negative video-caption pairs (i.e., and ), the video discriminator is trained to not only recognize each real video from synthetic ones but also classify semantically matched video-caption pair from mismatched ones.
where , and denotes the -th frame in , and , respectively.
Temporal coherence loss. Temporal coherence is one generic prior for video modeling, which reveals the intrinsic characteristic of video that the consecutive video frames are usually visually and semantically coherent. To incorporate this temporal coherence prior into TGANs-C for video generation, we consider two kinds of schemes on the basis of motion discriminator .
(1) Temporal coherence constraint loss. Motivated by (Mobahi et al., 2009), the similarity of two consecutive frames can be directly defined according to the Euclidean distances between their frame-level tensors, i.e., the magnitude of motion tensor:
Then, given the real-synthetic video triplet, we characterize the temporal coherence of the synthetic video as a constraint loss by accumulating the Euclidean distances over every two consecutive frames:
Please note that the temporal coherence constraint loss is designed only for optimizing generator network . By minimizing this loss of synthetic video, the generator network is enforced to produce temporally coherent frame sequence.
(2) Temporal coherence adversarial loss. Different from the first scheme formulating temporal coherence as a monotonous constraint in an unconditional manner, we further devise an adversarial loss to flexibly emphasize temporal consistency conditioning on the given caption. Similar to frame discriminator , the motion tensor in motion discriminator is first augmented with embedded sentence representation . Next, such concatenated tensor representation is leveraged to measure the final probability of classifying the temporal dynamics between consecutive frames as real ones conditioning on the given caption. Thus, given the real-synthetic video triplet and the conditioning caption , the temporal coherence adversarial loss is measured as
where , and denotes the motion tensor in , and , respectively. By minimizing the temporal coherence adversarial loss, the temporal discriminator is trained to not only recognize the temporal dynamics across synthetic frames from real ones but also align the temporal dynamics with the matched caption.
2.3. Optimization
The overall training objective function of TGANs-C integrates the video-level matching-aware loss in Eq.(3), frame-level matching-aware loss in Eq.(4) and temporal coherence constraint loss/temporal coherence adversarial loss in Eq.(6)/Eq.(7). As our TGANs-C is a variant of the GANs architecture, we train the whole architecture in a two-player minimax game mechanism. For the discriminator network , we update its parameters according to the following overall loss
where is the set of real-synthetic video triplets, and denotes the discriminator network ’s overall adversarial loss in unconditional scheme (i.e., TGANs-C with temporal coherence Constraint loss (TGANs-C-C)) and conditional scheme (i.e., TGANs-C with temporal coherence Adversarial loss (TGANs-C-A)), respectively. By minimizing this term, the discriminator network is trained to classify both videos and frames with correct sources, and simultaneously align videos and frames with semantically matching captions. Moreover, for TGANs-C-A, the discriminator network is additionally enforced to distinguish the temporal dynamics across frames with correct sources and also align the temporal dynamics with the matched captions.
For the generator network , its parameters are adjusted with the following overall loss
where and denotes the generator network ’s overall adversarial loss in TGANs-C-C and TGANs-C-A, respectively. The generator network is trained to fool the discriminator network on videos/frames source prediction with its synthetic videos/frames and meanwhile align synthetic videos/frames with the conditioning captions. Moreover, for TGANs-C-C, the consecutive synthetic frames are enforced to be similar in an unconditional scheme, while for TGANs-C-A, it additionally aims to fool on temporal dynamics source prediction with the synthetic videos in a conditional scheme. The training process of TGANs-C is given in Algorithm 1.
3. Testing Epoch
After the optimization of TGANs-C, we can obtain the learnt generator network . Thus, given a test caption , the bi-LSTM is first utilized to contextually embed the input word sequence, followed by a LSTM-based encoder to achieve the sentence representation . The sentence representation is then concatenated with the random noise variable as in Eq.(2) and finally fed into the generator network to produce the synthetic video .
Experiments
We evaluate and compare our proposed TGANs-C with state-of-the-art approaches by conducting video generation task on three datasets of progressively increasing complexity: Single-Digit Bouncing MNIST GIFs (SBMG) (Mittal et al., 2016), Two-digit Bouncing MNIST GIFs (TBMG) (Mittal et al., 2016), and Microsoft Research Video Description Corpus (MSVD) (Chen and Dolan, 2011). The first two are recently released GIF-based datasets consisting of MNIST (LeCun et al., 1998) digits moving frames and the last is a popular video captioning benchmark of YouTube videos.
SBMG. Similar to priors works (Srivastava et al., 2015; Shi et al., 2015) in generating synthetic dataset, SBMG is produced by having single handwritten digit bouncing inside a frame. It is composed of 12,000 GIFs and every GIF is 16 frames long, which contains a single digit moving left-right or up-down. The starting position of the digit is chosen uniformly at random. Each GIF is accompanied with single sentence describing the digit and its moving direction, as shown in Figure 3(a).
TBMG. TBMG is an extended synthetic dataset of SBMG which contains two handwritten digits bouncing. The generation process is the same as SBMG and the two digits within each GIF move left-right or up-down separately. Figure 3(b) shows two exemplary GIF-caption pairs in TBMG.
MSVD. MSVD contains 1,970 video snippets collected from YouTube. There are roughly 40 available English descriptions per video. In experiments, we manually filter out the videos about cooking and generate a subset of 518 cooking videos. Following the settings in (Guadarrama et al., 2013), our cooking subset is split with 363 videos for training and 155 for testing. Since video generation is a challenging problem, we assembled this subset with cooking scenario to better diagnose pros and cons of models. We randomly select two examples from this subset and show them in Figure 3(c).
2. Experimental Settings
Parameter Settings. We uniformly sample frames for each GIF/video and each word in the sentence is represented as “one-hot” vector. The architecture of our TGANs-C is mainly developed based on (Radford et al., 2016; Reed et al., 2016). We resize all the GIFs/videos in three datasets with pixels. In particular, for sentence encoding, the dimension of the input and hidden layers in bi-LSTM and LSTM-based encoder are all set to 256. For the generator network , the dimension of random noise variable is 100 and the dimension of sentence embedding in generator network is 256. For the discriminator network , we set the size of video-level tensor in video discriminator as and the size of frame-level tensor in frame discriminator is as .
Implementation Details. We mainly implement our proposed method based on Theano (Al-Rfou et al., 2016), which is one of widely adopted deep learning frameworks. Following the standard settings in (Radford et al., 2016), we train our TGANs-C models on all datasets by utilizing Adam optimizer with a mini-batch size of 64. All weights were initialized from a zero-centered Normal distribution with standard deviation 0.02 and the slope of the leak was set to 0.2 in the LeakyReLU. We set the learning rate and momentum as 0.0002 and 0.9, respectively.
Evaluation Metric. For the quantitative evaluation of video generation, we adopt Generative Adversarial Metric (GAM) (Im et al., 2016) which can directly compare two generative adversarial models by having them engage in a “battle” against each other. Given two generative adversarial models and , two kinds of ratios between the discriminative scores of the two models are measured as:
where denotes the classification error rate and is the testing set. The test ratio shows which model generalizes better on test data and the sample ratio reveals which model can fool the other model more easily. Finally, the GAM evaluation metric judges the winner as:
3. Compared Approaches
To empirically verify the merit of our TGANs-C, we compared the following state-of-the-art methods.
(1) Synchronized Deep Recurrent Attentive Writer (Sync-DRAW) (Mittal et al., 2016): Sync-DRAW is a VAEs-based model for video generation conditioning on captions which utilizes Recurrent VAEs to model spatio-temporal relationship and a separate attention mechanism to capture local saliency.
(2) Generative Adversarial Network for Video (VGAN) (Vondrick et al., 2016): The original VGAN attempts to leverage the spatio-temporal convolutional architecture to design a GANs-based generative model for video generation in an unconditioned manner. Here we additionally incorporate the matching-aware loss into the discriminator network of basic VGAN and enable this baseline to generate videos conditioning on captions.
(3) Generative Adversarial Network with Character-Level Sentence encoder (GAN-CLS) (Reed et al., 2016): GAN-CLS is originally designed for image synthesis from text descriptions by utilizing DC-GAN and a hybrid character-level convolutional-recurrent neural network for text encoding. We directly extend this architecture by replacing 2D convolutions with 3D spatio-temporal convolutions for text-conditional video synthesis.
(4) Temporal GANs conditioning on Captions (TGANs-C) is our proposal in this paper which includes two runs in different schemes: TGANs-C with temporal coherence constraint loss (TGANs-C-C) and TGANs-C with temporal coherence adversarial loss (TGANs-C-A). Two slightly different settings of TGANs-C are named as TGANs-C1 and TGANs-C2. The former is trained with only video-level matching-aware loss, while the latter is more similar to TGANs-C that only excludes the temporal coherence loss.
4. Optimization Analysis
Different from the traditional discriminative models which have a particularly well-behaved gradient, our TGANs-C is optimized with a complex two-player minimax game. Hence, we depict the evolution of the generator network at the training stage to illustrate the convergence of our TGANs-C. Concretely, we randomly sample one random noise variable and caption before training, and then leverage them to produce synthetic videos via the generator networks of TGANs-C-A at different iterations on TBMG. As shown in Figure 4, the quality of synthetic videos does improve as the iterations increase. Specifically, after 9,000 iterations, the generater network consistently synthesizes plausible videos by reproducing the visual appearances and temporal dynamics of handwritten digits conditioning on the caption.
5. Qualitative Evaluation
We then visually examine the quality of the results and compare among our four internal TGANs-C runs on SBMG and TBMG datasets. The examples of generated videos are shown in Figure 5. Given the input sentence of “digit 3 is moving up and down” in Figure 5(a), all the four runs can interpret the temporal track of forming single-digit bouncing videos. TGANs-C1 which only judges real or fake on video level and aligns video with the caption performs the worst among all the models and the predicted frames tend to be blurry. By additionally distinguishing frame-level realness and optimizing frame-caption matching, TGANs-C2 is capable of producing videos in which each frame is clear but the shape of the digit sometimes changes over time. Compared to TGANs-C2, TGANs-C-C emphasizes the coherence across adjacent frames by further regularizing the similarity in between. As a result, the frames generated by TGANs-C-C are more consistent than TGANs-C2 particularly of the digit in the frames, but on the other hand, the temporal coherence constraint exploited in TGANs-C-C is in a brute-force manner, making the generated videos monotonous and not that real. TGANs-C-A, in comparison, is benefited from the mechanism of adversarially modeling temporal connections. The chance that a video is gradually formed as real is better.
Figure 5(b) shows the generated videos by our four TGANs-C runs conditioning on the caption of “digit 1 is left and right and digit 9 is up and down.” Similar to the observations on single-digit bouncing videos, the four runs could also model the temporal dynamics of two-digit bouncing scenarios. When taking temporal smoothness into account, the quality of the videos generated by TGANs-C-C and TGANs-C-A is enhanced, as compared to the videos produced by TGANs-C1 and TGANs-C2. In addition, TGANs-C-A generates more realistic videos than TGANs-C-C, verifying the effectiveness of learning temporal coherence in an adversarial fashion.
Next, we compare with the three baselines on MSVD dataset. In view that TGANs-C-A consistently performs the best in our internal comparisons, we refer to this run as TGANs-C in the following evaluations. The comparisons of generated videos by different approaches are shown in Figure 6. We can easily observe that the videos generated by our TGANs-C have higher quality compared to the other models. The created frames by Sync-DRAW are very blurry since VAEs are biased towards generating smooth frames and the method does not present all the objects in the frames. The approach of VGAN generates the frames which tend to be fairly sharp. However, the background of the frames is stationary as VGAN enforces a static background and moving foreground, making it vulnerable to produce videos with background movement. Compared to GAN-CLS which only involves video-level matching-aware discriminator, our TGANs-C takes the advantages of additionally exploring frame-level matching-aware discriminator and temporal coherence across frames, and thus generates more realistic videos.
6. Human Evaluation
To better understand how satisfactory are the videos generated from different methods, we also conducted a human study to compare our TGANs-C against three approaches, i.e., Sync-DRAW, VGAN and GAN-CLS. A total number of 30 evaluators (15 females and 15 males) from different education backgrounds, including computer science (8), management (4), business (4), linguistics (4), physical education (1), international trade (1) and engineering (8), are invited and a subset of 500 sentences is randomly selected from testing set of MSVD dataset for the subjective evaluation.
We show all the evaluators the four videos generated by each approach plus the given caption and ask them to rank all the videos from 1 to 4 (good to bad) with respect to the three criteria: 1) Reality: how realistic are these generated videos? 2) Relevance: whether the videos are relevant to the given caption? 3) Coherence: judge the temporal connection and readability of the videos. To make the annotation as objective as possible, the four generated videos conditioning on each sentence are assigned to three evaluators and the final ranking is averaged on the three annotations. Furthermore, we average the ranking on each criterion of all the generated videos by each method and obtain three metrics. Table 1 lists the results of the user study on MSVD dataset. Overall, our TGANs-C is clearly the winner across all the three criteria.
7. Quantitative Evaluation
To further quantitatively verify the effectiveness of our proposed model, we compare our TGANs-C with two generative adversarial baselines (i.e., VGAN and GAN-CLS) in terms of GAM evaluation metric on MSVD dataset. As the method of Sync-DRAW produces videos by VAEs-based architecture rather than generative adversarial scheme, it is excluded in this comparison. The quantitative results are summarized in Table 2. Overall, considering the “battle” between our TGANs-C and the other two baselines, the sample ratios are both less than one, indicating that TGANs-C can produce more authentic synthetic videos and fool the other two models more easily. The results basically verify the advantages of exploiting frame-level realness, frame-caption matching and the temporal coherence across adjacent frames for video generation. Moreover, when comparing between the two 3D-based baselines, GAN-CLS beats VGAN easily. This somewhat reveals the weakness of VGAN, where the architecture is devised with the brute-force assumption that the background is stationary and only foreground moves, making it hard to mimic the real-word videos with dynamic background. Another important observation is that for the “battle” between each two runs, the test ratio is consistently approximately equal to one. This assures that none of the discriminator networks in these runs is overfitted more than the other, i.e., the corresponding sample ratios are applicable and not biased for evaluating generative adversarial models.
Conclusions
Synthesizing images or videos will be crucial for the next generation of multimedia systems. In this paper, we have presented the Temporal GANs conditioning on Captions (TGANs-C) architecture, succeeded in generating videos that correspond to a given input caption. Our model expands on adversarial learning paradigm from three aspects. First, we extend 2D generator network to 3D for explicitly modeling spatio-temporal connections in videos. Second, in addition to naive discriminator network which only judges fake or real, ours further evaluate whether the generated videos or frames match the conditioning caption. Finally, to guarantee the adjacent frames coherently formed over time, the motion information between consecutive real or generated frames is taken into account in the discriminator network. Extensive quantitative and qualitative experiments conducted on three datasets validate our proposal and analysis. Moreover, our approach creates videos with better quality by a user study from 30 human subjects.
Future works will focus, first of all, on improving visual discriminability of our model, i.e., synthesize higher resolution videos. A promising route to explore will be that of decomposing the problem into several stages, where the shape or basic color based on the given caption is sketched in the primary stages and the advanced stages rectify the details of videos. Second, how to generate videos conditioning on open-vocabulary caption is expected. Last but not least, extending our framework to audio domain should be also interesting.
Acknowledgments. This work was supported in part by 973 Program under contract No. 2015CB351803 and NSFC under contract No. 61325009.