MoCoGAN: Decomposing Motion and Content for Video Generation

Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, Jan Kautz

Introduction

Deep generative models have recently received an increasing amount of attention, not only because they provide a means to learn deep feature representations in an unsupervised manner that can potentially leverage all the unlabeled images on the Internet for training, but also because they can be used to generate novel images necessary for various vision applications. As steady progress toward better image generation is made, it is also important to study the video generation problem. However, the extension from generating images to generating videos turns out to be a highly challenging task, although the generated data has just one more dimension – the time dimension.

We argue video generation is much harder for the following reasons. First, since a video is a spatio-temporal recording of visual information of objects performing various actions, a generative model needs to learn the plausible physical motion models of objects in addition to learning their appearance models. If the learned object motion model is incorrect, the generated video may contain objects performing physically impossible motion. Second, the time dimension brings in a huge amount of variations. Consider the amount of speed variations that a person can have when performing a squat movement. Each speed pattern results in a different video, although the appearances of the human in the videos are the same. Third, as human beings have evolved to be sensitive to motion, motion artifacts are particularly perceptible.

Recently, a few attempts to approach the video generation problem were made through generative adversarial networks (GANs) . Vondrick et al. hypothesize that a video clip is a point in a latent space and proposed a VGAN framework for learning a mapping from the latent space to video clips. A similar approach was proposed in the TGAN work . We argue that assuming a video clip is a point in the latent space unnecessarily increases the complexity of the problem, because videos of the same action with different execution speed are represented by different points in the latent space. Moreover, this assumption forces every generated video clip to have the same length, while the length of real-world video clips varies. An alternative (and likely more intuitive and efficient) approach would assume a latent space of images and consider that a video clip is generated by traversing the points in the latent space. Video clips of different lengths correspond to latent space trajectories of different lengths.

In addition, as videos are about objects (content) performing actions (motion), the latent space of images should be further decomposed into two subspaces, where the deviation of a point in the first subspace (the content subspace) leads content changes in a video clip and the deviation in the second subspace (the motion subspace) results in temporal motions. Through this modeling, videos of an action with different execution speeds will only result in different traversal speeds of a trajectory in the motion space. Decomposing motion and content allows a more controlled video generation process. By changing the content representation while fixing the motion trajectory, we have videos of different objects performing the same motion. By changing motion trajectories while fixing the content representation, we have videos of the same object performing different motion as illustrated in Fig. 1.

In this paper, we propose the Motion and Content decomposed Generative Adversarial Network (MoCoGAN) framework for video generation. It generates a video clip by sequentially generating video frames. At each time step, an image generative network maps a random vector to an image. The random vector consists of two parts where the first is sampled from a content subspace and the second is sampled from a motion subspace. Since content in a short video clip usually remains the same, we model the content space using a Gaussian distribution and use the same realization to generate each frame in a video clip. On the other hand, sampling from the motion space is achieved through a recurrent neural network where the network parameters are learned during training. Despite lacking supervision regarding the decomposition of motion and content in natural videos, we show that MoCoGAN can learn to disentangle these two factors through a novel adversarial training scheme. Through extensive qualitative and quantitative experimental validations with comparison to the state-of-the-art approaches including VGAN and TGAN , as well as the future frame prediction methods including Conditional-VGAN (C-VGAN) and Motion and Content Network (MCNET) , we verify the effectiveness of MoCoGAN.

Video generation is not a new problem. Due to limitations in computation, data, and modeling tools, early video generation works focused on generating dynamic texture patterns . In the recent years, with the availability of GPUs, Internet videos, and deep neural networks, we are now better positioned to tackle this intriguing problem.

Various deep generative models were recently proposed for image generation including GANs , variational autoencoders (VAEs) , and PixelCNNs . In this paper, we propose the MoCoGAN framework for video generation, which is based on GANs.

Multiple GAN-based image generation frameworks were proposed. Denton et al. showed a Laplacian pyramid implementation. Radford et al. used a deeper convolution network. Zhang et al. stacked two generative networks to progressively render realistic images. Coupled GANs learned to generate corresponding images in different domains, later extended to translate an image from one domain to a different domain in an unsupervised fashion . InfoGAN learned a more interpretable latent representation. Salimans et al. proposed several GAN training tricks. The WGAN and LSGAN frameworks adopted alternative distribution distance metrics for more stable adversarial training. Roth et al. proposed a special gradient penalty to further stabilize training. Karras et al. used progressive growing of the discriminator and the generator to generate high resolution images. The proposed MoCoGAN framework generates a video clip by sequentially generating images using an image generator. The framework can easily leverage advances in image generation in the GAN framework for improving the quality of the generated videos. As discussed in Section 1, extended the GAN framework to the video generation problem by assuming a latent space of video clips where all the clips have the same length.

Recurrent neural networks for image generation were previously explored in . Specifically, some works used recurrent mechanisms to iteratively refine a generated image. Our work is different to in that we use the recurrent mechanism to generate motion embeddings of video frames in a video clip. The image generation is achieved through a convolutional neural network.

The future frame prediction problem studied in is different to the video generation problem. In future frame prediction, the goal is to predict future frames in a video given the observed frames in the video. Previous works on future frame prediction can be roughly divided into two categories where one focuses on generating raw pixel values in future frames based on the observed ones , while the other focuses on generating transformations for reshuffling the pixels in the previous frames to construct future frames . The availability of previous frames makes future frame prediction a conditional image generation problem, which is different to the video generation problem where the input to the generative network is only a vector drawn from a latent space. We note that used a convolutional LSTM encoder to encode temporal differences between consecutive previous frames for extracting motion information and a convolutional encoder to extract content information from the current image. The concatenation of the motion and content information was then fed to a decoder to predict future frames.

2 Contributions

We propose a novel GAN framework for unconditional video generation, mapping noise vectors to videos.

We show the proposed framework provides a means to control content and motion in video generation, which is absent in the existing video generation frameworks.

We conduct extensive experimental validation on benchmark datasets with both quantitative and subjective comparison to the state-of-the-art video generation algorithms including VGAN and TGAN to verify the effectiveness of the proposed algorithm.

Generative Adversarial Networks

GANs consist of a generator and a discriminator. The objective of the generator is to generate images resembling real images, while the objective of the discriminator is to distinguish real images from generated ones.

In practice, (1) is solved by alternating gradient update.

Motion and Content Decomposed GAN

Learning.

1 Categorical Dynamics

Experiments

We conducted extensive experimental validation to evaluate MoCoGAN. In addition to comparing to VGAN and TGAN , both quantitatively and qualitatively, we evaluated the ability of MoCoGAN to generate 1) videos of the same object performing different motions by using a fixed content vector and varying motion trajectories and 2) videos of different objects performing the same motion by using different content vectors and the same motion trajectory. We then compared a variant of the MoCoGAN framework with state-of-the-art next frame prediction methods: VGAN and MCNET . Evaluating generative models is known to be a challenging task . Hence, we report experimental results on several datasets, where we can obtain reliable performance metrics:

Shape motion. The dataset contained two types of shapes (circles and squares) with varying sizes and colors, performing two types of motion: one moving from left to right, and the other moving from top to bottom. The motion trajectories were sampled from Bezier curves. There were 4,0004,000 videos in the dataset; the image resolution was 64×6464\times 64 and video length was 16.

Facial expression. We used the MUG Facial Expression Database for this experiment. The dataset consisted of 8686 subjects. Each video consisted of 5050 to 160160 frames. We cropped the face regions and scaled to 96×9696\times 96. We discarded videos containing fewer than 6464 frames and used only the sequences representing one of the six facial expressions: anger, fear, disgust, happiness, sadness, and surprise. In total, we trained on 1,2541,254 videos.

Tai-Chi. We downloaded 4,5004,500 Tai Chi video clips from YouTube. For each clip, we applied a human pose estimator and cropped the clip so that the performer is in the center. Videos were scaled to 64×6464\times 64 pixels.

Human actions. We used the Weizmann Action database , containing 8181 videos of 9 people performing 9 actions, including jumping-jack and waving-hands. We scaled the videos to 96×9696\times 96. Due to the small size, we did not conduct a quantitative evaluation using the dataset. Instead, we provide visual results in Fig. 1 and Fig. 4(a).

UCF101 . The database is commonly used for video action recognition. It includes 13,22013,220 videos of 101 different action categories. Similarly to the TGAN work , we scaled each frame to 85×6485\times 64 and cropped the central 64×6464\times 64 regions for learning.

The details of the network designs are given in the supplementary materials. We used ADAM for training, with a learning rate of 0.0002 and momentums of 0.5 and 0.999. Our code will be made public.

1 Video Generation Performance

We compared MoCoGAN to VGAN and TGANThe VGAN and TGAN implementations are provided by their authors. using the shape motion and facial expression datasets. For each dataset, we trained a video generation model and generated 256 videos for evaluation. The VGAN and TGAN implementations can only generate fixed-length videos (32 frames and 16 frames correspondingly). For a fair comparison, we generated 16 frames using MoCoGAN, and selected every second frame from the videos generated by VGAN, such that each video has 16 frames in total.

For quantitative comparison, we measured content consistency of a generated video using the Average Content Distance (ACD) metric. For shape motion, we first computed the average color of the generated shape in each frame. Each frame was then represented by a 3-dimensional vector. The ACD is then given by the average pairwise L2 distance of the per-frame average color vectors. For facial expression videos, we employed OpenFace , which outperforms human performance in the face recognition task, for measuring video content consistency. OpenFace produced a feature vector for each frame in a face video. The ACD was then computed using the average pairwise L2 distance of the per-frame feature vectors.

We computed the average ACD scores for the 256 videos generated by the competing algorithms for comparison. The results are given in Table 1. From the table, we found that the content of the videos generated by MoCoGAN was more consistent, especially for the facial expression video generation task: MoCoGAN achieved an ACD score of 0.201, which was almost 40% better than 0.322 of VGAN and 34% better than 0.305 of TGAN. Fig. 3 shows examples of facial expression videos for competing algorithms.

Furthermore, we compared with TGAN and VGAN by training on the UCF101 database and computing the inception score similarly to . Table 2 shows comparison results. In this experiment we used the same MoCoGAN model as in all other experiments. We observed that MoCoGAN was able to learn temporal dynamics better, due to the decomposed representation, as it generated more realistic temporal sequences. We also noted that TGAN reached the inception score of 11.85 with WGAN training procedure and Singular Value Clipping (SVC), while MoCoGAN showed a higher inception score 12.42 without such tricks, supporting that the proposed framework is more stable than and superior to the TGAN approach.

User study.

We conducted a user study to quantitatively compare MoCoGAN to VGAN and TGAN using the facial expression and Tai-Chi datasets. For each algorithm, we used the trained model to randomly generate 80 videos for each task. We then randomly paired the videos generated by the MoCoGAN with the videos from one of the competing algorithms to form 80 questions. These questions were sent to the workers on Amazon Mechanical Turk (AMT) for evaluation. The videos from different algorithms were shown in random order for a fair comparison. Each question was answered by 3 different workers. The workers were instructed to choose the video that looks more realistic. Only the workers with a lifetime HIT (Human Intelligent Task) approval rate greater than 95% participated in the user study.

We report the average preference scores (the average number of times, a worker prefers an algorithm) in Table 3. From the table, we find that the workers considered the videos generated by MoCoGAN more realistic most of the times. Compared to VGAN, MoCoGAN achieved a preference score of 84.2% and 75.4% for the facial expression and Tai-Chi datasets, respectively. Compared to TGAN, MoCoGAN achieved a preference score of 54.7% and 68.0% for the facial expression and Tai-Chi datasets, respectively. In Fig. 3, we visualize the facial expression and Tai-Chi videos generated by the competing algorithms. We find that the videos generated by MoCoGAN are more realistic and contained less content and motion artifacts.

Qualitative evaluation.

We conducted a qualitative experiment to demonstrate our motion and content decomposed representation. We sampled two content codes and seven motion codes, giving us 14 videos. Fig. 4(b) shows an example randomly selected from this experiment. Each row has the same content code, and each column has the same motion code. We observed that MoCoGAN generated the same motion sequences for two different content samples.

2 Categorical Video Generation

To evaluate the performance, we computed the ACD of the generated videos. A smaller ACD means the generated faces over the 96 frames were more likely to be from the same person. Note that the ACD reported in this subsection are generally larger than the ACD reported in Table 1, because the generated videos in this experiment are 6 times longer and contain 6 facial expressions versus 1. We also used the motion control score (MCS) to evaluate MoCoGAN’s capability in motion generation control. To compute MCS, we first trained a spatio-temporal CNN classifier for action recognition using the labeled training dataset. During test time, we used the classifier to verify whether the generated video contained the action. The MCS is then given by testing accuracy of the classifier. A model with larger MCS offers better control over the action category.

3 Image-to-video Translation

Conclusion

We presented the MoCoGAN framework for motion and content decomposed video generation. Given sufficient video training data, MoCoGAN automatically learns to disentangle motion from content in an unsupervised manner. For instance, given videos of people performing different facial expressions, MoCoGAN learns to separate a person’s identity from their expression, thus allowing us to synthesize a new video of a person performing different expressions, or fixing the expression and generating various identities. This is enabled by a new generative adversarial network, which generates a video clip by sequentially generating video frames. Each video frame is generated from a random vector, which consists of two parts, one signifying content and one signifying motion. The content subspace is modeled with a Gaussian distribution, whereas the motion subspace is modeled with a recurrent neural network. We sample this space in order to synthesize each video frame. Our experimental evaluation supports that the proposed framework is superior to current state-of-the-art video generation and next frame prediction methods.

References

Appendix A Network Architecture

Appendix B Additional Qualitative Results