A Good Image Generator Is What You Need for High-Resolution Video Synthesis
Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski, Xi Peng, Dimitris N. Metaxas, Sergey Tulyakov
Introduction
Video synthesis seeks to generate a sequence of moving pictures from noise. While its closely related counterpart—image synthesis—has seen substantial advances in recent years, allowing for synthesizing at high resolutions (Karras et al., 2017), rendering images often indistinguishable from real ones (Karras et al., 2019), and supporting multiple classes of image content (Zhang et al., 2019), contemporary improvements in the domain of video synthesis have been comparatively modest. Due to the statistical complexity of videos and larger model sizes, video synthesis produces relatively low-resolution videos, yet requires longer training times. For example, scaling the image generator of Brock et al. (2019) to generate videos requires a substantial computational budgetWe estimate that the cost of training a model such as DVD-GAN (Clark et al., 2019) once requires $30K.. Can we use a similar method to attain higher resolutions? We believe a different approach is needed.
There are two desired properties for generated videos: (i) high quality for each individual frame, and (ii) the frame sequence should be temporally consistent, i.e. depicting the same content with plausible motion. Previous works (Tulyakov et al., 2018; Clark et al., 2019) attempt to achieve both goals with a single framework, making such methods computationally demanding when high resolution is desired. We suggest a different perspective on this problem. We hypothesize that, given an image generator that has learned the distribution of video frames as independent images, a video can be represented as a sequence of latent codes from this generator. The problem of video synthesis can then be framed as discovering a latent trajectory that renders temporally consistent images. Hence, we demonstrate that (i) can be addressed by a pre-trained and fixed image generator, and (ii) can be achieved using the proposed framework to create appropriate image sequences.
To discover the appropriate latent trajectory, we introduce a motion generator, implemented via two recurrent neural networks, that operates on the initial content code to obtain the motion representation. We model motion as a residual between continuous latent codes that are passed to the image generator for individual frame generation. Such a residual representation can also facilitate the disentangling of motion and content. The motion generator is trained using the chosen image discriminator with contrastive loss to force the content to be temporally consistent, and a patch-based multi-scale video discriminator for learning motion patterns. Our framework supports contemporary image generators such as StyleGAN2 (Karras et al., 2019) and BigGAN (Brock et al., 2019).
We name our approach as MoCoGAN-HD (Motion and Content decomposed GAN for High-Definition video synthesis) as it features several major advantages over traditional video synthesis pipelines. First, it transcends the limited resolutions of existing techniques, allowing for the generation of high-quality videos at resolutions up to . Second, as we search for a latent trajectory in an image generator, our method is computationally more efficient, requiring an order of magnitude less training time than previous video-based works (Clark et al., 2019). Third, as the image generator is fixed, it can be trained on a separate high-quality image dataset. Due to the disentangled representation of motion and content, our approach can learn motion from a video dataset and apply it to an image dataset, even in the case of two datasets belonging to different domains. It thus unleashes the power of an image generator to synthesize high quality videos when a domain (e.g., dogs) contains many high-quality images but no corresponding high-quality videos (see Fig. 4). In this manner, our method can generate realistic videos of objects it has never seen moving during training (such as generating realistic pet face videos using motions extracted from images of talking people). We refer to this new video generation task as cross-domain video synthesis. Finally, we quantitatively and qualitatively evaluate our approach, attaining state-of-the-art performance on each benchmark, and establish a challenging new baseline for video synthesis methods.
Related Work
Video Synthesis. Approaches to image generation and translation using Generative Adversarial Networks (GANs) (Goodfellow et al., 2014) have demonstrated the ability to synthesize high quality images (Radford et al., 2016; Zhang et al., 2019; Brock et al., 2019; Donahue & Simonyan, 2019; Jin et al., 2021). Built upon image translation (Isola et al., 2017; Wang et al., 2018b), works on video-to-video translation (Bansal et al., 2018; Wang et al., 2018a) are capable of converting an input video to a high-resolution output in another domain. However, the task of high-fidelity video generation, in the unconditional setting, is still a difficult and unresolved problem. Without the strong conditional inputs such as segmentation masks (Wang et al., 2019) or human poses (Chan et al., 2019; Ren et al., 2020) that are employed by video-to-video translation works, generating videos following the distribution of training video samples is challenging. Earlier works on GAN-based video modeling, including MDPGAN (Yushchenko et al., 2019), VGAN (Vondrick et al., 2016), TGAN (Saito et al., 2017), MoCoGAN (Tulyakov et al., 2018), ProgressiveVGAN (Acharya et al., 2018), TGANv2 (Saito et al., 2020) show promising results on low-resolution datasets. Recent efforts demonstrate the capacity to generate more realistic videos, but with significantly more computation (Clark et al., 2019; Weissenborn et al., 2020). In this paper, we focus on generating realistic videos using manageable computational resources. LDVDGAN (Kahembwe & Ramamoorthy, 2020) uses low dimensional discriminator to reduce model size and can generate videos with resolution up to , while we decrease training cost by utilizing a pre-trained image generator. The high-quality generation is achieved by using pre-trained image generators, while the motion trajectory is modeled within the latent space. Additionally, learning motion in the latent space allows us to easily adapt the video generation model to the task of video prediction (Denton et al., 2017), in which the starting frame is given (Denton & Fergus, 2018; Zhao et al., 2018; Walker et al., 2017; Villegas et al., 2017b; a; Babaeizadeh et al., 2017; Hsieh et al., 2018; Byeon et al., 2018), by inverting the initial frame through the generator (Abdal et al., 2020), instead of training an extra image encoder (Tulyakov et al., 2018; Zhang et al., 2020).
Interpretable Latent Directions. The latent space of GANs is known to consist of semantically meaningful vectors for image manipulation. Both supervised methods, either using human annotations or pre-trained image classifiers (Goetschalckx et al., 2019; Shen et al., 2020), and unsupervised methods (Jahanian et al., 2020; Plumerault et al., 2020), are able to find interpretable directions for image editing, such as supervising directions for image rotation or background removal (Voynov & Babenko, 2020; Shen & Zhou, 2020). We further consider the motion vectors in the latent space. By disentangling the motion trajectories in an unsupervised fashion, we are able to transfer the motion information from a video dataset to an image dataset in which no temporal information is available.
Contrastive Representation Learning is widely studied in unsupervised learning tasks (He et al., 2020; Chen et al., 2020a; b; Hénaff et al., 2020; Löwe et al., 2019; Oord et al., 2018; Misra & Maaten, 2020). Related inputs, such as images (Wu et al., 2018) or latent representations (Hjelm et al., 2019), which can vary while training due to data augmentation, are forced to be close by minimizing differences in their representation during training. Recent work (Park et al., 2020) applies noise-contrastive estimation (Gutmann & Hyvärinen, 2010) to image generation tasks by learning the correspondence between image patches, achieving performance superior to that attained when using cycle-consistency constraints (Zhu et al., 2017; Yi et al., 2017). On the other hand, we learn an image discriminator to create videos with coherent content by leveraging contrastive loss (Hadsell et al., 2006) along with an adversarial loss (Goodfellow et al., 2014).
Method
In this section, we introduce our method for high-resolution video generation. Our framework is built on top of a pre-trained image generator (Karras et al., 2020a; b; Zhao et al., 2020a; b), which helps to generate high-quality image frames and boosts the training efficiency with manageable computational resources. In addition, with the image generator fixed during training, we can disentangle video motion from image content, and enable video synthesis even when the image content and the video motion come from different domains.
where and denote the hidden state and cell state respectively, and is a noise vector sampled from the normal distribution to model the motion diversity at timestamp .
Motion Disentanglement. Prior work (Tulyakov et al., 2018) applies as the motion code for the frame to be generated, while the content code is fixed for all frames. However, such a design requires a recurrent network to estimate the motion while preserving consistent content from the latent vector, which is difficult to learn in practice. Instead, we propose to use a sequence of motion residuals for estimating the motion trajectory. Specifically, we model the motion residual as the linear combination of a set of interpretable directions in the latent space (Shen & Zhou, 2020; Härkönen et al., 2020). We first conduct principal component analysis (PCA) on randomly sampled latent vectors from to get the basis . Then, we estimate the motion direction from the previous frame to the current frame by using and as follows:
where is a 2-layer MLP that serves as a mapping function.
2 Contrastive Image Discriminator
Content Matching. To learn content similarity between frames within a video, we use the image discriminator as a feature extractor and train it with a form of contrastive loss known as InfoNCE (Oord et al., 2018). The goal is that pairs of images with the same content should be close together in embedding space, while images containing different content should be far apart.
The choice of positive pairs in Eqn. 6 is specifically designed for cross-domain video synthesis, as videos of arbitrary content from the image domain is not available. In the case that images and videos are from the same domain, the positive and negative pairs are easier to obtain. We randomly select and augment two frames from a real video to create positive pairs sharing the same content, while the negative pairs contain augmented images from different real videos.
Full Objective. The overall loss function for training motion generator, video discriminator, and image discriminator is thus defined as:
Experiments
In this section, we evaluate the proposed approach on several benchmark datasets for video generation. We also demonstrate cross-domain video synthesis for various image and video datasets.
We conduct experiments on three datasets including UCF-101 (Soomro et al., 2012), FaceForensics (Rössler et al., 2018), and Sky Time-lapse (Xiong et al., 2018) for unconditional video synthesis. We use StyleGAN2 as the image generator. Training details can be found in Appx. B.
The quantitative results are shown in Tab. 2. Our method achieves state-of-the-art results for both IS and FVD, and outperforms existing works by a large margin. Interestingly, this result indicates that a well-trained image generator has learned to represent rich motion patterns, and therefore can be used to synthesize high-quality videos when used with a well-trained motion generator.
FaceForensics is a dataset containing news videos featuring various reporters. We use all the images from training videos, with a resolution of , to learn an image generator, and sequences of consecutive frames to train motion generator. Note that our network can generate even longer continuous sequences, e.g. frames (Fig. 12 in Appx.), though only frames are used for training.
We show the FVD between generated and real video clips ( frames in length) for different methods in Tab. 2. Additionally, we use the Average Content Distance (ACD) from MoCoGAN (Tulyakov et al., 2018) to evaluate the identity consistency for these human face videos. We calculate ACD values over videos. We also report the two metrics for ground truth (GT) videos. To get FVD of GT videos, we randomly sample two groups of real videos and compute the score. Our method achieves better results than TGANv2 (Saito et al., 2020). Both methods have low FVD values, and can generate complex motion patterns close to the real data. However, the much lower ACD value of our approach, which is close to GT, demonstrates that the videos it synthesizes have much better identity consistency than the videos from TGANv2. Qualitative examples in Fig. 2 illustrate different motions patterns learned from the dataset. Furthermore, we perform perceptual experiments using Amazon Mechanical Turk (AMT) by presenting a pair of videos from the two methods to users and asking them to select a more realistic one. Results in Tab. 2 indicate our method outperforms TGANv2 in 73.6% of the comparisons.
Sky Time-Lapse is a video dataset consisting of dynamic sky scenes, such as moving clouds. The number of video clips for training and testing is and , respectively. We resize images to and train the model to generate frames. We compare our methods with the two recent approaches of MDGAN (Xiong et al., 2018) and DTVNet (Zhang et al., 2020), which are specifically designed for this dataset. In Tab. 3, we report the FVD for all three methods. It is clear that our approach significantly outperforms the others. Example sequences are shown in Fig. 3.
2 Cross-Domain Video Generation
To demonstrate how our approach can disentangle motion from image content and transfer motion patterns from one domain to another, we perform several experiments on various datasets. More specifically, we use the StyleGAN2 model, pre-trained on the FFHQ (Karras et al., 2019), AFHQ-Dog (Choi et al., 2020), AnimeFaces (Branwen, 2019), and LSUN-Church (Yu et al., 2015) datasets, as the image generators. We learn human facial motion from VoxCeleb (Nagrani et al., 2020) and time-lapse transitions in outdoor scenes from TLVDB (Shih et al., 2013). In these experiments, a pair such as (FFHQ, VoxCeleb) indicates that we synthesize videos with image content from FFHQ and motion patterns from VoxCeleb. We generate videos with a resolution of and for FFHQ, for AFHQ-Dog and AnimeFaces, and for LSUN-Church. Qualitative examples for (FFHQ, VoxCeleb), (LSUN-Church, TLVDB), (AFHQ-Dog, VoxCeleb), and (AnimeFaces, VoxCeleb) are shown in Fig. 4, depicting high-quality and temporally consistent videos (more videos, including results with BigGAN as the image generator, are shown in the Appendix).
We also demonstrate how the motion and content are disentangled in Fig. 6 and Fig. 6, which portray generated videos with the same identity but performing diverse motion patterns, and the same motion applied to different identities, respectively. We show results from (AFHQ-Dog, VoxCeleb) (first two rows) and (AnimeFaces, VoxCeleb) (last two rows) in these two figures.
3 Ablation Analysis
4 Long Sequence Generation
Due to the limitation of computational resources, we train MoCoGAN-HD to synthesize consecutive frames. However, we can generate longer video sequences during inference by applying the following two ways.
Motion Generator Unrolling. For motion generator, we can run the LSTM decoder for more steps to synthesize long video sequences. In Fig. 7, we show a synthesized video example of frames using the model trained on the FaceForensics dataset. Our method is capable to synthesize videos with more frames than the number of frames used for training.
Motion Interpolation. We can do interpolation on the motion trajectory directly to synthesize long videos. Fig. 8 shows an interpolation example of -frame on (AFHQ-Dog, VoxCeleb) dataset.
Conclusion
In this work, we present a novel approach to video synthesis. Building on contemporary advances in image synthesis, we show that a good image generator and our framework are essential ingredients to boost video synthesis fidelity and resolution. The key is to find a meaningful trajectory in the image generator’s latent space. This is achieved using the proposed motion generator, which produces a sequence of motion residuals, with the contrastive image discriminator and video discriminator. This disentangled representation further extends applications of video synthesis to content and motion manipulation and cross-domain video synthesis. The framework achieves superior results on a variety of benchmarks and reaches resolutions unattainable by prior state-of-the-art techniques.
References
Appendix A Additional Details for the Framework
For BigGAN (Brock et al., 2019), we sample the latent code directly from the space of .
A.2 Additional Details for the Discriminators
A.2.2 Image Discriminator
Here we describe in more detail the image augmentation and memory bank techniques used for conducting contrastive learning.
Image Augmentation. We perform data augmentation on images to create positive and negative pairs. We normalize the images to and apply the following augmentation techniques.
Affine. We augment each image with an affine transformation defined with three random parameters: rotation , translation , and scale .
Brightness. We add a random value to all channels of each image.
Color. We add a random value to one randomly-selected channel of each image.
Cutout (DeVries & Taylor, 2017). We mask out pixels in a random subregion of each image to . Each subregion starts at a random point and with size , where and is the image resolution.
Flipping. We horizontally flip the image with the probability of .
Memory Bank. It has been shown that contrastive learning benefits from large batch-sizes and negative pairs (Chen et al., 2020b). To increase the number of negative pairs, we incorporate the memory mechanism from MoCo (He et al., 2020), which designates a memory bank to store negative examples. More specifically, we keep an exponential moving average of the image discriminator, and its output of fake video frames are buffered as negative examples. We use a memory bank with a dictionary size of .
Appendix B More Details for Experiments
We also train an unconditional BigGAN model on the FFHQ dataset using the public PyTorch codehttps://github.com/ajbrock/BigGAN-PyTorch. We train a model with resolution and select the last checkpoint as the image generator.
Training Time. We train each image generator for UCF-101, FaceForensics, Sky Time-lapse, and AFHQ-Dog in less than 2 days using 8 Tesla V100 GPUs. For FFHQ, AnimeFaces, and LSUN-Church, we use the released models with no training cost. The training time for video generators ranges from days depending on the datasets (Due to the memory issue, the training for generating videos with resolution of was done on 8 Quadro RTX 8000, with 5 days). The total training time for all the datasets is days and the estimated cost for training on Google Cloud is \sim$$2.3K.
Video Prediction. For video prediction, we predict consecutive frames, given the first frame from a test video clip as the input. We find the inverse latent code for by minimizing the following objective:
AMT Experiments. We present more details on the AMT experiments for different experimental settings and datasets. For each experiment, we run iterations to get the averaged score.
FaceForensics, Ours vs TGANv2. We randomly select videos from each method and ask users to select the better one from a pair of videos.
Sky Time-lapse, Ours vs DTVNet. We compare our method with DTVNet on the video prediction task. The testing set of Sky Time-lapse dataset includes short video clips. Considering that many of these video clips share similar content and are sampled from long videos, we select short videos with different content for testing. For these videos, we perform inversion (Eqn. 8) on the first frame to get the latent code and generate videos. For DTVNet, we use the first frame directly as input to produce their results. We ask users to chose the one with better video quality from a pair of videos generated by our method and DTVNet. The results shown in Tab. 8 demonstrate the clear advantage of our approach.
Cross-Domain Video Generation. We provide more details on the image and video datasets.
FFHQ (Karras et al., 2019) consists of high-quality face images at resolution with considerable variation in terms of age, ethnicity, and background.
AFHQ-Dog (Choi et al., 2020) contains high-quality dog images at resolution with both training and testing sets.
AnimeFaces (Branwen, 2019) includes anime face images at resolution.
LSUN-Church (Yu et al., 2015) includes in-the-wild church images at resolution.
VoxCeleb (Nagrani et al., 2020) consists of short clips of human speech, extracted from interview videos uploaded to YouTube.
TLVDB (Shih et al., 2013) includes time-lapse videos, covering a wide range of landscapes and cityscapes.
For the video datasets, we randomly select consecutive frames from training videos and select every other frame to form a 16-frame sequence for training.
Appendix C More Video Results
In this section, we provide more qualitative video results generated by our approach. We show the thumbnail from each video in the figures. Full resolution videos are in the supplementary material. We also provide an HTML page to visualize these videos.
UCF-101. In Fig. 10, we show videos generated by our approach on the UCF-101 dataset.
FaceForensics. In Fig. 10, we show the generated videos for FaceForensics. In Fig. 12 and Fig. 12, we show that our approach can generate long consecutive results, and frames respectively, even when trained with -frame clips. In Fig. 14, we demonstrate that our approach can generate diverse motion patterns using the same content code. In Fig. 14, we apply the same motion codes with different content to get the synthesized videos.
Sky Time-lapse. Fig. 16 shows the generated videos for the Sky Time-lapse dataset.
(FFHQ, VoxCeleb). Fig. 16, Fig. 18, and Fig. 18 present the generated videos that have motion patterns from VoxCeleb and content from FFHQ, with resolutions of , , and , respectively. We use BigGAN as the generator for Fig. 16 and StyleGAN2 for Fig. 18 and Fig. 18.
(AFHQ-Dog, VoxCeleb). Fig. 20 presents the generated videos that have motion patterns from VoxCeleb and content from AFHQ-Dog. The videos have a resolution of . In Fig. 20, we show the interpolation between every two frames to get longer sequences.
(AnimeFaces, VoxCeleb). Fig. 22 shows the generated videos that have motion patterns from VoxCeleb and content from AmimeFaces. The videos have a resolution of .
(LSUN-Church, TLVDB). Fig. 22 presents the generated videos that have time-lapse changing style from TLVDB and content from LSUN-Church.
Appendix E Limitations
Our framework requires a well-trained image generator for frame synthesis. In order to synthesize high-quality and temporally coherent videos, an ideal image generator should satisfy two requirements: R1. The image generator should synthesize high-quality images, otherwise the video discriminator can easily tell the generated videos as the image quality is different from the real videos. R2. The image generator should be able to generate diverse image contents to include enough motion modes for sequence modeling.
Example of R1. UCF-101 is a challenging dataset even for the training of an image generator. In Tab. 7, the StyleGAN2 model trained on UCF-101 has FID , which is much worse than the others. We hypothesis the reason is that UCF-101 dataset has many categories, but within each category, it includes relatively a small amount of videos and these videos share very similar content. Such observation is also discussed in DVDGAN (Clark et al., 2019). Although we can achieve state-of-the-art performance on UCF-101 dataset, the quality of the generated videos is not as good as other datasets (Fig. 10), and the quality of synthesized videos is still not close to real videos.
Example of R2. We test our method on BAIR Robot Pushing Dataset (Ebert et al., 2017). We train a StyleGAN2 image generator with using the frames from BAIR videos. The image generator has FID as . Based on the image generator, we train a video generation model that can synthesize frames. An example of synthesized video is shown in Fig. 24 (more videos are in the supplementary materials). We can see our method can successfully model shadow changing, the robot arm moving, but it struggles to decouple the robot arm from some small objects in the background, which we show analysis follows.
Inspired by previous work (Härkönen et al., 2020), we further investigate the latent space of the image generator by considering the information contained in each PCA component. Fig. 25 shows the percentage of total variance captured by top PCA components. The image generator on BAIR compresses most of the information on a few components. Specially, the top PCA components captures of the variance. In contrast, the latent space of the image generator trained on FFHQ (and FFHQ 1024 for high-resolution image synthesis) uses PCA components to capture information. This implies the BAIR generator models the dataset in a low-dimension space, and such generator increases the difficulty for fully disentangling all the objects in images for manipulation.
Moreover, we visualize the video synthesis results by moving along the top PCA components. Let denote the PCA component. Given content code , we synthesize a -frame video clip by using the following sequence as input: . In Fig. 26, we show the video synthesis results by moving along the top PCA directions. It can be seen that: 1) changing the later components (the and later rows) of BAIR only make small changes; 2) the first components of BAIR have entangled semantic meaning, while the components in FFHQ have more disentangled meaning ( row, rotation; row, smile). This indicates the image generator of BAIR may not cover enough (disentangled) motion modes, and it might be hard for the motion generator to fully disentangle all the contents and motion with only a few dominating PCA components, while for the image generator trained on FFHQ, it is much easier for disentangling foreground and background.