A Good Image Generator Is What You Need for High-Resolution Video Synthesis

Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski, Xi Peng, Dimitris N. Metaxas, Sergey Tulyakov

Introduction

Video synthesis seeks to generate a sequence of moving pictures from noise. While its closely related counterpart—image synthesis—has seen substantial advances in recent years, allowing for synthesizing at high resolutions (Karras et al., 2017), rendering images often indistinguishable from real ones (Karras et al., 2019), and supporting multiple classes of image content (Zhang et al., 2019), contemporary improvements in the domain of video synthesis have been comparatively modest. Due to the statistical complexity of videos and larger model sizes, video synthesis produces relatively low-resolution videos, yet requires longer training times. For example, scaling the image generator of Brock et al. (2019) to generate 256×256256\times 256 videos requires a substantial computational budgetWe estimate that the cost of training a model such as DVD-GAN (Clark et al., 2019) once requires >> $30K.. Can we use a similar method to attain higher resolutions? We believe a different approach is needed.

There are two desired properties for generated videos: (i) high quality for each individual frame, and (ii) the frame sequence should be temporally consistent, i.e. depicting the same content with plausible motion. Previous works (Tulyakov et al., 2018; Clark et al., 2019) attempt to achieve both goals with a single framework, making such methods computationally demanding when high resolution is desired. We suggest a different perspective on this problem. We hypothesize that, given an image generator that has learned the distribution of video frames as independent images, a video can be represented as a sequence of latent codes from this generator. The problem of video synthesis can then be framed as discovering a latent trajectory that renders temporally consistent images. Hence, we demonstrate that (i) can be addressed by a pre-trained and fixed image generator, and (ii) can be achieved using the proposed framework to create appropriate image sequences.

To discover the appropriate latent trajectory, we introduce a motion generator, implemented via two recurrent neural networks, that operates on the initial content code to obtain the motion representation. We model motion as a residual between continuous latent codes that are passed to the image generator for individual frame generation. Such a residual representation can also facilitate the disentangling of motion and content. The motion generator is trained using the chosen image discriminator with contrastive loss to force the content to be temporally consistent, and a patch-based multi-scale video discriminator for learning motion patterns. Our framework supports contemporary image generators such as StyleGAN2 (Karras et al., 2019) and BigGAN (Brock et al., 2019).

We name our approach as MoCoGAN-HD (Motion and Content decomposed GAN for High-Definition video synthesis) as it features several major advantages over traditional video synthesis pipelines. First, it transcends the limited resolutions of existing techniques, allowing for the generation of high-quality videos at resolutions up to 1024×10241024\times 1024. Second, as we search for a latent trajectory in an image generator, our method is computationally more efficient, requiring an order of magnitude less training time than previous video-based works (Clark et al., 2019). Third, as the image generator is fixed, it can be trained on a separate high-quality image dataset. Due to the disentangled representation of motion and content, our approach can learn motion from a video dataset and apply it to an image dataset, even in the case of two datasets belonging to different domains. It thus unleashes the power of an image generator to synthesize high quality videos when a domain (e.g., dogs) contains many high-quality images but no corresponding high-quality videos (see Fig. 4). In this manner, our method can generate realistic videos of objects it has never seen moving during training (such as generating realistic pet face videos using motions extracted from images of talking people). We refer to this new video generation task as cross-domain video synthesis. Finally, we quantitatively and qualitatively evaluate our approach, attaining state-of-the-art performance on each benchmark, and establish a challenging new baseline for video synthesis methods.

Related Work

Video Synthesis. Approaches to image generation and translation using Generative Adversarial Networks (GANs) (Goodfellow et al., 2014) have demonstrated the ability to synthesize high quality images (Radford et al., 2016; Zhang et al., 2019; Brock et al., 2019; Donahue & Simonyan, 2019; Jin et al., 2021). Built upon image translation (Isola et al., 2017; Wang et al., 2018b), works on video-to-video translation (Bansal et al., 2018; Wang et al., 2018a) are capable of converting an input video to a high-resolution output in another domain. However, the task of high-fidelity video generation, in the unconditional setting, is still a difficult and unresolved problem. Without the strong conditional inputs such as segmentation masks (Wang et al., 2019) or human poses (Chan et al., 2019; Ren et al., 2020) that are employed by video-to-video translation works, generating videos following the distribution of training video samples is challenging. Earlier works on GAN-based video modeling, including MDPGAN (Yushchenko et al., 2019), VGAN (Vondrick et al., 2016), TGAN (Saito et al., 2017), MoCoGAN (Tulyakov et al., 2018), ProgressiveVGAN (Acharya et al., 2018), TGANv2 (Saito et al., 2020) show promising results on low-resolution datasets. Recent efforts demonstrate the capacity to generate more realistic videos, but with significantly more computation (Clark et al., 2019; Weissenborn et al., 2020). In this paper, we focus on generating realistic videos using manageable computational resources. LDVDGAN (Kahembwe & Ramamoorthy, 2020) uses low dimensional discriminator to reduce model size and can generate videos with resolution up to 512×512512\times 512, while we decrease training cost by utilizing a pre-trained image generator. The high-quality generation is achieved by using pre-trained image generators, while the motion trajectory is modeled within the latent space. Additionally, learning motion in the latent space allows us to easily adapt the video generation model to the task of video prediction (Denton et al., 2017), in which the starting frame is given (Denton & Fergus, 2018; Zhao et al., 2018; Walker et al., 2017; Villegas et al., 2017b; a; Babaeizadeh et al., 2017; Hsieh et al., 2018; Byeon et al., 2018), by inverting the initial frame through the generator (Abdal et al., 2020), instead of training an extra image encoder (Tulyakov et al., 2018; Zhang et al., 2020).

Interpretable Latent Directions. The latent space of GANs is known to consist of semantically meaningful vectors for image manipulation. Both supervised methods, either using human annotations or pre-trained image classifiers (Goetschalckx et al., 2019; Shen et al., 2020), and unsupervised methods (Jahanian et al., 2020; Plumerault et al., 2020), are able to find interpretable directions for image editing, such as supervising directions for image rotation or background removal (Voynov & Babenko, 2020; Shen & Zhou, 2020). We further consider the motion vectors in the latent space. By disentangling the motion trajectories in an unsupervised fashion, we are able to transfer the motion information from a video dataset to an image dataset in which no temporal information is available.

Contrastive Representation Learning is widely studied in unsupervised learning tasks (He et al., 2020; Chen et al., 2020a; b; Hénaff et al., 2020; Löwe et al., 2019; Oord et al., 2018; Misra & Maaten, 2020). Related inputs, such as images (Wu et al., 2018) or latent representations (Hjelm et al., 2019), which can vary while training due to data augmentation, are forced to be close by minimizing differences in their representation during training. Recent work (Park et al., 2020) applies noise-contrastive estimation (Gutmann & Hyvärinen, 2010) to image generation tasks by learning the correspondence between image patches, achieving performance superior to that attained when using cycle-consistency constraints (Zhu et al., 2017; Yi et al., 2017). On the other hand, we learn an image discriminator to create videos with coherent content by leveraging contrastive loss (Hadsell et al., 2006) along with an adversarial loss (Goodfellow et al., 2014).

Method

In this section, we introduce our method for high-resolution video generation. Our framework is built on top of a pre-trained image generator (Karras et al., 2020a; b; Zhao et al., 2020a; b), which helps to generate high-quality image frames and boosts the training efficiency with manageable computational resources. In addition, with the image generator fixed during training, we can disentangle video motion from image content, and enable video synthesis even when the image content and the video motion come from different domains.

where h\mathbf{h} and c\mathbf{c} denote the hidden state and cell state respectively, and ϵt\epsilon_{t} is a noise vector sampled from the normal distribution to model the motion diversity at timestamp tt.

Motion Disentanglement. Prior work (Tulyakov et al., 2018) applies ht\mathbf{h}_{t} as the motion code for the frame to be generated, while the content code is fixed for all frames. However, such a design requires a recurrent network to estimate the motion while preserving consistent content from the latent vector, which is difficult to learn in practice. Instead, we propose to use a sequence of motion residuals for estimating the motion trajectory. Specifically, we model the motion residual as the linear combination of a set of interpretable directions in the latent space (Shen & Zhou, 2020; Härkönen et al., 2020). We first conduct principal component analysis (PCA) on mm randomly sampled latent vectors from Z\mathcal{Z} to get the basis V\mathbf{V}. Then, we estimate the motion direction from the previous frame zt−1\mathbf{z}_{t-1} to the current frame zt\mathbf{z}_{t} by using ht\mathbf{h}_{t} and V\mathbf{V} as follows:

where HH is a 2-layer MLP that serves as a mapping function.

2 Contrastive Image Discriminator

Content Matching. To learn content similarity between frames within a video, we use the image discriminator as a feature extractor and train it with a form of contrastive loss known as InfoNCE (Oord et al., 2018). The goal is that pairs of images with the same content should be close together in embedding space, while images containing different content should be far apart.

The choice of positive pairs in Eqn. 6 is specifically designed for cross-domain video synthesis, as videos of arbitrary content from the image domain is not available. In the case that images and videos are from the same domain, the positive and negative pairs are easier to obtain. We randomly select and augment two frames from a real video to create positive pairs sharing the same content, while the negative pairs contain augmented images from different real videos.

Full Objective. The overall loss function for training motion generator, video discriminator, and image discriminator is thus defined as:

Experiments

In this section, we evaluate the proposed approach on several benchmark datasets for video generation. We also demonstrate cross-domain video synthesis for various image and video datasets.

We conduct experiments on three datasets including UCF-101 (Soomro et al., 2012), FaceForensics (Rössler et al., 2018), and Sky Time-lapse (Xiong et al., 2018) for unconditional video synthesis. We use StyleGAN2 as the image generator. Training details can be found in Appx. B.

The quantitative results are shown in Tab. 2. Our method achieves state-of-the-art results for both IS and FVD, and outperforms existing works by a large margin. Interestingly, this result indicates that a well-trained image generator has learned to represent rich motion patterns, and therefore can be used to synthesize high-quality videos when used with a well-trained motion generator.

FaceForensics is a dataset containing news videos featuring various reporters. We use all the images from 704704 training videos, with a resolution of 256×256256\times 256, to learn an image generator, and sequences of 1616 consecutive frames to train motion generator. Note that our network can generate even longer continuous sequences, e.g. 6464 frames (Fig. 12 in Appx.), though only 1616 frames are used for training.

We show the FVD between generated and real video clips (1616 frames in length) for different methods in Tab. 2. Additionally, we use the Average Content Distance (ACD) from MoCoGAN (Tulyakov et al., 2018) to evaluate the identity consistency for these human face videos. We calculate ACD values over 256256 videos. We also report the two metrics for ground truth (GT) videos. To get FVD of GT videos, we randomly sample two groups of real videos and compute the score. Our method achieves better results than TGANv2 (Saito et al., 2020). Both methods have low FVD values, and can generate complex motion patterns close to the real data. However, the much lower ACD value of our approach, which is close to GT, demonstrates that the videos it synthesizes have much better identity consistency than the videos from TGANv2. Qualitative examples in Fig. 2 illustrate different motions patterns learned from the dataset. Furthermore, we perform perceptual experiments using Amazon Mechanical Turk (AMT) by presenting a pair of videos from the two methods to users and asking them to select a more realistic one. Results in Tab. 2 indicate our method outperforms TGANv2 in 73.6% of the comparisons.

Sky Time-Lapse is a video dataset consisting of dynamic sky scenes, such as moving clouds. The number of video clips for training and testing is 35,39235,392 and 2,8152,815, respectively. We resize images to 128×128128\times 128 and train the model to generate 1616 frames. We compare our methods with the two recent approaches of MDGAN (Xiong et al., 2018) and DTVNet (Zhang et al., 2020), which are specifically designed for this dataset. In Tab. 3, we report the FVD for all three methods. It is clear that our approach significantly outperforms the others. Example sequences are shown in Fig. 3.

2 Cross-Domain Video Generation

To demonstrate how our approach can disentangle motion from image content and transfer motion patterns from one domain to another, we perform several experiments on various datasets. More specifically, we use the StyleGAN2 model, pre-trained on the FFHQ (Karras et al., 2019), AFHQ-Dog (Choi et al., 2020), AnimeFaces (Branwen, 2019), and LSUN-Church (Yu et al., 2015) datasets, as the image generators. We learn human facial motion from VoxCeleb (Nagrani et al., 2020) and time-lapse transitions in outdoor scenes from TLVDB (Shih et al., 2013). In these experiments, a pair such as (FFHQ, VoxCeleb) indicates that we synthesize videos with image content from FFHQ and motion patterns from VoxCeleb. We generate videos with a resolution of 256×256256\times 256 and 1024×10241024\times 1024 for FFHQ, 512×512512\times 512 for AFHQ-Dog and AnimeFaces, and 256×256256\times 256 for LSUN-Church. Qualitative examples for (FFHQ, VoxCeleb), (LSUN-Church, TLVDB), (AFHQ-Dog, VoxCeleb), and (AnimeFaces, VoxCeleb) are shown in Fig. 4, depicting high-quality and temporally consistent videos (more videos, including results with BigGAN as the image generator, are shown in the Appendix).

We also demonstrate how the motion and content are disentangled in Fig. 6 and Fig. 6, which portray generated videos with the same identity but performing diverse motion patterns, and the same motion applied to different identities, respectively. We show results from (AFHQ-Dog, VoxCeleb) (first two rows) and (AnimeFaces, VoxCeleb) (last two rows) in these two figures.

3 Ablation Analysis

4 Long Sequence Generation

Due to the limitation of computational resources, we train MoCoGAN-HD to synthesize 1616 consecutive frames. However, we can generate longer video sequences during inference by applying the following two ways.

Motion Generator Unrolling. For motion generator, we can run the LSTM decoder for more steps to synthesize long video sequences. In Fig. 7, we show a synthesized video example of 6464 frames using the model trained on the FaceForensics dataset. Our method is capable to synthesize videos with more frames than the number of frames used for training.

Motion Interpolation. We can do interpolation on the motion trajectory directly to synthesize long videos. Fig. 8 shows an interpolation example of 3232-frame on (AFHQ-Dog, VoxCeleb) dataset.

Conclusion

In this work, we present a novel approach to video synthesis. Building on contemporary advances in image synthesis, we show that a good image generator and our framework are essential ingredients to boost video synthesis fidelity and resolution. The key is to find a meaningful trajectory in the image generator’s latent space. This is achieved using the proposed motion generator, which produces a sequence of motion residuals, with the contrastive image discriminator and video discriminator. This disentangled representation further extends applications of video synthesis to content and motion manipulation and cross-domain video synthesis. The framework achieves superior results on a variety of benchmarks and reaches resolutions unattainable by prior state-of-the-art techniques.

References

Appendix A Additional Details for the Framework

For BigGAN (Brock et al., 2019), we sample the latent code directly from the space of Z\mathcal{Z}.

A.2 Additional Details for the Discriminators

A.2.2 Image Discriminator

Here we describe in more detail the image augmentation and memory bank techniques used for conducting contrastive learning.

Image Augmentation. We perform data augmentation on images to create positive and negative pairs. We normalize the images to [−1,1]\left[-1,1\right] and apply the following augmentation techniques.

Affine. We augment each image with an affine transformation defined with three random parameters: rotation αr∈U(−180,180)\alpha_{r}\in\mathcal{U}(-180,180), translation αt∈U(−0.1,0.1)\alpha_{t}\in\mathcal{U}(-0.1,0.1), and scale αs∈U(0.95,1.05)\alpha_{s}\in\mathcal{U}(0.95,1.05).

Brightness. We add a random value αb∼U(−0.5,0.5)\alpha_{b}\sim\mathcal{U}(-0.5,0.5) to all channels of each image.

Color. We add a random value αc∼U(−0.5,0.5)\alpha_{c}\sim\mathcal{U}(-0.5,0.5) to one randomly-selected channel of each image.

Cutout (DeVries & Taylor, 2017). We mask out pixels in a random subregion of each image to . Each subregion starts at a random point and with size (αmH,αmW)(\alpha_{m}H,\alpha_{m}W), where αm∼U(0,0.25)\alpha_{m}\sim\mathcal{U}(0,0.25) and (H,W)(H,W) is the image resolution.

Flipping. We horizontally flip the image with the probability of 0.50.5.

Memory Bank. It has been shown that contrastive learning benefits from large batch-sizes and negative pairs (Chen et al., 2020b). To increase the number of negative pairs, we incorporate the memory mechanism from MoCo (He et al., 2020), which designates a memory bank to store negative examples. More specifically, we keep an exponential moving average of the image discriminator, and its output of fake video frames are buffered as negative examples. We use a memory bank with a dictionary size of 4,0964,096.

Appendix B More Details for Experiments

We also train an unconditional BigGAN model on the FFHQ dataset using the public PyTorch codehttps://github.com/ajbrock/BigGAN-PyTorch. We train a model with resolution 128×128128\times 128 and select the last checkpoint as the image generator.

Training Time. We train each image generator for UCF-101, FaceForensics, Sky Time-lapse, and AFHQ-Dog in less than 2 days using 8 Tesla V100 GPUs. For FFHQ, AnimeFaces, and LSUN-Church, we use the released models with no training cost. The training time for video generators ranges from 1.5∼31.5\sim 3 days depending on the datasets (Due to the memory issue, the training for generating videos with resolution of 1,024×1,0241,024\times 1,024 was done on 8 Quadro RTX 8000, with 5 days). The total training time for all the datasets is 1.5∼51.5\sim 5 days and the estimated cost for training on Google Cloud is 0.7K0.7K\sim$$2.3K.

Video Prediction. For video prediction, we predict consecutive frames, given the first frame x\mathbf{x} from a test video clip as the input. We find the inverse latent code z^1\hat{\mathbf{z}}_{1} for x1\mathbf{x}_{1} by minimizing the following objective:

AMT Experiments. We present more details on the AMT experiments for different experimental settings and datasets. For each experiment, we run 55 iterations to get the averaged score.

FaceForensics, Ours vs TGANv2. We randomly select 300300 videos from each method and ask users to select the better one from a pair of videos.

Sky Time-lapse, Ours vs DTVNet. We compare our method with DTVNet on the video prediction task. The testing set of Sky Time-lapse dataset includes 2,8152,815 short video clips. Considering that many of these video clips share similar content and are sampled from 148148 long videos, we select 148148 short videos with different content for testing. For these videos, we perform inversion (Eqn. 8) on the first frame to get the latent code and generate videos. For DTVNet, we use the first frame directly as input to produce their results. We ask users to chose the one with better video quality from a pair of videos generated by our method and DTVNet. The results shown in Tab. 8 demonstrate the clear advantage of our approach.

Cross-Domain Video Generation. We provide more details on the image and video datasets.

FFHQ (Karras et al., 2019) consists of 70,00070,000 high-quality face images at 1024×10241024\times 1024 resolution with considerable variation in terms of age, ethnicity, and background.

AFHQ-Dog (Choi et al., 2020) contains 5,2395,239 high-quality dog images at 512×512512\times 512 resolution with both training and testing sets.

AnimeFaces (Branwen, 2019) includes 2,232,4622,232,462 anime face images at 512×512512\times 512 resolution.

LSUN-Church (Yu et al., 2015) includes 126,227126,227 in-the-wild church images at 256×256256\times 256 resolution.

VoxCeleb (Nagrani et al., 2020) consists of 22,49622,496 short clips of human speech, extracted from interview videos uploaded to YouTube.

TLVDB (Shih et al., 2013) includes 463463 time-lapse videos, covering a wide range of landscapes and cityscapes.

For the video datasets, we randomly select 3232 consecutive frames from training videos and select every other frame to form a 16-frame sequence for training.

Appendix C More Video Results

In this section, we provide more qualitative video results generated by our approach. We show the thumbnail from each video in the figures. Full resolution videos are in the supplementary material. We also provide an HTML page to visualize these videos.

UCF-101. In Fig. 10, we show videos generated by our approach on the UCF-101 dataset.

FaceForensics. In Fig. 10, we show the generated videos for FaceForensics. In Fig. 12 and Fig. 12, we show that our approach can generate long consecutive results, 3232 and 6464 frames respectively, even when trained with 1616-frame clips. In Fig. 14, we demonstrate that our approach can generate diverse motion patterns using the same content code. In Fig. 14, we apply the same motion codes with different content to get the synthesized videos.

Sky Time-lapse. Fig. 16 shows the generated videos for the Sky Time-lapse dataset.

(FFHQ, VoxCeleb). Fig. 16, Fig. 18, and Fig. 18 present the generated videos that have motion patterns from VoxCeleb and content from FFHQ, with resolutions of 128×128128\times 128, 256×256256\times 256, and 1024×10241024\times 1024, respectively. We use BigGAN as the generator for Fig. 16 and StyleGAN2 for Fig. 18 and Fig. 18.

(AFHQ-Dog, VoxCeleb). Fig. 20 presents the generated videos that have motion patterns from VoxCeleb and content from AFHQ-Dog. The videos have a resolution of 512×512512\times 512. In Fig. 20, we show the interpolation between every two frames to get longer sequences.

(AnimeFaces, VoxCeleb). Fig. 22 shows the generated videos that have motion patterns from VoxCeleb and content from AmimeFaces. The videos have a resolution of 512×512512\times 512.

(LSUN-Church, TLVDB). Fig. 22 presents the generated videos that have time-lapse changing style from TLVDB and content from LSUN-Church.

Appendix E Limitations

Our framework requires a well-trained image generator for frame synthesis. In order to synthesize high-quality and temporally coherent videos, an ideal image generator should satisfy two requirements: R1. The image generator should synthesize high-quality images, otherwise the video discriminator can easily tell the generated videos as the image quality is different from the real videos. R2. The image generator should be able to generate diverse image contents to include enough motion modes for sequence modeling.

Example of R1. UCF-101 is a challenging dataset even for the training of an image generator. In Tab. 7, the StyleGAN2 model trained on UCF-101 has FID 45.6345.63, which is much worse than the others. We hypothesis the reason is that UCF-101 dataset has many categories, but within each category, it includes relatively a small amount of videos and these videos share very similar content. Such observation is also discussed in DVDGAN (Clark et al., 2019). Although we can achieve state-of-the-art performance on UCF-101 dataset, the quality of the generated videos is not as good as other datasets (Fig. 10), and the quality of synthesized videos is still not close to real videos.

Example of R2. We test our method on BAIR Robot Pushing Dataset (Ebert et al., 2017). We train a 64×6464\times 64 StyleGAN2 image generator with using the frames from BAIR videos. The image generator has FID as 6.126.12. Based on the image generator, we train a video generation model that can synthesize 1616 frames. An example of synthesized video is shown in Fig. 24 (more videos are in the supplementary materials). We can see our method can successfully model shadow changing, the robot arm moving, but it struggles to decouple the robot arm from some small objects in the background, which we show analysis follows.

Inspired by previous work (Härkönen et al., 2020), we further investigate the latent space of the image generator by considering the information contained in each PCA component. Fig. 25 shows the percentage of total variance captured by top PCA components. The image generator on BAIR compresses most of the information on a few components. Specially, the top 2020 PCA components captures 85%85\% of the variance. In contrast, the latent space of the image generator trained on FFHQ (and FFHQ 1024 for high-resolution image synthesis) uses 100100 PCA components to capture 85%85\% information. This implies the BAIR generator models the dataset in a low-dimension space, and such generator increases the difficulty for fully disentangling all the objects in images for manipulation.

Moreover, we visualize the video synthesis results by moving along the top 2020 PCA components. Let ViV_{i} denote the ithi^{th} PCA component. Given content code z1z_{1}, we synthesize a 55-frame video clip by using the following sequence as input: {z1−2Vi,z1−Vi,z1,z1+Vi,z1+2Vi}\{z_{1}-2V_{i},z_{1}-V_{i},z_{1},z_{1}+V_{i},z_{1}+2V_{i}\}. In Fig. 26, we show the video synthesis results by moving along the top 2020 PCA directions. It can be seen that: 1) changing the later components (the 8th8^{th} and later rows) of BAIR only make small changes; 2) the first 77 components of BAIR have entangled semantic meaning, while the components in FFHQ have more disentangled meaning (2nd2^{nd} row, rotation; 20th20^{th} row, smile). This indicates the image generator of BAIR may not cover enough (disentangled) motion modes, and it might be hard for the motion generator to fully disentangle all the contents and motion with only a few dominating PCA components, while for the image generator trained on FFHQ, it is much easier for disentangling foreground and background.