StyleVideoGAN: A Temporal Generative Model using a Pretrained StyleGAN
Gereon Fox, Ayush Tewari, Mohamed Elgharib, Christian Theobalt
Introduction
Recent advances of generative adversarial networks (GANs), notably StyleGAN [Karras et al.(2018)Karras, Aila, Laine, and Lehtinen, Karras et al.(2019)Karras, Laine, and Aila, Karras et al.(2020)Karras, Laine, Aittala, Hellsten, Lehtinen, and Aila], show impressive capabilities in learning manifolds of photorealistic images at high resolution (up to ). This is especially true for images of human faces. However, these improvements are only starting to carry over to the domain of videos: While existing methods for videos show promising results in modeling content and motion [Tulyakov et al.(2018)Tulyakov, Liu, Yang, and Kautz, Saito et al.(2020)Saito, Saito, Koyama, and Kobayashi, Weissenborn et al.(2020)Weissenborn, Uszkoreit, and Täckström, Munoz et al.(2021)Munoz, Zolfaghari, Argus, and Brox, Tian et al.(2021)Tian, Ren, Chai, Olszewski, Peng, Metaxas, and Tulyakov], they usually are subject to at least a subset of the following limitations: small spatial resolution ( ); spatial artifacts ; constrained motion ; necessity of large amounts of training data ; large computational cost for training (memory, time; see Table 2 and Section 4) .
To address these problems, we present a novel approach to unconditional video generation: Our goal is to learn a generator for nearly photorealistic, high-resolution videos (up to ), by training on a video dataset that contains no more than 10 minutes of video footage. In addition, we limit the training footage to depict only a single subject. However, we want the trained model can generate motion not only for the training subject, but for many different random subjects. As the proving ground for our method we chose the generation of portrait videos, because (1) portraits are an attractive target for animation, and (2) high-quality training data and StyleGAN models for this domain are readily available. To demonstrate that our method is also applicable to other domains very different from portrait, we also show how it can be applied to the generation of complex hand motion.
The ideas described so far already allow us to generate very high resolution videos with a minimal demand of training data and computational resources, by training a recurrent Wasserstein GAN [Arjovsky et al.(2017)Arjovsky, Chintala, and Bottou] on temporal volumes of 25 time steps, more than what most previous methods can afford. However, in order to have our model generate videos of longer duration, our generator needs to continue its output sequence beyond 25 time steps at test time. This can be achieved by making the generator a recurrent neural network (RNN). Previous work [Tian et al.(2021)Tian, Ren, Chai, Olszewski, Peng, Metaxas, and Tulyakov] has pointed out, though, that just using an RNN is not sufficient. We validate this observation by demonstrating that a vanilla RNN may tend to “loop”, i.e., repeat the same motion over and over. We address this problem with a novel “gradient angle penalty” term which successfully prevents looping. In summary, we make the following contributions:
We present a novel approach for unconditional video generation that is supervised in the latent space of a pretrained image generator, without having to render video frames at training time, leading to large savings in computational resources.
We present a novel “gradient angle penalty” loss that helps in the generation of videos that are longer than the temporal window seen at training time.
We demonstrate that our approach is applicable even to domains as complicated as hand motion, where there are more challenging articulations and self-occlusions.
Related Work
Several methods have been proposed for learning a generative model of videos [Vondrick et al.(2016)Vondrick, Pirsiavash, and Torralba, Saito et al.(2017)Saito, Matsumoto, and Saito, Denton and Birodkar(2017), Tulyakov et al.(2018)Tulyakov, Liu, Yang, and Kautz, Acharya et al.(2018)Acharya, Huang, Paudel, and Gool, Yushchenko et al.(2019)Yushchenko, Araslanov, and Roth, Clark et al.(2019)Clark, Donahue, and Simonyan, Saito et al.(2020)Saito, Saito, Koyama, and Kobayashi, Kahembwe and Ramamoorthy(2020), Weissenborn et al.(2020)Weissenborn, Uszkoreit, and Täckström, Munoz et al.(2021)Munoz, Zolfaghari, Argus, and Brox, Ye et al.(2020)Ye, Han, Lin, Guoqiang, and He, Menapace et al.(2021)Menapace, Lathuiliere, Tulyakov, Siarohin, and Ricci, Chai et al.(2020)Chai, Liu, Liu, Han, and He, Tian et al.(2021)Tian, Ren, Chai, Olszewski, Peng, Metaxas, and Tulyakov, Hong et al.(2021)Hong, Uh, and Byun, Wang et al.(2021)Wang, Bremond, and Dantcheva]. While such approaches show interesting results, they are limited to low spatial resolutions such as 128x128 [Tulyakov et al.(2018)Tulyakov, Liu, Yang, and Kautz, Munoz et al.(2021)Munoz, Zolfaghari, Argus, and Brox, Ye et al.(2020)Ye, Han, Lin, Guoqiang, and He, Wang et al.(2021)Wang, Bremond, and Dantcheva], 256x256 [Acharya et al.(2018)Acharya, Huang, Paudel, and Gool, Clark et al.(2019)Clark, Donahue, and Simonyan, Saito et al.(2020)Saito, Saito, Koyama, and Kobayashi, Hong et al.(2021)Hong, Uh, and Byun] or 512x512 [Kahembwe and Ramamoorthy(2020)]. An exception is the work of Tian et al\bmvaOneDot [Tian et al.(2021)Tian, Ren, Chai, Olszewski, Peng, Metaxas, and Tulyakov], which can generate videos at 1024x1024. Furthermore, most approaches struggle with generating realistic videos of long durations. We now discuss these existing approaches in more detail:
Saito et al\bmvaOneDot [Saito et al.(2017)Saito, Matsumoto, and Saito] presented an approach for video synthesis using Wasserstein GAN losses and a novel parameter clipping method. The network architecture, like ours, uses a shared image generator for each frame. However, the image generator is not pretrained, and the loss functions are defined on the final video space, and not the latent space of the image generator. Saito et al\bmvaOneDotdemonstrate results on videos upto 128x128 resolution. Tulyakov et al\bmvaOneDot [Tulyakov et al.(2018)Tulyakov, Liu, Yang, and Kautz] decompose the generation of videos into a content part and a motion part. Their “MoCoGAN” is trained in an unsupervised manner using motion and content discriminators. Yushchenko et al\bmvaOneDot [Yushchenko et al.(2019)Yushchenko, Araslanov, and Roth] formulated video generation by means of Markov Decision Processes and extended into the framework of Tulyakov et al\bmvaOneDot [Tulyakov et al.(2018)Tulyakov, Liu, Yang, and Kautz]. Acharya et al\bmvaOneDot [Acharya et al.(2018)Acharya, Huang, Paudel, and Gool] learned to progressively grow the generative model starting from low-resolution and short duration, which enabled the synthesis of videos at 256x256 resolution for the first time. Clark et al\bmvaOneDot [Clark et al.(2019)Clark, Donahue, and Simonyan] divided their discriminator into a spatial component and a temporal component, which also allows the generation of videos at resolutions up to 256x256 and duration up to 48 frames. More recently, Saito et al\bmvaOneDot [Saito et al.(2020)Saito, Saito, Koyama, and Kobayashi] proposed a memory-efficient approach for training that scales linearly with resolution. Instead of directly training on the full temporal window, it uses a stack of sub-generators that are trained on different temporal resolutions. Earlier sub-generators process high frame-rate input with low resolution information while the later sub-generators process low frame-rate input with high resolution information.
Weissenborn et al\bmvaOneDot [Weissenborn et al.(2020)Weissenborn, Uszkoreit, and Täckström] proposed an autoregressive video generation model that generalizes the Transformer architecture [Vaswani et al.(2017)Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, and Polosukhin] using a three-dimensional self-attention mechanism. To reduce computational complexity, images are produced as sequences of smaller, sub-scaled image slices, akin to [Menick and Kalchbrenner(2019)]. Munoz et al\bmvaOneDot [Munoz et al.(2021)Munoz, Zolfaghari, Argus, and Brox] modelled frames as points in a latent space, without any 3D convolutions. Their generator consists of a sequence generator and an image generator. Adversarial losses contain a 2D discriminator and a 3D discriminator.
Very recently, and most related to our work, Tian et al\bmvaOneDot [Tian et al.(2021)Tian, Ren, Chai, Olszewski, Peng, Metaxas, and Tulyakov] formulated video generation as the problem of finding a suitable trajectory through the latent space of a pretrained and fixed image generator [Karras et al.(2020)Karras, Laine, Aittala, Hellsten, Lehtinen, and Aila, Brock et al.(2019)Brock, Donahue, and Simonyan], for example StyleGAN. Despite this commonality and their ability to produce output at resolution as a result, there is a number of important differences between their approach and ours:
Their design inherently relies on the image generator being available for forward and backward passes at training time, which increases the required amounts of GPU memory and computation time immensely, compared to our approach.
Their method requires a diverse training set. When trained on a single video (like our method is), their results show very limited motion, as we demonstrate in Section 4
Method
We train a Wasserstein GAN [Arjovsky et al.(2017)Arjovsky, Chintala, and Bottou], consisting of generator and critic .
Note that although we focus on demonstrating our pipeline on portrait videos, no part of it other than the preprocessing step is inherently face-specific.
Generator
Our generator is a stack of 4 GRU cells that processs the “per-time-step randomness” . To initialize the GRU memory, we have the MLP “hallucinate” some memory contents for the first three cells, whereas the last is initialized with :
Only at test time, not at training time, do we forward to StyleGAN for rendering of actual video frames. Note that that might be considerably larger than the 25 time steps used at training time. In Section 4 we report numbers for .
Critic
Loss Terms & Training
By training we minimize , where is the usual WGAN-GP loss [Gulrajani et al.(2017)Gulrajani, Ahmed, Arjovsky, Dumoulin, and Courville] (with ) and is a novel gradient angle penalty (with ): When training our model only with the WGAN-GP loss we have observed (Section 4) that synthesizing videos for can lead to motion that seems to be “looping”, i.e. the same motion pattern is repeated over and over. Based on observations reported in previous work [Tian et al.(2021)Tian, Ren, Chai, Olszewski, Peng, Metaxas, and Tulyakov] we suspect that learns to simply ignore the “per-time-step randomness” and rely exclusively on , without modifying it much in the course of the sequence. There seems to be a tendency to make determine the entire course of the sequence, which makes looping very likely. To counteract this, we present a new loss formulation that makes sure that the gradient of the producer output with respect to is at least a certain fraction of the gradient with respect to . We set:
where is the normalized Euclidean distance between the last time step and the first time step generated by . The function here normalizes the components of the difference vector according to running statistics that are tracked during training, such that we can expect the distribution of each component to have mean 0 and variance 1. The angle will be close to if the output of depends mostly on , which we want to prevent. Unless stated otherwise, we trained our models with Adam [Kingma and Ba(2015)] for 350 epochs. We exponentially average the weights of the generator throughout training using a momentum of .
The offset trick
Results
Each method was trained on the training video, depicting only one actor, as our task demands. We then generated two sets of videos from each model: The “Short” set consists of 2048 videos that are as long as the temporal window the particular method considered at training time (see column “” in Table 2). The “Long” set consists of 128 video segments, that all have at least 128 frames. The technique by Munoz et al\bmvaOneDot is not able to produce samples longer than its training window, which is why Table 1 contains no numbers for this set. FID scores are computed on 8,000 frames randomly sampled from the reference and generated sets. FVD scores are computed on 2048 videos from each of the two sets, with the duration of the videos again equal to the default temporal window length of each method.
We also evaluate the temporal consistency of facial identity using a variant of the Average Content Distance (ACD) [Tulyakov et al.(2018)Tulyakov, Liu, Yang, and Kautz]: For each generated frame, we obtain the identity features from a popular facial recognition library [fac()] and compute the average L2-distances between all pairs of frames of the video. This score is then averaged over all generated videos.
Video Generation
Fig. 5 shows sequences generated by three different models, each for the respective training identity. As shown in Fig. 1 however, the “offset trick” allows us to generate motion for randomly sampled identities as well, even though our training datasets always contain only 1 actor. All videos are synthesized at a resolution of and even though our method was trained only on a temporal window of frames, we can easily generate videos that are much longer, e.g. frames. The quality of motion can only be judged in our supplemental video, not on paper.
Comparison to Previous Methods
We compare to several previous techniques by training them all on the same dataset and evaluating the metrics described above.
The methods by Saito et al\bmvaOneDot [Saito et al.(2020)Saito, Saito, Koyama, and Kobayashi] and Tulyakov et al\bmvaOneDot [Tulyakov et al.(2018)Tulyakov, Liu, Yang, and Kautz] generate realistic motion, but are very limited in terms of spatial resolution ( and respectively). As we show in our supplemental video, the results for Munoz et al\bmvaOneDot [Munoz et al.(2021)Munoz, Zolfaghari, Argus, and Brox] include strong structural artifacts. All three methods are not capable of generalizing their output to random identities after being trained on only one subject, i.e. their training set would have to be much larger to make them generate the diversity of output identities Tian et al\bmvaOneDot and our method achieve. We attempted to compare to further methods [Weissenborn et al.(2020)Weissenborn, Uszkoreit, and Täckström, Ye et al.(2020)Ye, Han, Lin, Guoqiang, and He], but could not get access to their code.
Tab. 1 reports Fréchet distances for all methods. We have also computed ACD scores for our method (random actors) and for Tian et al\bmvaOneDot: While our method averages at (over 5 models) Tian et al\bmvaOneDot achieve a score of 0.54. We attribute this large difference to the very limited facial motion generated by Tian et al\bmvaOneDot, that of course makes it much easier to preserve the identity. For visual impressions of this observation please see our supplemental video.
Evaluation of the Gradient Angle Penalty
Table 1 shows that while removing from our training objective slightly improves the scores for “short” samples, it considerably increases the FVD scores for “long” samples. However, since the primary purpose of is to prevent looping, which is maybe not effectively captured by FVD, we have recorded a very short training sequence (one single sentence, spoken three times, 20 seconds in total) that provoked strong looping artifacts in 19 out of 20 independently trained models if was absent, but led to looping only in 6 out of 20 models that were trained with the loss in place. This suggests that is indeed making looping artifacts much less likely.
Proof of concept: Hands
To demonstrate that our method should in principle be applicable to content categories other than talking faces, we have conducted a proof-of-concept experiment for hands: We recorded the right hand of a subject for 1 hour, performing various types of motions (like showing numerals or performing a set of gestures), resulting in a dataset of around 100k frames. The only constraint was for the hand to always turn the palm to the camera and to never leave the recording space. This dataset we successfully used for training a StyleGAN model and the corresponding pSp inverter, both for resolution . With these models available, we were able to train our temporal model with a temporal window of time steps, on several test sequences (each about 8000 frames). Results are shown in Fig. 1 and in our supplemental video.
Proof of concept: Cars
Conclusion
We have presented a temporal GAN for the unconditional generation of high-quality videos. Based on embedding the footage of only 1 actor into the latent space of StyleGAN, we are able to train our model with a minimal amount of resources and can nevertheless generate diverse motion of arbitrary length for a great number of random actors at high spatial resolution. Although these abilities also have their limitations (see supplemental), we hope that our work can pave the way for future innovations in video generation.
We thank the the authors of [Tian et al.(2021)Tian, Ren, Chai, Olszewski, Peng, Metaxas, and Tulyakov] for training their model on the training data we sent them. We also thank Pramod Rao for his invaluable support in conducting the experiments for our evaluation section. This work was supported by the ERC Consolidator Grant 4DReply (770784).