Deep Video-Based Performance Cloning

Kfir Aberman, Mingyi Shi, Jing Liao, Dani Lischinski, Baoquan Chen, Daniel Cohen-Or

Introduction

Wouldn’t it be great to be able to have a video of yourself moonwalking like Michael Jackson, or moving like Mick Jagger? Unfortunately, not many of us are capable of delivering such exciting and well-executed performances. However, perhaps it is possible to use an existing video of Jackson or Jagger to convert a video of an imperfect, or even a completely unrelated, reference performance into one that looks more like “the real deal”? In this paper, we attempt to provide a solution for such a scenario, which we refer to as video-driven performance cloning.

Motion capture is widely used to drive animation. However, in this scenario, the target to be animated is typically an articulated virtual character, which is explicitly rigged for animation, rather than a target individual captured by a simple RGB video. Video-driven performance cloning, which transfers the motion of a character between two videos, has also been studied. Research has mainly focused on facial performance reenactment, where speech, facial expressions, and sometimes changes in head pose, are transferred from one video to another (or a to a still image), while keeping the identity of the target face. To accomplish the transfer, existing works typically leverage dense motion fields and/or elaborate controllable parametric facial models, which are not yet available for the full human body. Cloning full body motion remains an open challenging problem, since the space of clothed and articulated humans is much richer and more varied than that of faces and facial expressions.

The emergence of neural networks and their startling performance and excellent generalization capabilities have recently driven researchers to develop generative models for video and motion synthesis. In this paper, we present a novel two-branch generative network that is trained to clone the full human body performance captured in a driving video to a target individual captured in a given reference video.

Given an ordinary video of a performance by the target actor, its frames are analyzed to extract the actor’s 2D body poses, each of which is represented as a set of part confidence maps, corresponding to the joints of the body skeleton [Cao et al., 2017]. The resulting pairs of video frames and the extracted poses are used to train a generative model to translate such poses to frames, in a self-supervised manner. Obviously, the reference video contains only a subset of the possible poses and motions that one might like to generate when cloning another performance. Thus, the key idea of our approach is to improve the generalization capability of our model by also training it with unpaired data. The unpaired data consists of pose sequences extracted from other relevant videos, which could depict completely different individuals.

The approach outlined above is realized as a deep network with two branches, both of which train the same space-time conditional generator (using shared weights). One branch uses paired training data with a combination of reconstruction and adversarial losses. The second branch uses unpaired data with a combination of temporal consistency loss and an adversarial loss. The adversarial losses used in this work are patchwise (Markovian) making it possible to generate new poses of the target individual by “stitching together” pieces of poses present in the reference video.

Following a review of related work in Section 2, we present are poses-to-video translation network in Section 3. Naturally, the success of our approach depends greatly on the richness of the paired training data (the reference video), the relevance of the unpaired pose sequences used to train the unpaired branch, and the degree of similarity between poses in the driving video and those which were used to train the network. In Section 4 we define a similarity metric between poses, which is suitable for our approach, and is able to predict how well a given driving pose might be translated into a frame depicting the target actor.

Section 5 describes and reports some qualitative and quantitative results and comparisons. Several performance cloning examples are included in the supplementary video. The cloned performances are temporally coherent, and preserve the target actor’s identity well. Our results exhibit some of the artifacts typical of GAN-generated imagery. Nevertheless, we believe that they are quite promising, especially considering some of the challenging combinations of reference and driving videos that they have been tested with.

Related Work

Conditional GANs [Mirza and Osindero, 2014] have proven effective for image-to-image translation tasks which are aimed at converting between images from two different domains. The structure of the main generator often incorporates an encoder-decoder architecture [Hinton and Salakhutdinov, 2006], with skip connections [Ronneberger et al., 2015], and adversarial loss functions [Goodfellow et al., 2014]. Isola et al. demonstrate impressive image translation results obtained by such an architecture, trained using paired data sets. Corresponding image pairs between different domains are often not available, but in some cases it is possible to use a translation operator that given an image in one domain can automatically produce a corresponding image from another. Such self-supervision have been successfully used for tasks such as super-resolution [Ledig et al., 2016], colorization [Zhang et al., 2016], etc. Recent works show that it is even possible to train high-resolution GANs [Karras et al., 2017] as well as conditional GANs [Wang et al., 2017].

However, in many tasks, the requirement for paired training data poses a significant obstacle. This issue is tackled by CycleGAN [Zhu et al., 2017], DualGAN [Yi et al., 2017], UNIT [Liu et al., 2017], and Hoshen and Wolf . These unsupervised image-to-image translation techniques only require two sets of unpaired training samples.

In this work, we are concerned with translation of 2D human pose sequences into photorealistic videos depicting a particular target actor. In order to achieve this goal we propose a novel combination of self-supervised paired training with unpaired training, both of which make use of (different) customized loss functions.

Human Pose Image Generation

Recently, a growing number of works aim to synthesize novel poses or views of humans, while preserving their identity. Zhao et al. learn to generate multiple views of a clothed human from only a single view input by combining variational inference and GANs. This approach is conditioned by the input image and the target view, and does not provide control over the articulation of the person.

In contrast, Ma et al. [2017a] condition their person image generation model on both a reference image and the specified target pose. Their two-stage architecture, trained on pairs of images capturing the same person in different poses, first generates a coarse image with the target pose, which is then refined. Siarohin et al. attempt to improve this approach and extend it to other kinds of deformable objects for which sufficiently many keypoints that capture the pose may be extracted.

The above works require paired data, do not address temporal coherence, and in some cases their ability to preserve identity is limited, because of reliance on a single conditional image of the target individual. In contrast, our goal is to generate a temporally coherent photorealistic video that consistently preserves the identity of a target actor. Our approach combines self-supervised training using paired data extracted from a video of the target individual with unpaired data training, and explicitly accounts for temporal coherence.

Ma et al. [2017b] propose to learn a representation that disentangles background, foreground, and pose. In principle, this enables modifying each of these factors separately. However, new latent features are generated by sampling from Gaussian noise, which is not suitable for our purpose and does not consider temporal coherence.

Deep Video Generation

Video generation is a highly challenging problem, which has recently been tackled with deep generative networks. Vondrick et al. propose a generative model that predicts plausible futures from static images using a GAN consisting of 3D deconvolutional layers, while Saito et al. use a temporal GAN that decomposes the video generation process into a temporal generator and an image generator. These methods do not support detailed control of the generated video, as required in our scenario.

Walker et al. attempt to first predict the motion of the humans in an image, and use the inferred future poses as conditional information for a GAN that generates the frame pixels. In contrast, we are not concerned with future pose prediction; the poses are obtained from a driving video, and we focus instead on the generation of photorealistic frames of the target actor performing these poses.

The MoCoGAN [Tulyakov et al., 2017] video generation framework maps random vectors to video frames, with each vector consisting of a content part and a motion part. This allows one to generate videos with the same motion, but different content. In this approach the latent content subspace is modeled using a Gaussian distribution, which does not apply to our scenario, where the goal is to control the motion of a specific target individual.

Performance Reenactment

Over the years, multiple works addressed video-based facial reenactment, where pose, speech, and/or facial expressions are transferred from a source performance to an image or a video of a target actor, e.g., [Averbuch-Elor et al., 2017; Suwajanakorn et al., 2015; Thies et al., 2016; Vlasic et al., 2005; Dale et al., 2011; Kemelmacher-Shlizerman et al., 2010; Garrido et al., 2014]. These works typically rely on dense motion fields or leverage the availability of detailed and fully controllable elaborate parametric facial models. Such models are not yet available for full human bodies, and we are not aware of methods capable of performing full body motion reenactment from video data.

Methods for performance driven animation [Xia et al., 2017], use a video or a motion-captured performance to animate 3D characters. In this scenario, the target is typically an articulated 3D character, which is already rigged for animation. In contrast, our target is a real person captured using plain video.

Poses-to-Video Translation Network

The essence of our approach is to learn how to translate coherent sequences of 2D human poses into video frames, depicting an individual whose appearance closely resembles a target actor captured in a provided reference video. Such a translator might be used, in principle, to clone arbitrary video performances of other individuals by extracting their human poses from various driving videos and feeding them into the translator.

More specifically, we are given a reference video, capturing a range of motions of a target actor, whom we’d like to use in order to clone a performance of another individual, as well as one or more videos of other individuals. We first train a deep neural network using paired data: pairs of consecutive frames and their corresponding poses extracted from the reference video. Since any given reference video contains only a subset of the poses and the motions that one might like to be able to generate, we gradually improve the generalization capability of this network by training it with unpaired data, consisting of pairs of consecutive poses extracted from videos of other individuals, while requiring temporal coherency between consecutive frames. An overview of our architecture is shown in Figure 2.

Let XX and YY denote the domains of video frames and human poses, respectively, and let P:X→Y{\bf P}:X\rightarrow Y be an operator which extracts the pose of a human in a given image. We first generate the paired data set that consists of frames of a target actor, xi∈Xx_{i}\in X, and their corresponding estimated poses, P(xi)∈Y{\bf P}(x_{i})\in Y. In addition, we are given some additional sequences of poses, yi∈Yy_{i}\in Y, which were extracted from additional videos. While the frames from which these poses were extracted are available as well, they typically capture other individuals, and frames of the target individual are not available for these poses.

The goal of our network, described in more detail in the next section, is to learn the mapping between arbitrary pose sequences to videos that appear to depict the target performer. The mapping is first learned using the paired data, and then extended to be able to translate the unpaired poses as well, in a temporally coherent manner.

2. Network Architecture

The core of our network consists of two branches, both of which train the same (shared weights) conditional generator module G{\bf G}, as depicted in Figure 2. The architecture of the generator is adapted from Isola et al. to accept a pose volume as input. The first, paired data branch (blue) of our network is aimed at learning the inverse mapping P−1{\bf P}^{-1} from the target actor poses to the reference video frames, supervised using paired training data. The second, unpaired data branch (red) utilizes temporal consistency and adversarial losses in order to learn to translate unpaired pose sequences extracted from one or more additional videos into temporally coherent frames that are ideally indistinguishable from those contained in the reference video.

where ϕ(⋅)\phi(\cdot) denotes the VGG-extracted features. In practice, we use the features from the relu1_1 layer, which has the same spatial resolution as the input images. Minimizing the distance between these low-level VGG features implicitly requires similarity of the underlying native pixel values, but also forces G{\bf G} to better reproduce small scale patterns and textures, compared to a metric defined directly using RGB pixels

For the second part of the loss we employ a Markovian discriminator [Isola et al., 2016; Li and Wand, 2016], Dp(y,x){\bf D}_{\text{p}}(y,x), trained to tell whether frame xx conditioned by pose y=P(x)y={\bf P}(x) is real or fake. This further encourages G{\bf G} to generate frames indistinguishable from the target actor in the provided input pose. The discriminator is trained to minimize a standard conditional GAN loss:

while the generator is trained to minimize the negation of the second term above. The two terms together constitute the reconstruction loss of the paired branch with

The purpose of this branch is to improve the generalization capability of the generator. The goal is to ensure that the generator is able to translate sequences of unseen poses, extracted from a driving video, into sequences of frames that look like reference video frames, and are temporally coherent.

The training data in this branch is unpaired, meaning that we do not provide a ground truth frame for each pose. Instead, the input consists of temporally coherent sequences of poses, extracted from one or more driving videos. The loss used to train this branch is a combination of a temporal coherence loss and a Markovian adversarial loss (PatchGAN) [Isola et al., 2016; Li and Wand, 2016].

The temporal coherence loss is computed using optical flow fields extracted from pairs of successive driving video frames. Ideally, we’d like the flow fields between pairs of successive generated frames to be the same, indicating that their temporal behavior follows that of the input driving sequence. However, we found that comparing smooth and textureless optical flow fields is not sufficiently effective for training. Similarly to Chen et al. , we enforce temporal coherency by using the original optical flow fields to warp the generated frames, and comparing the warped results to the generated consecutive ones.

More formally, let yi,yi+1y_{i},y_{i+1} denote a pair of consecutive poses, and fif_{i} denote the optical flow field between the consecutive frames that these poses were extracted from. The temporal coherence loss is defined as:

where W(x,f){\bf W}(x,f) is the frame obtained by warping xx with a flow field ff, using bi-linear interpolation. Due to potential differences in body size between the driving and target actors, the differences between the frames are weighted by the spatial map α(yi+1)\alpha(y_{i+1}), defined by a Gaussian falloff from the limbs connecting the joint locations. In other words, the differences are weighted more strongly near the skeleton, where we want to guarantee smooth motion.

Similarly to the paired branch, here we also apply an adversarial loss. However, here the discriminator, Ds(x){\bf D}_{\text{s}}(x), receives a single input and its goal is to tell whether it is a real frame from the reference video, or a generated one, without being provided the condition yy:

In addition, another goal of Ds{\bf D}_{\text{s}} is to prevent over smoothing effects that might exist due to the temporal loss. Because the poses come from other videos, they may be very different from the ones in the reference video. However, we assume that, given a sufficiently rich reference video, the generated frames should consist from pieces (patches) from the target video frames. Thus, we use a PatchGAN discriminator [Isola et al., 2016; Li and Wand, 2016], which determines whether a frame is real by examining patches, rather than the entire frame.

The two branches constitute our complete network, whose total loss is given by

where λs\lambda_{\text{s}} is a global weight for the unpaired branch, and λtc\lambda_{\text{tc}} weighs the temporal coherency term.

As mentioned earlier, the architecture of the generator G{\bf G} is adapted from Isola et al. , by extending it to take pose sequences as input, instead of images. The input tensor is of size JN×H×WJN\times H\times W, where NN is the number of consecutive poses and the output tensor is 3N×H×W3N\times H\times W. For all the results in this paper we used N=2N=2 (pairs of poses). The network progressively downsamples the input tensor, encoding it into a compact latent code and then decodes it into a reconstructed frame. Both of the discriminators, Dp{\bf D}_{\text{p}} and Ds{\bf D}_{\text{s}} share the same PatchGAN architecture [Isola et al., 2016], but were extended to support space time volumes. The only difference between the two discriminators is that Dp{\bf D}_{\text{p}} is extended to receive N(3+J)N(3+J) input channels while Ds{\bf D}_{\text{s}} receives 3N3N input channels, since it does not receive a conditional pose with each frame.

In the first few epochs, the training is only done by the paired branch using paired training data (Section 3.1). In this stage, we set λs=0\lambda_{s}=0, letting the network learn the appearance of the target actor, conditioned by a pose. Next, the second branch is enabled and joins the training process using unpaired training data. Here, we set λs=0.1\lambda_{s}=0.1 and λtc=10\lambda_{\text{tc}}=10. This stage ensures that the network learns to translate new motions containing unseen poses into smooth and natural video sequences.

We use the Adam optimizer [Kinga and Adam, 2015] and a batch size of 8. Half of each batch consists of paired data, while the other half is unpaired. Each half is fed into the corresponding network branch. The training data typically consists of about 3000 frames. We train the network for 150 epochs using a learning rate of 0.0002 and first momentum of 0.999. The network is trained from scratch for each individual target actor where the weights are randomly initialized based on a Normal distribution N(0,0.2)\mathcal{N}(0,0.2).

Typically, 3000 frames of a target actor video, are sufficient to train our network. However, the amount of data is only one of the factors that one should consider. A major factor that influences the results is our parametric model. Unlike face models [Thies et al., 2016] which explicitly control many aspects, such as expression and eye gaze, our model only specifies the positions of joints. Many details, such as hands, shoes, face, etc., are not explicitly controlled by the pose. Thus, given a new unseen pose, the network must generate these details based on the connections between pose and appearance that has been seen in the paired training data. Hence, unseen input poses that are similar to those that exist in the paired data set, will yield better performance. In the next section, we propose a new metric suitable for measuring the similarity between poses in our context.

Pose Metrics

As discussed earlier, the visual quality of the frames generated by our network, depends strongly on the similarity of the poses in the driving video to the poses present in the paired training data, extracted from the reference video. However, there is no need for each driving pose to closely match a single pose in the training data. Due to the use of a PatchGAN discriminator, a frame generated by our network may be constructed by combining together local patches seen during the training process. Thus, similarity between driving and reference poses should be measured in a local and translation invariant manner. In this section we propose a pose metrics that attempt to measure similarity in such a manner.

As explained in Section 3.1, poses are extracted as stacks of part confidence maps [Cao et al., 2017]. A 2D skeleton may be defined by estimating the position of each joint (the maximal value in each confidence map), and connecting the joints together into a predefined set of limbs. We consider 12 limbs (3 for each leg and 3 for each arm and shoulder), as depicted in Figure 4, for example. To account for translation invariance, our descriptor consists of only the orientation and length of each limb. The length of the limb is used to detect difference in scale that might affect the quality of the results, and the orientation is used to detect rotated limbs that haven’t been seen by the network during training. Fang et al. measure the distance between two poses using the normalized coordinates of the joints. This means that two limbs of the same size and orientation but shifted, will be considered different. In contrast, our descriptor is invariant to translation.

Specifically, our skeleton descriptor consists of vectors that measure the 2D displacement (in the image plane) between the beginning and the end of each limb. Namely, for every limb l∈Ll\in L we calculate Δxl=x1l−x2l\Delta x^{l}=x^{l}_{1}-x^{l}_{2} and Δyl=y1l−y2l\Delta y^{l}=y^{l}_{1}-y^{l}_{2}, where (x1l,y1l)(x^{l}_{1},y^{l}_{1}) and (x2l,y2l)(x^{l}_{2},y^{l}_{2}) are the image coordinates of the first and second joint of the limb, respectively. The descriptor pip_{i} consist of L=12L=12 such displacements: pi={pi1,…,piL}p_{i}=\{p^{1}_{i},\ldots,p^{L}_{i}\}, where pil=(Δxil,Δyil)p^{l}_{i}=(\Delta x_{i}^{l},\Delta y_{i}^{l}).

Given two pose skeleton descriptors, pi,pjp_{i},p_{j}, we define the distance between their poses as the average of distances between their corresponding limbs,

Figure 4 depicts the above distance between a reference pose (a) and different driving poses (b)–(d). Note that as the driving poses diverge from the reference pose, the distances defined by Eq. (7) increase. Limbs whose distance is above a threshold, dl(pi,pj)>γd_{l}(p_{i},p_{j})>\gamma, are indicated in gray. These limbs are unlikely to be reconstructed well from the frame corresponding to the pose in (a). Other limbs, which have roughly the same orientation and length as those in (a) are not indicated as problematic, even when significant translation exists (e.g., the bottom left limb of pose (c)).

Using the above metric we can estimate how well a given driving pose pip_{i}, might be reconstructed by a network trained using a specific reference sequence vv. The idea is simply to find, for each limb in pip_{i}, its nearest neighbor in the entire sequence vv, and compute the average distance to these nearest-neighbor limbs:

In other words, this metric simply measures whether each limb is present in the training data with a similar length and orientation. It should be noted that the frequency of the appearance of the limb in the training data set is not considered by this metric. Ideally, it should be also taken into account, as it might influence on the reconstruction quality.

The above metric is demonstrated in Figure 5, (a)–(c) show well reconstructed frames from poses whose distance to the reference sequence is small. The pose in (d) does not have good nearest neighbors for the arm limbs, since the reference video did not contain horizontal arm motions. Thus, the reconstruction quality is poor, as predicted by the larger value of our metric.

However, it should be emphasized that small pose distances alone cannot guarantee good reconstruction quality. Additional factors should be considered, such as differences in body shape and clothing, which are not explicitly represented by poses, as well as the accuracy of pose extraction, differences in viewpoint, and more.

Results and Evaluation

We implemented our performance cloning system in PyTorch, and performed a variety of experiments on a PC equipped with an Intel Core i7-6950X/3.0GHz CPU (16 GB RAM), and an NVIDIA GeForce GTX Titan Xp GPU (12 GB). Training our network typically takes about 4 hours for a reference video with a 256×\times256 resolution. When cloning a performance, the driving video frames must be first fed forward through the pose extractor P{\bf P} at 35ms per frame, and the resulting pose must be fed forward through the generator G{\bf G} at 45ms per frame. Thus, it takes a total of 80ms per cloned frame, which means that a 25fps video performance may be cloned in real time if two or more GPUs are used.

To carry out our experiments, we captured videos of 8 subjects, none of which are professional dancers. The subjects were asked first to imitate a set of simple motions (stretches, boxing motions, walking, waving arms) for two minutes, and then perform a free-style dance for one more minute in front of the camera. The two types of motions are demonstrated in the supplemental video. In all these videos we extracted the performer using an interactive video segmentation tool [Fan et al., 2015], since this is what our network aims to learn and to generate. We also gathered several driving videos from the Internet, and extracted the foreground performance from these videos as well.

Figure 6 shows four of our subjects serving as target individuals for cloning three driving poses extracted from the three frames in the leftmost column. The supplementary video shows these individuals cloning the same dance sequence with perfect synchronization.

Below, we begin with a qualitative comparison of our approach to two baselines (Section 5.1), and then proceed to a quantitative evaluation and comparison with the baselines (Section 5.2). It should be noted that, to our knowledge, there are no existing approaches that directly target video performance cloning. Thus, the baselines we compare to consist of existing state-of-the-art methods for unpaired and paired image translation.

We compare our video performance cloning system with state-of-the-art methods for unpaired and paired image translation techniques. The comparison is done against image translation methods since there is no other method that we’re aware of that is directly applicable to video performance cloning. Thus, we first compare our results to CycleGAN [Zhu et al., 2017], which is capable of learning a mapping between unpaired image sets from two domains. In our context, the two domains are frames depicting the performing actor from the driving video and frames depicting the target actor from the reference video.

It is well known that CycleGAN tends to transform colors and texture of objects, rather than modifying their geometry, and thus it tends to generate geometrically distorted results in our setting. More specifically, it attempts to replace the textures of the driving actor by those of the target one, but fails to properly modify the body geometry. This is demonstrated in Figure 7 (2nd column from the left).

Next, we compare our results to the pix2pix framework [Isola et al., 2016]. This framework learns to translate images using a paired training dataset, much like we do in our paired branch. To apply it in our context, we extend the pix2pix architecture to support translation of 3D part confidence maps, rather than images, and train it using a paired set of poses and frames from the reference video. Since the training uses a specific and limited set of motions, it may be seen that unseen poses cause artifacts. In addition, temporal coherency is not taken into consideration and some flickering exist in the results, which are included in the supplemental video.

Figure 7 shows a qualitative comparison between the results of the two baselines described above and the results of our method (rightmost column). It may be seen that our method generates frames of higher visual quality. The driving frame in this comparison is taken from a music video clip found on the web.

We next compare our method to Ma el al. [2017a], which is aimed at person image generation conditioned on a reference image and a target pose. We used a pertained version of their network that was trained on the “deep fashion” dataset. This dataset consists of models in different poses on white background, which closely resembles our video frames. Figure 8 compares a sample result of Ma et al. to our result for the same input pose using our network, trained on a reference video of the depicted target actor. It may be seen that the model of Ma et al. does not succeed in preserving the identity of the actor. The result of their model will likely improve after training with the same reference video as our network, but it is unlikely to be able to translate a pose sequence into a smooth video, since it does not address temporal coherence.

2. Quantitative Evaluation

We proceed with a quantitative evaluation of our performance cloning quality. In order to be able compare our results to ground truth frames, we evaluate our approach in a self-reenactment scenario. Specifically, given a captured dance sequence by the target actor, we use only two thirds of the video to train our network. The remaining third is then used as the driving video, as well as the ground truth result. Ideally, the generator should be able to reconstruct frames that are as close as possible to the input ones.

We have observed that although the last third of the video captures the same individual dancing as the first two thirds, which were used for training, the target actors seldom repeat the exact same poses. Thus, the generator cannot merely memorize all the frames seen during training and retrieve them at test time. This is demonstrated qualitatively in Figure 9: it may be seen that for the two shown driving poses, the nearest neighbors in the training part of the video are close, but not identical. In contrast, the frames generated by our system reproduce the driving pose more faithfully.

We performed the self-reenactment test for three different actors, and report the mean square error (MSE) between the RGB pixel values of the ground truth and the generated frames in Table 1. We also report the MSE for frames generated by the two baselines, CycleGAN [Zhu et al., 2017] and pix2pix [Isola et al., 2016]. Our approach consistently results in a lower error than these baselines. As a reference, we also report the error when the self-reenactment is done using the nearest-neighbor poses extracted from the training data. As pointed out earlier, the target actors seldom repeat their poses, and the errors are large.

We also examine the importance of the training set. In this experiment, we trained our translation network with 1200, 2000, and 3400 frames of the reference sequence of one of the actors (actor1). The visual quality of the results is demonstrated in Figure 10. The MSE error is 18.5 (1200 frames), 16.7 (2000 frames), and 15.4 (3400 frames). Not surprisingly, the best results are obtained with the full training set.

Finally, we examine the benefit of having the unpaired branch. We first train our network in a self reenactment setting, as previously described, using only the paired branch. Next, we train the full configuration, and measure the MSE error for both versions. We repeat the experiment for 3 different reference videos and the reported errors, on average, are 17.5 for paired branch training only vs. 16.2 when both branches are used. Visually, the result generated after training with only the paired branch exhibit more temporal artifacts. These artifacts are demonstrated in the supplementary video.

3. Online performance cloning

Once our network has been properly trained, we can drive the target actor by feeding the network with unseen driving videos. In particular, given the quick feed-forward time, with two or more GPUs working in parallel it should be possible to clone a driving performance as it is being captured, with just a small buffering delay. Naturally, the quality of the result depends greatly on the diversity of the paired data that the network was trained with.

Discussion and Future work

We have presented a technique for video-based performance cloning. We achieve performance cloning by training a generator to produce frame sequences of a target performer that imitate the motions of an unseen driving video. The key challenge in our work is to learn to synthesize video frames of the target actor that was captured performing motions which may be quite different from those performed in the driving video. Our method leverages the recent success of deep neural networks in developing generative models for visual media. However, such models are still in their infancy, and they are just beginning to reach the desired visual quality. This is manifested by various artifacts when generating fine details, such as in human faces, or in regions close to edges or high frequencies. Recent efforts in progressive GANs provide grounds for the belief that stronger neural networks, trained on more and higher resolution data will eventually be able to produce production quality images and videos.

Our method mainly focuses on two challenging aspects, dealing with unseen poses and generating smooth motions. Nevertheless, there is of course a direct correlation between the synthesis success and the similarity between the poses in the reference video and the driving videos. Moreover, the sizes and proportions of the performers’ skeletons should also be similar. With our current approach, it is impossible to transfer the performance of a grownup to a child, or vice versa.

Our current implementation bypasses several challenging issues in dealing with videos. We currently do not attempt to generate the background, and extract the performances from the reference video using an interactive tool. Furthermore, we heavy rely on the accuracy of the pose extraction unit, which is a non-trivial system on its own, and is also based on a deep network, that has its own limitations.

Encouraged by our results, we plan to continue and improve various aspects, including the ones mentioned above. In particular, we would like to consider developing an auxiliary universal adversarial discriminator that can tell how likely a synthesized frame is or how realistic a synthesized motion is. Such a universal discriminator might be trained on huge human motion repositories like the CMU motion capture library. We are also considering applying local discriminators that focus on critical parts, such as faces, hands, and feet. Synthesizing the fine facial details is important for the personal user experience, since we believe that the primary application of our method is a tool that enhances an amateur performance to match a perfect one, performed by a professional idol, as envisioned in beginning of this paper.

acknowledgements

We would like to thank the various subjects (Oren, Danielle, Adam, Kristen, Yixin, Fang) that were part of the data capturing and helped us to promote our research by cloning their performances.

References