The Pose Knows: Video Forecasting by Generating Pose Futures

Jacob Walker, Kenneth Marino, Abhinav Gupta, Martial Hebert

Introduction

Consider the image in Figure 1. Given the context of the scene and perhaps a few past frames of the video, we can infer what likely action this human will perform. This man is outside in the snow with skis. What is he going to do in the near future? We can infer he will move his body forward towards the viewer. Visual forecasting is a fundamental part of computer vision with applications ranging from human computer interaction to anomaly detection. If computers can anticipate events before they occur, they can better interact in a real-time environment. Forecasting may also serve as a pretext task for representation learning .

Given this goal of forecasting, how do we proceed? How can we predict events in a data-driven way without relying on explicit semantic classes or human-labeled data? In order to forecast, we first must determine what is active in the scene. Second, we then need to understand how the structure of the active object will deform and move over time. Finally, we need to understand how the pixels will change given the action of the object. All of these steps have a level of uncertainty; however, the second step may have far more uncertainty than the other two. In Figure 1, we can already tell what is active in this scene, the skier, and given a description of the man’s motion, we can give a good guess as to how that motion will play out at the pixel level. He is wearing dark pants and a red coat, so we would expect the colors of his figure to still be fairly coherent throughout the motion. However, the way he skis forward is fairly uncertain. He is moving towards the viewer, but he might move to the left or right as he proceeds. Models that either try to directly forecast pixels or pixel motion are forced to perform all of these tasks simultaneously. What makes the problem harder for a complete end-to-end approach is that it has to simultaneously learn the underlying structure (what pixels move together), the underlying physics and dynamics (how the pixels move) and the underlying low-level rendering factors (such as illumination). Forecasting models may instead benefit if they explicitly separate the structure of objects from their low-level pixel appearance.

The most common agent in videos is a human. In terms of obtaining the underlying structure, there have been major advances in human pose estimation in images, making 2D human pose a viable “free” signal in video. In this paper, we exploit these advances to self-label video and aid forecasting. We propose a new approach to video forecasting by leveraging a more tractable space—human pose—as intermediate representation. Finally, we combine the strengths of VAE with those of GANs. The VAE estimates the probability distribution over future poses given a few initial frames. We can then forecast different plausible events in pose space. Given this structure, we then can use a Generative Adversarial Network to fill in the details and map to pixels, generating a full video. Our approach does not rely on any explicit class labels, human labeling, or any prior semantic information beyond the presence of humans. We provide experimental results that show our model is able to account for the uncertainty in forecasting and generate plausible videos.

Related Work

Activity Forecasting: Much work in activity forecasting has focused on predicting future semantic action classes or more generally semantic information . One way to move beyond semantic classes is to forecast an underlying aspect of human activity—human motion. However, the focus in recent work has been in specific data domains such as pedestrian trajectories in outdoor scenes or pose prediction on human-labeled mocap data . In our paper, we aim to rely on as few semantic assumptions as possible and move towards approaches that can utilize large amounts of unlabeled data in unconstrained settings. The only assumption we make on our data is that there is at least one detectable human in the scene. While the world of video consists of more than humans, we find that the great majority of video data in computer vision research focuses on human actions .

Generative Models: Our paper incorporates ideas from recent work in generative models of images. This body of work views images as samples from a distribution and seeks to build parametric models (usually CNNs) that can sample from these distributions to generate novel images. Variational Autoencoders (VAEs) are one such approach which have been employed in a variety of visual domains. These include modeling faces and handwritten digits . Furthermore, Generative Adversarial Networks (GANs) , have shown promise as well, generating almost photo-realistic images for particular datasets. There is also a third line of work including PixelCNNs and PixelRNNs which model the conditional distribution of pixels given spatial context. In our paper, we combine the advantages of VAEs with GANs. VAEs are inherently designed to estimate probability distributions of inputs, but utilizing them for estimating pixel distributions often leads to blurry results. On the other hand, GANs can produce sharp results, especially when given additional structure . Our VAE estimates a probability distribution over the more tractable space of pose while a GAN conditions on this structure to produce pixel videos.

Forecasting Video: In the last few years there have been a great number of papers focusing specifically on data-driven forecasting in videos. One line of work directly predicts pixels, often incorporating ideas from generative models. Many of these papers used LSTMs , VAEs , or even a PixelCNN approach . While these approaches work well in constrained domains such as moving MNIST characters, they lead to blurring when applied to more realistic datasets. A more promising direction for direct pixel prediction may be the use of adversarial loss . These methods seem to yield better results for unconstrained, realistic videos, but they still struggle with blurriness and uninterpretable outputs.

Given the difficulty of modeling direct pixels in video, many have resorted to pixel motion for forecasting. This seems reasonable, as motion trajectories are much more tractable than direct pixel appearances. These approaches can generate interpretable results for short time spans, but over longer time spans they are untenable. They depend on warping existing pixels in the scene. However, this general approach is a conceptual dead-end for video prediction—all it can do is move existing pixels. These methods cannot model occluded pixels coming into frame or model changes in pixel appearance.

Modeling low-level pixel space is difficult, and motion-based approaches are inherently limited. How then can we move forward with data-driven forecasting? Perhaps we can use some kind of intermediate representation that is more tractable than pixels. One paper explored this idea using HOG patches as an intermediate representation for forecasting. However, this work focused on specific domains involving cars or pedestrians and could only model rigid objects and rough appearances. In this paper, we use an intermediate representation which is now easy to acquire from video—human pose. Human pose is still visually meaningful, representing interpretable structure for the actions human perform in the visual world. It is also fairly low dimensional—many 2D human pose models only have 18 joints. Estimating a probability distribution over this space is going to be far more tractable than pixels. Yet human pose can still serve as a proxy for pixels. Given a video of a moving skeleton, it is then an easier task to fill in the details and output a final pixel video. We find that training a Video-GAN on completely unconstrained videos leads to results that are many times visually uninterpretable. However, when given prior structure of pose, performance improves dramatically.

Methodology

In this paper we break down the process of video forecasting into two steps. We first predict the high-level movement in pose space using the Pose-VAE. Then we use this structure to predict a final pixel level video with the Pose-GAN.

The first step in our pipeline is forecasting in pure pose space. At time tt, given a series of past poses P1..tP_{1..t} and the last frame of in input video XtX_{t}, we want to predict the future poses up to time step TT, Pt+1..TP_{t+1..T}. Pt∈R36P_{t}\in\mathcal{R}^{36} is a 2D pose as timestep tt represented by the (x,y)(x,y) locations of 1818 key-points. We actually predict a series of pose velocities Yt+1..TY_{t+1..T}. Given the pose velocities and an initial pose we can then construct the future pose sequence.

To accomplish this forecasting task, we build upon ideas related to sequential encoder-decoder networks . As in these papers we can use an LSTM to encode the past information sequence. We call this the Past Encoder which takes in the past information XtX_{t}, P1..tP_{1..t}, and Y1..tY_{1..t} and encodes it in a hidden representation HtH_{t}. We also have Past Decoder module to reconstruct the past information from the hidden state. Given this encoding HtH_{t} of the past, it would be tempting to use another LSTM to simply produce the future sequence of poses similar to . However, forecasting the future is not a deterministic problem; there may be multiple plausible outcomes of a video. Forecasting actually requires estimating a probability distribution over possible events. To solve this problem, we use a probabilistic Future Decoder. Our probabilistic decoder is nothing but a conditional variational autoencoder where the future velocity Yt+1Y_{t+1} is predicted given the past information HtH_{t}, the current pose Pt+1P_{t+1} (estimated from PtP_{t} and YtY_{t}), and the random latent vector zt+1z_{t+1}. The hidden states of the Future Decoder are updated using the standard LSTM update rules.

Variational Autoencoders: A Variational Autoencoder attempts to estimate the probability distribution P(Y∣z)P(Y|z) of its input data YY given latent variables zz. An encoder Q(z∣Y)Q(z|Y) learns to encode the inputs into a stochastic latent variable zz. The decoder P(Y∣z)P(Y|z) then reconstructs the inputs based on what is sampled from zz. During training, zz is regularized to match N(0,1)\mathcal{N}(0,1) through KL-Divergence. During testing we can then sample our distribution of YY by first sampling z∼N(0,1)z\sim\mathcal{N}(0,1) and then feeding our sample through a neural network P(Y∣z)P(Y|z) to create a sample from the distribution of Y. Another interpretation is that the decoder PP transforms the latent random variable z∼N(0,1)z\sim\mathcal{N}(0,1) into random variable Y∼P(Y∣z)Y\sim P(Y|z).

In our case, we want to estimate a distribution of future pose velocities given the past. Thus we aim to “encode” the future into latent variables z=[zt+1,zt+2,...zT]z=[z_{t+1},z_{t+2},...z_{T}]. Concretely, we wish to learn a way to estimate the distribution P(Yt+1..T∣z,Ht)P(Y_{t+1..T}|z,H_{t}) of future pose velocities Yt+1..TY_{t+1..T} given our encoded knowledge of the past HtH_{t}. Thus we need to train a “Future Encoder” that learns an encoding for latent variables z∼Q(z∣Yt+1..T,Ht)z\sim Q(z|Y_{t+1..T},H_{t}), where QQ is trained to match N(0,1)\mathcal{N}(0,1) as closely as possible. During testing, as in , we sample z∼N(0,1)z\sim\mathcal{N}(0,1) and feed sampled zz values into the future decoder network to output different possible forecasts.

Past Encoder-Decoder: Figure 3 shows the Past Encoder-Decoder. The Past Encoder takes as input a frame XtX_{t}, a series of previous poses P1..tP_{1..t}, and the previous pose velocities Y1..tY_{1..t}. We apply a convolutional neural network on XtX_{t}. The units from the pose information and the image features are concatenated and then fed into an LSTM. After encoding the entire sequence, we use the hidden state of the LSTM at step tt, HtH_{t} to condition the Future Decoder. To enforce that HtH_{t} encodes the pose velocity, the hidden state of of the encoding LSTM is fed into a decoder LSTM, the Past Decoder, which is trained through Euclidean loss to reconstruct Y1..tY_{1..t} in reverse order. This enforces that the network learns a “memory” of past inputs . The Past Decoder exists only as an aid for training, and at test time, only the Past Encoder is used.

Future Encoder-Decoder: Figure 4 shows the Future Encoder-Decoder. The Future Encoder-Decoder is composed of a VAE encoder (Future Encoder) and a VAE decoder (Future Decoder) both conditioned on past information HtH_{t}. The Future Encoder takes the future pose velocity Yt+1..TY_{t+1..T} and the past information HtH_{t} and encodes it as a mean and variance μ(Yt+1..T,Ht)\mu(Y_{t+1..T},H_{t}) and σ(Yt+1..T,Ht)\sigma(Y_{t+1..T},H_{t}). We then sample a latent variable z∼Q(z∣Yt+1..T,Ht)=N(μ,σ)z\sim Q(z|Y_{t+1..T},H_{t})=\mathcal{N}(\mu,\sigma). During testing, we sample zz from a standard normal, so during training we incorporate a KL-divergence loss such that QQ matches N(0,1)\mathcal{N}(0,1) as closely as possible. Given the latent variable zz and the past information HtH_{t}, the Future Decoder recovers an approximation of the future pose sequence Y^t+1..T(z,Ht)\hat{Y}_{t+1..T}(z,H_{t}). The training loss for this network is the usual VAE Loss. It is Euclidean distance from the pose trajectories combined with KL-divergence loss of QQ from N(0,1)\mathcal{N}(0,1).

At every future step tft_{f}, the Future Decoder takes in ztfz_{t_{f}}, as well as the current pose PtfP_{t_{f}} and outputs the pose motion YtfY_{t_{f}}. At training time, we use the ground truth poses, but at test time, we recover the future poses by simply adding the pose trajectory information Ptf+1=Ptf+YtfP_{t_{f}+1}=P_{t_{f}}+Y_{t_{f}}.

Implementation Details: We train our network with Adam Solver at a learning rate of 0.001 and β1\beta_{1} of 0.9. For the KL-divergence loss we set λ=0.00025\lambda=0.00025 for 60000 iterations and then set λ=0.0005\lambda=0.0005 for an additional 20000 iterations of training. Every timestep tt represents 0.2 second. We conditioned the past on 2 timesteps and predict for 5 timesteps. For the convolutional network over the image network, we used an architecture almost identical to AlexNet with the exception of a smaller (7x7) receptive field at the bottom layer and the addition of batch normalization layers. All layers in the entire network were trained from scratch. The LSTM units consist of two layers, both 1024 units. The Future Encoder is a simple single hidden layer network with ReLU activations and a hidden size of 512512.

2 Pose-GAN

Generative Adversarial Networks: Once we sample a pose prediction from our Pose-VAE, we can then render a video of a moving skeleton. Given an input image and a video of the skeleton, we train a Generative Adversarial Network to predict a pixel level video of future events. As described in , GANs consist of two models pitted against each other: a generator GG and a discriminator DD. The generator GG takes the input skeleton video and image and attempts to generate a realistic video. The discriminator DD, trained as a binary classifier, attempts to classify videos as either real or generated. During training, GG will try to generate videos which fool DD, while DD will attempt to distinguish the fake videos generated by GG from ones sampled from the future video frames. Following the work of we do not use any noise variables for the adversarial network. All the noise is contained in the Pose-VAE through zz.

Where VV are videos, MM is the batch size, II is an input image, and STS_{T} is a video of a pose skeleton, lrl_{r} is the real label (11), and lfl_{f} is the fake label (). Inside the batch MM, half of videos VV are generated, and the rest are real. The loss function ll here is the binary entropy loss.

Given our Pose-VAE, we can now generate plausible pose motions given a very short clip input. For each sample, we can render a video of a skeleton visualizing how a human will deform over the last frame of the input image. Recent work in adversarial networks has shown that GANs benefit from given structure . In particular, showed that GANs improve on generating humans when given initial keypoints. In this paper we build on this work by extending this idea to Conditional Video GANs. Given an image and a generated skeleton video, we train a GAN to generate a realistic video at the pixel level. Figure 5 shows the Pose-GAN network. The architecture of the discriminator DD is nearly identical to that of .

Implementation Details: The Pose-GAN consists of five volumetric convolutional layers with receptive fields of 4, stride of 2, and padding of 1. At each layer LeakyReLU units and Batch Normalization are used. The only difference is that the input is a 64x80 video. For the generator GG, we first encode the input using a series of five Volumentric Convolutional Layers with receptive fields of 4, stride of 2, and padding of 1. We use LeakyReLU and Batch Normalization at each layer. In order to handle the modified aspect ratio of the input (80x64), the fifth layer has a receptive field of 6 in the spatial dimensions. The top five layers are the same but in reverse, gradually increasing the spatial and temporal resolution to 64x80 pixels at 32 frames. Our training parameters are identical to , except that we set our regularization parameter α=1000\alpha=1000. Similar to , we utilize skip layers for the top part of the network. For the top five layers, ReLU activation and Batch Normalization is used. The final layer is sent through a TanH function in order to scale the outputs.

Experiments

We evaluate our model on UCF-101 in both pose space and video space. We utilized the training split described in which uses a large portion for training data. This split leaves out one video group for testing and uses the rest for training. In total we use around 1500 one-second clips for testing. To label the data we utilize the pose detector of Cao et al. and use the videos above an average confidence threshold. We perform temporal smoothing over the pose detections as a post-processing step.

First we evaluate how well our Pose-VAE is able to forecast actions in pose space. There has been some prior work on forecasting pose in mocap datasets such as the H3.6M dataset . However, to the best of our knowledge there has been no evaluation on 2D pose forecasting on unconstrained, realistic video datasets such as UCF101. We compare our Pose-VAE against state-of-the-art baselines. First, we study the effects of removing the VAE from our Future Decoder. In that case, the forecasting model becomes a Encoder-Recurrent-Decoder network similar to . We also implemented a deterministic Structured RNN model for forecasting with LSTMs explictly modeling arms, legs and torso. Finally, we take a feed-forward VAE and apply it to pose trajectory prediction. In our case, the feed-forward VAE is conditioned on the image and past pose information, and it only predicts pose trajectories.

Quantitative Evaluations: For evaluation of pose forecasting, we utilize Euclidean distance from the ground-truth pose velocities. However, specifically taking the Euclidean distance over all the samples from our model in a given clip may not be very informative. Instead, we follow the evaluation proposed by . For a set number of samples nn, we see what is the best possible prediction made by the model and consider the error of closest sample from the ground-truth. We then measure how this minimum error changes as the sample size nn increases and the model is given more chances. We make our deterministic baselines stochastic by treating the output as a mean of a multivariate normal distribution. For these baselines, we derive the bandwidth parameters from the variance of the testing data. Attempting to use the optimal MLE bandwidth via gradient search led to inferior performance. We describe the possible reasons for this phenomenon in the results section.

2 Video Evaluation

We also evaluate the final video predictions of our method. These evaluations are far more difficult as pixel space is much higher-dimensional than pose space. However, we nonetheless provide quantitative and qualitative evaluations to compare our work to the current state of the art in pixel video prediction. Specifically, we compare our method to Video-GAN . For this baseline, we only make two small modifications to the original architecture—Instead of a single frame, we condition Video-GAN on 16 prior frames. We also adjust the aspect ratio of the network to output a 64x80 video.

We also propose a new evaluation metric based on the test statistic Maximum Mean Discrepancy (MMD) . MMD was proposed as a test statistic for a two sample test—given samples drawn from two distributions P and Q, we test whether or not the two distributions are equal.

While the MMD metric is based on a two sample test, and thus is a metric for how similar the generated distribution is from the ground truth, the Inception score is a rather an ad hoc metric measuring the entropy of the conditional label distribution and marginal label distribution. We present scores for both metrics, but we believe MMD to be a more statistically justifiable metric

The exact MMD statistic for a class of functions F\mathcal{F} is:

Some nice properties of this test statistic are that the empirical estimate is consistent and converges in O(1n)O(\frac{1}{\sqrt{n}}) where nn is the sample size. This is independent of the dimension of data . MMD has been used in generative models before, but as part of the training procedure rather than as an evaluation criteria. uses MMD as a loss function to train a generative network to produce better images. extends this by first training an autoencoder and then training the generative network to minimize MMD in the latent space of the autoencoder, achieving less noisy images.

We choose Gaussian kernels with bandwidth ranging from 10−410^{-4} to 10910^{9} and choose the maximum of the values generated from these bandwidths as the reported value since from eq. (4), we want the maximum distance out of all possible functions.

Like Inception score, we use semantic features instead of raw pixels or flow for comparison. However, we use the fc7fc7 feature space rather than the labels. We concatenate the fc7fc7 features from the rgb stream and the flow stream of our action classifier. This choice choice of semantic fc7fc7 features is supported by the results in which show that training MMD on a lower-dimensional latent space rather than the original image space generates better looking images.

Results

In Figure 7 we show the qualitative results of our model. The results are best viewed as videos; we strongly encourage readers to look at our \hrefhttp://www.cs.cmu.edu/ jcwalker/POS/POS.htmlvideos. In order to generate these results, for each scene we took 1000 samples from Pose-VAE and clustered the samples above a threshold into five clusters. The pose movement shown is the largest discovered cluster. We then feed the last input frame and the future pose movement into Pose-GAN to generate the final video. On the far right we show the last predicted frame by Pose-GAN. We find that our Pose-GAN is able to forecast a plausible motion given the scene. The skateboarder moves forward, and the man in the second row, who is jumproping, moves his arms to the right side. The man doing a pullup in the third row moves his body down. The drummer plays the drums, the man in the living room moves his arm down, and the bowler recovers to standing position from his throw. We find that our Pose-GAN is able to extrapolate the pixels based on previous information. As the body deforms, the general shading and color of the person is preserved in the forecasts. We also find that Pose-GAN, to a limited extent, is able to inpaint occluded background as humans move from their starting position. In Figure 7 we show a side-by-side qualitative comparison of our video generation to conditional Video-GAN. While Video-GAN shows compelling results when specifically trained and tested on a specific scene category , we discover that this approach struggles to generate interpretable results when trained on inter-class, unconstrained videos from the UCF101. We specifically find that fails to capture even the general structure of the original input clip in many cases.

2 Quantitative Results

We show the results of our quantitative evaluation on pose prediction in Figure 6. We find our method is able to outperform the baselines on Euclidean distance even with a small number of samples. The dashed lines for ERD and SRNN use the only the direct output as a mean—identical to sampling with variance 0. As expected, the Pose-VAE has a higher error with only a few samples, but as samples grow the error quickly decreases due to the stochastic nature of future pose motion. The solid lines for ERD and SRNN treat the output as a mean of a multivariate normal with variance derived from the testing data. Using the variance seems to worsen performance for these two baselines. This suggests that these deterministic baselines output one particularly incorrect motion for each of the examples, and the distribution of pose motion is not well modeled by Gaussian noise. We also find our recurrent Pose-VAE outperforms Feedforward-VAE . Interestingly, FF-VAE underperforms the mean of the two deterministic baselines. This is likely due to the fact that FF-VAE is forced to predict all timesteps simultaneously, while recurrent models are able to predict more refined motions in a sequential manner.

In Table 2 we show our quantitative results of pixel-level video prediction against . As the Inception score increases, the KL-Divergence between the prior distribution of labels and the conditional class label distribution given generated videos increases. Here we are effectively measuring how often the two stream action classifier detects particular classes with high confidence in the generated videos. We compute variances using bootstrapping. We find, not surprisingly, that real videos show the highest Inception score. In addition, we find that videos generated by our model have a higher Inception score than . This suggests that our model is able to generate videos which are more likely to have particular meaningful features detected by the classifier. In addition to Inception scores, we show the results of our MMD metric in Table 2. While Inception is measuring diversity, MMD is instead testing something slightly different. Given the distribution of two sets, we perform a statistical test measuring the difference of the distributions. We again compute a variance with bootstrapping. We find that, compared to the distribution of real videos, the distribution videos generated by are much further than the videos generated by ours.

Conclusion and Future Work

In this paper, we make great steps in pixel-level video prediction by exploiting pose as an essentially free source of supervision and combining the advantages of VAEs, GANs and recurrent networks. Rather than try to model the entire scene at once, we predict the high level dynamics of the scene by predicting the pose movements of the humans in the scenes with a VAE and then predict each pixel with a GAN. We find that our method is able to generate a distribution of plausible futures and outperform contemporary baselines. There are many future directions from this work. One possibility is to combine VAEs with the power of structured RNNs to improve performance. Another direction is to apply our model to representation learning for action recognition and early action detection; our method is unsupervised and thus could scale to large amounts of unlabeled video data.

Acknowledgements: We thank the NVIDIA Corporation for the donation of GPUs for this research. In addition, this work was supported by NSF grant IIS1227495.

References