Convolutional Sequence to Sequence Model for Human Dynamics

Chen Li, Zhen Zhang, Wee Sun Lee, Gim Hee Lee

Introduction

Understanding human motion is extremely important for various applications in computer vision and robotics, particularly for applications that require interaction with humans. For example, an unmanned vehicle must have the ability to predict human motion in order to avoid potential collision in a crowded street. Besides, applications such as sports analysis and medical diagnosis may also benefit from human motion modeling.

The biomechanical dynamics of human motion is extremely complicated. Although several analytic models have been proposed, they are limited to few simple actions such as standing and walking . For more complicated actions, data-driven methods are required to attain acceptable accuracy . In this paper, we focus on the human motion prediction task, using learning based methods from motion capture data.

Recently, along with the success of deep learning in various areas of computer vision and machine learning, deep recurrent neural network based models have been introduced in human motion prediction . In earlier works , it is often observed that there is a significant discontinuity between the first predicted frame of motion and the last frame of the true observed motion. Martinez et al. solved the problem by adding a residual unit in the recurrent network. However, their residual unit based model often converges to an undesired mean pose in the long-term predictions, i.e., the predictor gives static predictions similar to the mean of the ground truth of future sequences (see Figure 1). We believe that the mean pose problem is caused by the fact that it is difficult for recurrent models to learn to keep track of long-term information; the mean pose becomes a good prediction of the future pose when the model loses track of information from the distant past. For a chain-structured RNN model, it takes nn steps for two elements that are nn time steps apart to interact with each other; this may make it difficult for an RNN to learn a structure that is able to exploit long-term correlations .

The current state-of-the-art for human motion modeling is based on the sequence-to-sequence model, which is first proposed for machine translation . The sequence-to-sequence model consists of an encoder and a decoder, in which the encoder maps a given seed sequence to a hidden variable, and the decoder maps the hidden variable to the target sequence. A major difference between human motion prediction and other sequence-to-sequence tasks is that human motion is a highly constrained system by environment properties, human body properties and Newton’s Laws. As a result, the encoder needs to learn these constraints from a relative long seed sequence. However, RNN may not be able to learn these constraints accurately, and the accumulation of these errors in decoder may result in larger error in long-term prediction.

Furthermore, the human body is often not static stable during motion, and our central neural system must make multiple parts of our body coordinate with each other to stabilize the motion under gravity and other loads . Thus joints from different limbs have both temporal and spatial correlations. A typical example is human walking. During walking, most people tend to move left arm forward while moving right leg forward. However, RNN based methods have difficulties learning to capture this kind of spatial correlations well, and thus may generate some unrealistic predictions. Despite these limitations, the RNN based method is considered to be the current state-of-the-art as its performance is superior to other human motion prediction methods in terms of accuracy.

In this paper, we build a convolutional sequence-to-sequence model for the human motion prediction problem. Unlike previous chain-structured RNN models, the hierarchical structure of convolutional neural networks allows it to naturally model and learn both spatial dependencies as well as long-term temporal dependencies . We evaluate the proposed method on the Human3.6M and the CMU Motion Capture datasets. Experimental results show that our method can better avoid the long-term mean pose problem, and give more realistic predictions. The quantitative results also show that our algorithm outperforms state-of-the-art methods in terms of accuracy.

Related Works

The main task of this work is human motion prediction via convolutional models, while previous works mainly focus on RNN based models. We briefly review the literature as follows.

Data driven methods in human motion modeling face a series of difficulties including high-dimensionality, complicated non-linear dynamics and the uncertainty of human movement. Previously, Hidden Markov Model , linear dynamics system , Gaussian Process latent variable models etc. have been applied to model human motion. Due to limited computational resource, all these models have some trade-off between model capacity and inference complexity. Conditional Restricted Boltzmann Machine (CRBM) based method has also been applied to human motion modeling . However, CRBM requires a more complicated training process and it also requires sampling for approximate inference.

RNN based human motion prediction

Due to the success of recurrent models in sequence-to-sequence learning, a series of recurrent neural network based methods are proposed for the human motion prediction task . Most of these works have a recurrent network based encoder-decoder structure, where the encoder accepts a given motion frames or sequence and propagates an encoded hidden variable to the decoder, which then generates the future motion frame or series. The main differences in these works lie in their different encoder and decoder structures. For example, in the Encoder-Recurrent-Decoder (ERD) model , additional non-recurrent spatial encoder and decoder are are added to its recurrent part, which captures the temporal dependencies . In Structural-RNN , several RNNs are stacked together according to a hand-crafted spatial-temporal graph. Martinez et al. proposed a residual based model, which predicts the gradient of human motion rather than human motion directly, and used a standard sequence-to -sequence learning model with (GRU) cell as the encoder and decoder. In these RNN based models, fully-connected layers are used to learn a representation of human action, and the recurrent middle layers are applied to model the temporal dynamics. In contrast to previous RNN based models, we use a convolutional model to learn the spatial and the temporal dynamics at the same time and we show that this outperforms the state-of-the-art methods in human motion prediction.

Convolutional sequence-to-sequence model

The task of a sequence-to-sequence model is to generate a target sequence from a given seed sequence. Most sequence-to-sequence models consist of two parts, an encoder which encodes the seed sequence into a hidden variable and a decoder which generates the target sequence from the hidden variable. Although RNNs seem to be a natural choice for sequential data, convolutional models have also been adopted to sequence-to-sequence tasks such as machine translation. Kalchbrenner and Blunsom proposed the Recurrent Continuous Translation Model (RCTM) which uses a convolutional model as encoder to generate hidden variables and a RNN as decoder to generate target sequences, while later methods are fully convolutional sequence-to-sequence model. However, unlike machine translation, where only temporal correlations exist, there exist complicated spatial-temporal dynamics in human motion. Thus we design a convolutional sequence-to-sequence model that is suitable for complicated spatial-temporal dynamics.

Network Architecture

We adapt a multi-layer convolutional architecture, which has the advantage of expressing input sequences hierarchically. In particular, when we apply convolution to the input skeleton sequences, lower layers will capture dependencies between nearby frames and higher layers will capture dependencies between distant frames. Unlike the chain-structured RNN, the hierarchical structure of a multi-layer convolutional architecture is designed to capture long-term dependencies. Figure 2 shows an illustration of the architecture of our network, where the convolutional encoding module (CEM) plays the central role. We use the CEM module as long and short-term encoder. The long-term encoder is used to memorize a given motion sequence as the long-term hidden variable z⁡le\operatorname{\mathbf{z}}_{l}^{e}, and the short-term encoder is used to map a shorter sequence to the short-term hidden variable z⁡se\operatorname{\mathbf{z}}_{s}^{e}. Finally the hidden variables z⁡le\operatorname{\mathbf{z}}_{l}^{e} and z⁡se\operatorname{\mathbf{z}}_{s}^{e} are concatenated together and propagated to the decoder to give a prediction of next frame. The short-term encoder and decoder are applied recursively to produce the whole predicted sequence. We use convolutional layers with stride two in the CEM, thus two elements with distance nn are able to interact with each other in O⁡(log⁡n)\operatorname{\mathcal{O}}(\log n) operations. Furthermore, we use a rectangle convolution kernel to get a larger perception range in the spatial domain.

Similar to previous works , we also use an encoder-decoder model as a predictor to generate future motion sequences. However unlike previous works, we adapt a convolutional model for this sequence-to-sequence modeling task. Specifically, both the encoder and the decoder consist of similar convolutional structure, which computes a hidden variable based on a fixed number of inputs. There have been several convolutional sequence-to-sequence models that have been shown to give better performance than RNN based models in machine translation. However, these models mainly use convolution in the temporal domain to capture correlations, while in human motion there are also complicated spatial correlations between different body parts.

We first formalize the human motion prediction problem before giving more details of our convolutional model. Assume that we are given a series of seed human motion poses X⁡1:t=[x⁡1,x⁡2,…,x⁡t]\operatorname{\mathbf{X}}_{1:t}=[\operatorname{\mathbf{x}}_{1},\operatorname{\mathbf{x}}_{2},\dots,\operatorname{\mathbf{x}}_{t}], where each x⁡i∈\mathdsRL\operatorname{\mathbf{x}}_{i}\in\mathds{R}^{L} is a parameterization of human pose. The goal of human motion prediction is to generate a target prediction X⁡^(t+1):(t+T)\hat{\operatorname{\mathbf{X}}}_{(t+1):(t+T)} for the next TT frame poses.

We aim to capture the long-term information, such as categories of actions, human body properties (e.g. step length, step pace etc.), environmental constraints etc. from the seed human motion poses. To this end, a convolutional long-term encoder is used in our model. It maps the whole sequence X⁡1:t=[x⁡1,x⁡2,…,x⁡t]\operatorname{\mathbf{X}}_{1:t}=[\operatorname{\mathbf{x}}_{1},\operatorname{\mathbf{x}}_{2},\dots,\operatorname{\mathbf{x}}_{t}] to a hidden variable

where w⁡le\operatorname{\mathbf{w}}_{l}^{e} is the parameter of the long-term encoder hleh_{l}^{e}.

Our decoder has an encoder-decoder structure, which consists of a short-term encoder and a spatial decoder. The short-term encoder

where w⁡se\operatorname{\mathbf{w}}_{s}^{e} is the parameter. It maps a shorter sequence X⁡t−C+1:t\operatorname{\mathbf{X}}_{t-C+1:t}, which consists of CC neighboring frames of the current frame, to a hidden variable. Note that our short-term encoder is a sliding window of size CC, it only encodes the most recent CC frames. Finally, the long-term and short-term hidden variables z⁡le\operatorname{\mathbf{z}}_{l}^{e} and z⁡ee\operatorname{\mathbf{z}}_{e}^{e} are concatenated together as a input of the spatial decoder, which predicts the next pose x⁡^t+1\hat{\operatorname{\mathbf{x}}}_{t+1} as

where w⁡d\operatorname{\mathbf{w}}_{d} is the parameter of the spatial decoder hdh_{d}. To predict a sequence, the short-term encoder will slide one frame forward once a new frame is generated, thus the short-term encoder and decoder are applied recursively as

In our model, the long-term encoder and short-term encoder have the similar structure, i.e. the CEM, which includes 3 convolutional layers and 1 fully connected layers. The number of output channels for each convolutional layer is 64, 128 and 128, and the output number of the fully connected layer is 512. As discussed earlier, the CEM needs to capture long-term correlations in order to improve the prediction accuracy. Thus the stride of every convolutional layer is set to 22. With such convolutional layers, two elements of distance nn are able to interact with each other with a path length O⁡(log⁡(n))\operatorname{\mathcal{O}}(\log(n)), while O⁡(n)\operatorname{\mathcal{O}}(n) steps are required in a conventional RNN.

Furthermore, the perception range of the CEM in the spatial domain should be large enough to capture the spatial correlations of joints from different limbs. Hence, we use a rectangle 2×72\times 7 convolutional kernel (22 along temporal domain, and 77 along spatial domain) to enlarge the perception domain in the spatial domain. We use a simple two layer fully-connected neural network for the spatial decoder. The first layer maps a 1024 dimensional hidden variable to 512 dimensions and uses a leaky ReLU as the activation function. The second layer maps the hidden variable to one frame of human poses, and does not include an activation function. We also use a residual link in our network as suggested by previous works . This means that out decoder actually predicts the residual value rather than directly generates the next frame. Consequently, the output of our network consists of two parts:

hdh_{d} and wdw_{d} denote the decoder and its parameters.

In recurrent neural network based models (e.g. ), the encoded hidden variable often serves as the initial state of the decoding RNN. Thus during the long propagation path in RNN, the encoded information may vanish. However, our proposed model does not have this problem because the encoded hidden variable z⁡le\operatorname{\mathbf{z}}_{l}^{e} is always maintained. In recurrent neural networks, the model captures short-term dynamical information through variation of hidden states. In our model, the short-term dynamical information is captured by the short-term encoder from a short sequence. By using such a structure, our model is able to capture long-term invariant information and short-term dynamical information, and thus resulting in better performance in both long-term and short-term predictions.

2 Optimization

During training, we use the mean squared error of the predicted poses as the loss function:

Motivated by the generative adversarial network (GAN), we apply an adversarial regularizer for the proposed model, which mainly improves the qualitative performance. We train an additional discriminator to classify the generated and real sequences as follows

The discriminator DD is then used to encourage the generation of realistic sequences.

There are multiple choices of X⁡ˉt−C+k:t+k−1\bar{\operatorname{\mathbf{X}}}_{t-C+k:t+k-1} in (3.1), which may have different effects on the training results. In previous works, the corresponding part is often set to ground truth, or ground truth with noise . Besides setting X⁡ˉt−C+k:t+k−1\bar{\operatorname{\mathbf{X}}}_{t-C+k:t+k-1} as (3.1), it can also be set to

where η∈\eta\in is a manually specified parameter.

Note that the window size of the short-term encoder CC may also affect the results. The model may not capture enough short-term information when CC is too small. On the other hand, it may be a waste of computation when CC is too large since we already have the long-term encoder. Hence, the value of CC should be a trade-off between accuracy and computation. The effect of different window sizes are explored in our experiments.

Experiments

In this section, we apply the proposed convolutional model on several human motion prediction tasks. The proposed method is compared with several recent and state-of-the-art matching algorithms:

The Encoder-Recurrent-Decoder (ERD) method ;

An three layer LSTM with linear encoder and decoder (LSTM-3LR) ;

Stuctural Recurrent Neural Networks (SRNN) ;

Residual Recurrent Neural Networks (RRNN) ;

An three layer LSTM with an denoising auto encoder (LSTM-AE) .

Our model is implemented in tensorflow, and we used the ADAM optimizer to optimize over our model. The batch size is set to 64 and the learning rate is 0.0002. For more optimizing details, please refer to Section 3.2. Following the setting of previous works, the length of seed pose sequence is set to 50, and the length of target sequence is set to 25. We trained RRNN model based on the public available implementationhttps://github.com/una-dinosauria/human-motion-prediction. We quote the results from for ERD, LSTM-3LR, SRNN,, and for LSTM-AE.

ERD , LSTM-3LR and SRNN are action specific models, where they train a specific model for each action. On the other hand, the RRNN model considers the more challenging task of training a general model for multiple actions. In our experiments, we also train a single model for multiple actions.

1 Dataset and Preprosessing

In the experiments, we consider two datasets: the Human 3.6M dataset and the CMU Motion Capture dataset Available at http://mocap.cs.cmu.edu.

The Human 3.6M dataset is currently the largest available video pose dataset, which provides accurate 3D body joint locations recorded by a Vicon motion capture system. It is regarded as one of the most challenging datasets because of the large pose variations performed by different actors. There are 15 activity scenarios in total. Each action scenario includes 12 trials lasting between 3000 to 5000 frames. The 12 trials are categorized as 6 subjects, where each subject includes 2 trials. Each 3D pose consists of 32 joints plus a root orientation and displacement represented as an exponential map.

During the experiments, each pose would subtract to the mean pose over all trials and gets divided by the standard deviation. We eliminate the joint angle dimensions with constant standard deviation, which corresponds to joints with less than three degrees of freedom. Furthermore, the global rotation and translation are set to zero since our models are not trained with this information. Finally, the dimension of the input vector is set to 54. Similar to , we treat the two sequences in subject 5 as the test set and all others as the training set. For evaluation, we calculate the Euclidean error in terms of Euler angle. Specifically, we measure the Euclidean distance between our predictions and the ground truth in terms of Euler angle for each action, followed by calculating the mean value over all sequences which are randomly selected from the test set.

We also apply our model to the CMU Motion Capture dataset in order to test its generalization ability. There are five main categories in the dataset - “human interaction”, “interaction with environment”, “locomotion”, “physical activities & sports” and “situations & scenarios”. We choose some of the actions for our experiments based on some criteria. Firstly, we do not use data from the “human interaction” category since multiple subjects motion prediction is out of the scope of this paper. Secondly, action categories which include less than six trials are excluded on the consideration that we need enough data for each action to train our model. Lastly, some action categories in the dataset are actually combinations of other actions, e.g. actions in the subcategory “playground” consist of jump, climb and other actions which already exist in the dataset. We do not chose these action categories to avoid repetition. Finally, eight actions are selected for our experiments - running, walking and jumping from category “locomotion”, basketball and soccer from category “physical activities & sports”, wash windows from category “common behaviours and expressions”, traffic direction and basketball signals from category “communication gestures and signals”. We pre-process the data and evaluate the results in the same way as we did on the Human 3.6M dataset.

2 Evaluation on Human3.6M and CMU Datasets

We first report our results on all actions in the Human 3.6M dataset for both short-term prediction of 80 ms, 160 ms, 320 ms, 400 ms and long-term prediction of 1000 ms. Among the 15 actions in the dataset, the four actions “walking”, “eating”, “smoking” and “discussion” are commonly used in comparison of action specific human motion prediction methods. Thus we compare the accuracy of our method against four action specific methods ERD , LSTM-3LR , Structural RNN (SRNN) and LSTM-AE , as well as one general motion prediction method RRNN in Table 1. From the results, we can see that our method outperforms the others in most cases. We also provide the qualitative comparison results with the state-of-the-art RRNN method in Figure 3. Both RRNN and our model achieve good result on “walking” because of its periodic property, which makes the action easier to model. But for other aperiodic classes like “eating”, “smoking” and “discussion”, RRNN quickly converges to a mean pose – the predicted figure could not put its hands down in “eating” and raise its hand up in “discussion”.

Rather than maintaining a gesture in which one leg should be put on the other one in the action “smoking”, RRNN generates an implausible motion in real life that would cause the subject to go off balance. This further shows that it is very important to take the correlation between different body parts into consideration so that the predicted pose is more realistic. In comparison, our model predicts plausible motions for both “eating” and “smoking”. Furthermore, it is observed in the highly aperiodic action “discussion” that our model can still predict the correct motion trend, i.e. raising the hands while talking, even though this motion is not exactly the same as the ground truth. We compare our algorithm with the general human prediction model RRNN for the other 11 actions. The quantitative comparison results are provided in Table 2, which suggest that our algorithm outperforms RRNN in most cases. Additionally, our method outperforms RRNN on the average in both long and short-term predictions. The out-performance of our method becomes more significant for longer term predictions.

We only consider the more challenging task of training a general motion prediction model for all actions using the CMU Motion Caption dataset. Hence, we only show comparison results with the state-of-the-art RRNN method. For a fair comparison, both our model and RRNN are trained using the same settings on the Human3.6M dataset. The testing error of each action is given in Table 3 and the average testing error is given in Table 4. In the quantitative comparison, our method outperforms the RRNN method in several challenging actions such as jumping and running. The qualitative comparisons of running and jumping are also shown in Figure 5. In the qualitative comparisons, we can see that RRNN converges to mean pose for both running and jumping. On the other hand, our prediction for running is very close to the realistic one. For jumping, our model also predicts the correct motion trend, i.e. squatting followed by jumping, and the main error comes from the duration of squatting.

3 Ablation Study

The role of long-term encoder The long-term encoder in our model is used to capture long-term dependencies. We verify its effectiveness by removing it from our model. The results in Table 5 suggest that the average error gets larger without the long-term encoder, especially for long-term prediction of 1000 ms.

Rectangular kernel over spatial axis We use a rectangular kernel over the spatial axis (2×72\times 7 kernel) in our CEM in order to better capture the dependencies between different body parts. We also verify the effectiveness by comparing it with square kernel (4×44\times 4 kernel) and rectangular kernel over the temporal axis (7×27\times 2 kernel). The results in Table 6 indicate that the 2×72\times 7 kernel is the best choice.

Adversarial regularizer We use an adversarial regularizer in our model to help generate more plausible motion. In order to explore the role of the adversarial regularizer, we compare the performance of our model with and without the regularizer. The result in Figure 4 suggests that the adversarial regularizer helps to improve the performance of our model even though marginally. Moreover, since the adversarial regularizer is only used during training, it does not add complexity to our model during inference.

Different window size CC In our decoder, different window sizes CC result in different perception range. Intuitively, enlarging the window size may enlarge the perception range and results in better performance, but also requires more computation resources. We thus train three different models with window size C=5C=5, 1010 and 2020. In the right plot of Figure 4, we show average testing error over all 15 actions. The result suggests that there is not much improvement when the window size is larger than 10. Consequently, we set C=20C=20 for our model in view of the trade-off between accuracy and computational accuracy.

Conclusion

In this work, we proposed a convolutional sequence-to-sequence model for human motion prediction. We adopted two types of convolutional encoders in our model, namely the long-term encoder and short-term encoder, so that both distant and nearby temporal motion information can be used for future prediction. We demonstrated that our model performs better than existing state-of-the-art RNN based models, especially for long-term prediction tasks. Moreover, we show that our model can generate better predictions for complex actions due to the use of hierarchical convolutional structure for modeling complicated spatial-temporal correlations.

Acknowledgements

This work was partially supported by the Singapore MOE Tier 1 grant R-252-000-637-112 and the National University of Singapore AcRF grant R-252-000-639-114.

References