Dynamic Multiscale Graph Neural Networks for 3D Skeleton-Based Human Motion Prediction

Maosen Li, Siheng Chen, Yangheng Zhao, Ya Zhang, Yanfeng Wang, Qi Tian

Introduction

3D skeleton-based human motion prediction forecasts future poses given the past motions based on the human-body-skeleton. The motion prediction helps machines understand human behaviors, attracting considerable attention Fragkiadaki_2015_ICCV; Jain_2016_CVPR; Martinez_2017_CVPR; Butepage_2017_CVPR; Gui_2018_ECCV; Barsoum_2018_CVPR_Workshops. The related techniques can be widely applied to many computer vision and robotics scenarios, such as human-computer interaction 7102751; pmlr-v28-koppula13; Huang_eccv_2014; gui-2018-110272, autonomous driving SihengChen, and pedestrian tracking Alahi_2016_CVPR; Gupta_2014_CVPR; Bhattacharyya_2018_CVPR.

Many methods, including the conventional state-based methods Lehrmann_2014_CVPR; NIPS2005_2783; NIPS2006_3078; icml2009_129; NIPS2008_3567 and deep-network-based methods Fragkiadaki_2015_ICCV; Martinez_2017_CVPR; GhoshSAH17; abs-1810-09676; Gui_2018_ECCV; AAAI_Guo; Gopalakrishnan_2019_CVPR; quater; WangSpatio, have been proposed to achieve promising motion prediction. However, most methods did not explicitly exploit the relations or constraints between different body-components, which carry crucial information for motion prediction. A recent work Mao_2019_ICCV built graphs across body-joints for pairwise relation modeling; however, such a graph was still insufficient to reflect a functional group of body-joints. Another work WangSpatio builds pre-defined sturctures to aggregate body-joint features to represent fixed body-parts, while the model only considers the body physical constraints without exploiting the movement coordination and relations. For example, the action of ‘Walking’ tends to be understood based on the collaborative movements of abstract arms and legs, rather than the detailed locations of fingers and toes.

To model more comprehensive relations, we propose a new representation for a human body: a multiscale graph, whose nodes are body-components at various scales and edges are pairwise relations between components. To model a body at multiple scales, a multiscale graph consists of two types of sub-graphs: single-scale graphs, connecting body-components at the same scales, and cross-scale graphs, connecting body-components across two scales; see Figure 1. The single-scale graphs together provide a pyramid representation of a body skeleton. Each cross-scale graph is a bipartite graph, bridging one single-scale graph to another. For example, an “arm” node in a coarse-scale graph could connect to “hand” and “elbow“ nodes in a fine-scale graph. This multiscale graph is initialized by predefined physical connections and adaptively adjusted in training to be motion-sensitive. Overall, this multiscale representation provides a new potentiality to model body relations.

Based on the multiscale graph, we propose a novel model, called dynamic multiscale graph neural networks (DMGNN), which is action-category-agnostic and follows from an encoder-decoder framework to learn motion representations for prediction. The encoder contains a cascade of multiscale graph computational units (MGCU), where each is associated with a multiscale graph. One MGCU includes two key components: single-scale graph convolution block (SS-GCB), leveraging single-scale graphs to exact features at individual scales, and cross-scale fusion block (CS-FB), inferring cross-scale graphs to convert features from one scale to another and enable fusion across scales. The multiscale graph has adaptive and trainable inbuilt topology; it is also dynamic because the topology is changing from one MGCU to another; see the learned dynamic multiscale graphs in Figure 1. Notably, cross-scale graphs in CS-FBs are constructed adaptively to input motions, and reflect discriminative motion patterns for category-agnostic prediction.

As for the decoder, we adopt a graph-based gated recurrent unit (G-GRU) to sequentially produce predictions given the last estimated poses. The G-GRU utilizes trainable graphs to further enhance state propagation. We also use residual connections to stabilize the prediction. To learn richer motion dynamics, we introduce difference operators to extract multiple orders of motion differences as the proxies of positions, velocities, and accelerations. The architecture of DMGNN is illustrated in Figure 2.

To verify the superiority of our DMGNN, extensive experiments are conducted on two large-scale datasets: Human 3.6M 6682899 and CMU Mocap http://mocap.cs.cmu.edu/. The experimental results show that our model outperforms most state-of-the-art works for both short-term and long-term prediction in terms of both effectiveness and efficiency. The main contributions of this paper are as follow:

We propose dynamic multiscale graph neural networks (DMGNN) to extract deep features at multiple scales and achieve effective motion prediction;

We propose two key components: a multiscale graph computational unit, which leverages a multiscale graph to extract and fuse features across multiple scales, as well as a graph-based GRU to enhance state propagation for pose generation; and

We conduct extensive experiments to show that the proposed DMGNN outperforms most state-of-the-art methods for short and long-term motion prediction on two large datasets. We further visualize the learned graphs for interpretability and reasoning.

Related Work

Human motion prediction: To forecast motions, some traditional methods, e.g., hidden Markov models Lehrmann_2014_CVPR, Gaussian-process NIPS2005_2783 and random forests Lehrmann_2014_CVPR, were developed. Recently, deep networks are playing increasingly crucial roles: some recurrent-network-based models generated future poses step-by-step Fragkiadaki_2015_ICCV; Jain_2016_CVPR; Martinez_2017_CVPR; Walker_2017_ICCV; NIPS2016_6552; Gopalakrishnan_2019_CVPR; Liu_2019_CVPR; Gui_2018_ECCV; SymGNN; some feed-forward networks Li_2018_CVPR; Mao_2019_ICCV tried to reduce error accumulation for stable prediction; imitation-learning algorithm was also proposed Wang_2019_ICCV. However, these methods rarely considered enough relations from various scales, which carry comprehensive information for human behaviors understanding. In this work, we build dynamic multiscale graphs to capture rich multiscale relations and extract flexible semantics for motion prediction.

Graph deep learning: Graphs, expressing data associated with non-grid structures, preserve the dependencies among internal nodes AAAI1817135; Verma_2018_CVPR; valsesia2018learning. Many studies focused on graph representation learning and the relative applications Li2018learning; NIPS2016_6081; kipf_iclr2017; NIPS2017_6703; AAAI1817135; Si_2018_ECCV. Based on fixed graph structures, previous works explored propagating node features according to either the graph spectral domain NIPS2016_6081; kipf_iclr2017 or the graph vertex domain NIPS2017_6703. Several graph-based models have been employed for skeleton-based action recognition AAAI1817135; Li_cvpr_2019; Shi_2019_CVPR, motion prediction Mao_2019_ICCV and 3D pose estimation Zhao_2019_CVPR; Different from any previous works, our model considers multiscale graphs and corresponding operations.

Problem Formulation

To exploit rich body relations, we represent a body as a multiscale graph across multiscale body-components. Theorically, we could use arbitrary number of scales. Based on human nature, we specifically adopt 33 scales: the body-joint scale, the low-level-part scale, and the high-level-part scale. To initialize multiscale body graphs, we merge spatially nearby joints to coarser scales based on human prior; see Figure 3. With the multiscale graphs, we propose dynamic multiscale graph neural networks (DMGNN) to predict future poses in an end-to-end fashion.

Key Components

To construct our dynamic multiscale graph neural networks (DMGNN), we consider three basic components: a multiscale graph computational unit (MGCU), a graph-based GRU (G-GRU), and a difference operator.

The functionality of a MGCU is to extract and fuse features at multiple scales based on a multiscale graph, which is trained adaptively and individually. One MGCU includes two types of building blocks: single-scale graph convolution blocks, which leverage single-scale graphs to extract features at each scale, and cross-scale fusion blocks, which leverage cross-scale graphs to convert features from one scale to another and enable effective fusion across scales; see Figure 4. We now introduce each block in detail.

Cross-scale fusion block (CS-FB). To enable information diffusion across scales, we propose a cross-scale fusion block (CS-FB) which uses a cross-scale graph to convert features from one scale to another. A cross-scale graph is a bipartite graph that corresponds the nodes in one single-scale graph to the nodes in another single-scale graph. For example, the features of an “arm” node in the low-level-part scale s2s_{2} can potentially guide the feature learning of a “hand” node in the body-joint scale s1s_{1}. We aim to infer this cross-scale graph adaptively from data. Here we present CS-FB from s1s_{1} to s2s_{2} as an example.

2 Graph-based GRU

where rin(⋅)r_{\rm in}(\cdot), rhid(⋅)r_{\rm hid}(\cdot), uin(⋅)u_{\rm in}(\cdot), uhid(⋅)u_{\rm hid}(\cdot), cin(⋅)c_{\rm in}(\cdot) and chid(⋅)c_{\rm hid}(\cdot) are trainable linear mappings; WH\mathbf{W}_{\rm H} denotes the trainable weights. For each G-GRU cell, it applies a graph convolution on the hidden states for information propagation and produces the state for next frame.

3 Difference operator

Here we consider β=2\beta=2. The three elements reflects positions, velocities, and accelerations.

DMGNN Framework

Here we present the architecture of our DMGNN, which contains a multiscale graph-based encoder and a recurrent graph-based decoder for motion prediction.

2 Decoder

The decoder aims to predict future poses sequentially. The core of the decoder is the proposed graph-based GRU (G-GRU), which further propagates motion states for sequence regression. We first use the difference operator to extract three orders of differences as motion priors, and then feed them into G-GRU to update the hidden state. We next generate future pose displacement with an output function. Finally, we add the displacements to the input pose to predict the next frame. At frame tt, the decoder works as

where fpred(⋅)f_{\rm pred}(\cdot) represents an output function, implemented by MLPs. The initial state H(0)=H\mathbf{H}^{(0)}=\mathbf{H}, which is the final output of encoder.

3 Loss function

Experiments

Human 3.6m (H3.6M). H3.6M dataset 6682899 has 77 subjects performing 1515 different classes of actions. There are 3232 joints in each subject, and we transform the joint positions into the exponential maps and only use the joints with non-zero values (2020 joints remain). Along the time axis, we downsample all sequences by two. Following previous paradigms Martinez_2017_CVPR, the models are trained on 6 subjects and tested on the specific clips of the 5th subject.

CMU motion capture (CMU Mocap). CMU Mocap consists of 55 general classes of actions: ‘human interaction’, ‘interaction with environment’, ‘locomotion’, ‘physical activities & sports’, and ‘situations & scenarios’, where each subject has 3838 joints and we preserve 2626 joints with non-zero exponential maps. Be consistent with Li_2018_CVPR, we select 88 detailed actions: ‘basketball’, ‘basketball signal’, ‘directing traffic’, ‘jumping’, ‘running’, ‘soccer’, ‘walking’ and ‘washing window’. We evaluate our model with the same approach as we do for H3.6M.

Baseline methods. We compare the proposed DMGNN with many recent works, which learned motion patterns from pose vectors, e.g. Res-sup. Martinez_2017_CVPR, CSM Li_2018_CVPR, TP-RNN abs-1810-09676, AGED Gui_2018_ECCV, and Imit-L Wang_2019_ICCV, or separated bodies e.g. Skel-TNet AAAI_Guo, and Traj-GCN Mao_2019_ICCV. We reproduce, Res-sup., CSM and Traj-GCN based on their released codes. We also employ a naive baseline, ZeroV Martinez_2017_CVPR, which sets all predictions to be the last observed pose at t=0t=0.

2 Comparison to state-of-the-art methods

To validate the proposed DMGNN, we show the prediction performance for both short-term and long-term motion prediction on Human 3.6M (H3.6M) and CMU Mocap. We quantitatively evaluate various methods by the mean angle error (MAE) between the generated motions and ground-truths in angle space. We also illustrate the predicted samples for qualitative evaluation.

Short-term motion prediction. Short-term motion prediction aims to predict the future poses within 500 milliseconds. We compare DMGNN to state-of-the-art methods for predicting poses in 400 milliseconds on H3.6M dataset. We first test 44 representative actions: ‘Walking’, ‘Eating’, ‘Smoking’ and ‘Discussion’. Table 16 shows MAEs of DMGNN and some baselines. We also present the performance of several variants of DMGNN: we use fixed body-graphs in SS-GCBs (fixed As\mathbf{A}_{s}); the common GRU without a graph (no G-GRU); or only the joint-scale (S=1S=1) bodies. We see that, i) the complete DMGNN obtain the most precise prediction among all the variants; ii) compared to baselines, DMGNN has the lowest prediction MAEs on ‘Eating’ and ‘Smoking’, and obtains competitive results on ‘Walking’ and ‘Discussion’. Table 2 compares the proposed DMGNN with some recent baselines on the remaining 1111 actions in H3.6M. We see that DMGNN achieves the best performance in most actions (also for average MAEs).

Long-term motion prediction. Long-term motion prediction aims to predict the poses over 500 milliseconds, which is challenging due to the action variation and non-linearity movements. Table 3 presents the MAEs of various models for predicting 44 actions and average MAEs across the 44 actions in the future 560 ms and 1000 ms on H3.6M dataset. We see that DMGNN outperforms the competitors on actions ‘Eating’, and ‘Discussion’ at 560 ms, and obtains competitive performances on other cases.

We also train our DMGNN for short-term and long-term prediction on 88 classes of actions in CMU Mocap dataset. Table 4 shows the MAEs across the future 1000 ms. We see that DMGNN significantly outperforms the state-of-the-art methods on actions ‘Basketball’, ‘Basketball Signal’, ‘Running’ and ‘Walking’ and obtains competitive performance on the other actions.

Predicted sample visualization. We compare the synthesized samples of DMGNN to those of Res-sup., CSM and Traj-GCN on H3.6M. Figure 6 illustrates the future poses of ‘Taking Photo’ in 1000 ms with the frame interval of 80 ms. Comparing to baselines, we see that DMGNN completes the action accurately and reasonably, providing significantly better predictions. Res-sup. has large discontinuity between the last observed pose the first predicted one (red box); CSM and Traj-GCN have large errors after the 280th ms (blue box); three baselines give large posture errors in long-term (yellow box). We show more prediction images and videos in Appendix.

Effectiveness and efficiency test. We compare the running time costs of DMGNN to several latest models. Table 5 presents the running time of different methods for short and long-term motion prediction on H3.6M dataset. We see that DMGNN achieves the shortest running time while generating future poses over both 400 or 1000 ms, compared with the other competitors Martinez_2017_CVPR; Li_2018_CVPR; Mao_2019_ICCV. DMGNN takes only 29.1829.18 ms to generate motions in 400 ms, indicating that DMGNN with multiscale graphs has efficient operations.

3 Ablation study

We now investigate some crucial elements of DMGNN.

To verify the proposed multiscale representation, we employ various scales in DMGNN for 3D skeleton-based motion prediction. Besides the three scales in our model, we introduce additional two scales: s4s_{4}, which represents a body as Ms4=3M_{s_{4}}=3 parts: left limbs, right limbs and torso, and s5s_{5}, which contains Ms5=2M_{s_{5}}=2 parts: upper body and lower body; see illustrations of s4s_{4} and s5s_{5} in Appendix. Table 6 presents the MAEs with various scales. We see that, when we combine s1s_{1}, s2s_{2} and s3s_{3}, lowest prediction error is achieved. Notably, using two scales (s1,s2s_{1},s_{2} or s1,s3s_{1},s_{3}) is significant better than using only s1s_{1}; but involving too abstract scales (s4s_{4} or s5s_{5}) tends to hurt prediction.

To validate the effects of multiple MGCUs in the encoder, we tune the numbers of MGCUs from 11 to 66 and show the prediction errors and running time costs for short and long-term prediction on H3.6M, which are presented in Table 7. We see that, when we adopt 11 to 44 MGCUs, the prediction MAEs fall and time costs rise continuously; when we use 55 or 66 MGCUs, the prediction errors are stably low, but the time costs rise higher. Therefore, we select to use 44 MGCUs, resulting in precise prediction and high running efficiency.

Here, we evaluate 1) the effectiveness of using relative features during cross-scale graph inference in CS-FBs; 2) different numbers of CS-FBs in a sequence of 44 MGCUs. For 00 CS-FB, the model only fuses all scales at the end of the encoder. Table 8 presents the average MAEs with different CS-FBs and relative-feature mechanisms across 400 ms on H3.6M. We see that 1) using relative features leads to lower MAEs, validating the effectiveness of such augmented features; 2) 22 CS-FBs leads to the best prediction performance. The intuition is that 0 or 1 CS-FB fuse insufficiently and 3 CS-FBs tend to fuse redundant information to confuse the model.

Effect of λ\lambda in final fusion. The hyper-parameter λ\lambda in the final fusion (3) balances the influence between joint-scale and more abstract scales. Figure 7 illustrates the average MAE with different body scales and CS-FBs for short-term prediction on H3.6M. We see that the performance reach its best when we use 33 scales, 22 hierarchical CS-FBs and λ=0.6\lambda=0.6, even though it is robust to the change of λ\lambda.

We study the effects of various orders of motion differences fed into the encoder and decoder of our model. We evaluate DMGNN with combinations of 0,1,20,1,2-orders of pose differences. Table 9 presents the MAEs of DMGNN with various input differences for short-term motion prediction. We see that the proposed DMGNN obtains the lowest MAEs when it adopts the 0,1,20,1,2-orders of motion differences. This indicates that high-order differences improve the prediction performance significantly.

4 Analysis of category-agnostic property

Here we validate that DMGNN can learn discriminative motion features for category-agnostic prediction.

We first visualize the learned cross-scale graphs for different actions to test the discriminative power. Figure 8 shows the graphs in two CS-FBs on ‘Walking’ and ‘Directions’ in H3.6M. For each action, we show some strong relations from detailed scales to the right arms in coarse scales. We see that i) for each action, the CS-FBs capture diverse ranges of a human body: the graph in the first CS-FB focuses on nearby body-components; the second CS-FB captures more global and action-related effects; i.e. hands and feet affects arms during walking; and ii) the cross-scale graphs are different for various actions, especially in the second CS-FB, capturing distinct patterns.

We next conduct action classification on the intermediate representations to test the discriminative power. We isolatedly train a two-layer MLP to classify each dynamic cross-scale graph. We also classify the outputs from the encoders of DMGNN, Res-sup. (class-aware) and TP-RNN (class-agnostic). Table 10 presents the average classification accuracies on 1515 categories of actions. We see that the cross-scale graph in the second CS-FB is more informative than the one in the first CS-FB for action recognition. Comparing to baselines, DMGNN obtains the highest the classification accuracies on encoder representation, indicating that DMGNN captures discriminative information for class-agnostic prediction.

Conclusion

We build dynamic mutiscale graphs to represent a human body and propose dynamic multiscale graph neural networks (DMGNN) with an encoder-decoder framework for 3D skeleton-based human motion prediction. In the encoder, We develop multiscale graph computational units (MGCU) to extract features; in the decoder, we develop a graph-based GRU (G-GRU) for pose generation. The results show that the proposed model outperforms most state-of-the-art methods for both short and long-trem prediction in terms of both effectiveness and efficiency.

Acknowledgement: This work is supported by the National Key Research and Development Program of China (No. 2019YFB1804304), SHEITC (No. 2018-RGZN-02046), NSFC (No. 61521062), 111 plan (No. B07022), and STCSM (No. 18DZ2270700).

References

Detailed Architecture

Here we show the detailed structure of the proposed DMGNN. We first show the structure of the encoder, including the single-scale graph convolution block (SS-GCB) and cross-scale fusion block (CS-FB). We then show the structure of the decoder, including the graph-based gated recurrent unit (G-GRU).

Single-scale graph convolution block (SS-GCB). SS-GCB consists of a graph convolution and a temporal convolution. Table 11 presents the structures of four cascaded SS-GCB at scale ss in the encoder of DMGNN.

We see that we use four SS-GCBs to extract spatio-temporal motion features. In each SS-GCB, we employ ReLU, batch normalization, and dropout operations. We use stride 2 to downsample along the temporal dimension.

Cross-scale fusion block (CS-FB) We use CS-FB to fuse multiscale features. Table 12 presents the structure of the first CS-FB to fuse the feature from s1s_{1} to s2s_{2}.

We first use a temporal convolution to shrink the temporal dimension and obtain a compact feature vector for each body-component; we then use four MLPs to learn the feature embeddings for two body-scales, respectively; we finally calculate the inner product of these two embeddings and employ a softmax to calculate the corresponding edge weight in a cross-scale graph.

Total architecture In summary, we show the total architecture of the encoder, which combine SS-GCBs at multiple scales and CS-FB across scales. Table 13 presents the structure of the encoder.

We see that we use four MGCUs, where the first two MGCUs use SS-GCBs and CS-FBs to learn the features from multiscale bodies and the last two MGCUs only use SS-GCB to extract features.

2 Decoder

Graph-based Gated Recurrent Unit (G-GRU) G-GRU is one of the key components in the proposed decoder for synthesizing precise and reasonable future poses. Table 14 presents the structure of the G-GRU at time stamp tt.

We see that we take the historical motion state and the online 3D skeleton-based information as inputs and introduce the graph convolution to propagate the motion information to produce the motion state at the next frame. The hidden dimension of the G-GRU is 256.

Total architecture Here, we show the total architecture of the decoder, which combines the proposed G-GRU and an MLP-formed output function. Table 15 presents the structure the decoder at time stamp tt.

We see that, given the hidden motion state and current input information, we use a G-GRU and an MLP-formed output function fpredf_{\rm pred} to model the displacement of motions between two consecutive frames, and we emply residual connections to obtain the estimated poses. The hidden dimensions are 256.

Quantitative Comparison with more Baselines

In our paper submission, we only compare DMGNN to several state-of-the-art works, while many other methods has been developed. Here we compare DMGNN to as many previous methods as possible. Table 16 presents the MAE of many methods for short-term motion prediction on 4 representative actions of Human 3.6M

We see that, the proposed DMGNN outperforms the state-of-the-art methods on most actions. Notably, we have cited all of baselines presented in Table 16 in our paper submission.

Coarser Body-scales in Ablation Studies

In the first experiment of ablation studies (‘effects of multiple scales’), we initialize two coarser body-scales (s4s_{4} and s5s_{5}) besides the effective three scales (s1s_{1}, s2s_{2} and s3s_{3}) that used in our DMGNN. Here we present s4s_{4} and s5s_{5} in details.

To initialize s4s_{4}, we average the input features of three body-components: left-body, head-and-torso, and right-body as the nodes of corresponding body-graph. We build two initial edges to respectively connect head-and-torso with left-body and right-body. To initialize s5s_{5}, we average the input features of two body-components: upper-body and lower-body as the graph nodes. We build an edge between these two body-components. Figure 9 illustrates the two coarser body-scales as well as the body-joint scale on Human 3.6M 6682899. We name s4s_{4} as ‘Left-right-body scale’ and name s5s_{5} as ‘Up-low-body scale’.

Effects of Numbers and Positions of CS-FBs

In our DMGNN, we employ CS-FBs with aggregating relative features at different MGCUs to fuse various levels of motion features across different scales; see Equation (2a) in the submission. Here we further investigate the effects of numbers and positions of CS-FBs at cascaded MGCUs. In the four MGCUs, we use one to four CS-FBs at different MGCUs, and we obtain the average prediction MAEs of different model variants.

Table 17 presents the average MAEs of DMGNN with different numbers of CS-FBs at different MGCUs on H3.6M for short-term motion prediction. We also compare the performance of CS-FBs with or without aggregating relative information from all the body-components (‘with relative’ or ‘without relative’). We denote the numbers of CS-FBs at the column ‘Number’ and denote the CS-FB positions as MGCU indices at column ‘Position’.

We see that 1) when we aggregate global relative information to in the CS-FB, we obtain lower MAEs than the module without relative information aggregation; 2) when we use two CS-FBs with relative information aggregation at the 1st and 2nd MGCUs, DMGNN produces the most precise predictions across different model variants; 3) fusing multiscale features at first few MGCUs outperforms fusing at last ones. The reason behind could be, if we use only one CS-FB, we cannot fuse rich features for comprehensive pattern learning; if we use too many CS-FBs, the capacity of the network become much larger, leading to overfitting.

More Generated Motion Samples

To further demonstrate the effectiveness of the ASGNN, we illustrate more predicted samples on both Human 3.6M 6682899 and CMU Mocap http://mocap.cs.cmu.edu/ dataset.

We first illustrate two generated motions of the actions of ‘Posing’ and ‘Waiting’ on Human 3.6 dataset (H3.6). We compare the DMGNN with three models: Res-sup. Martinez_2017_CVPR, CSM Li_2018_CVPR and Traj-GCN Mao_2019_ICCV.

Figure 10 illustrates the predicted poses of ‘Posing’ in Human 3.6M in 1000 ms.

We see that the proposed DMGNN could well model the posture, such as stretched bodies and arms; however, Res-sup predicts the motion with large discontinuity between the last observed pose the first predicted one (red box); CSM and Traj-GCN tends to have large errors after the 400th ms (blue box); all the baselines produce unreasonable poses at the 1000th ms (yellow box), which are far from the ground truth.

We also predict the action of ‘Waiting’ in Human 3.6M in a long term with different methods. The results are illustrated in Figure 11.

We see that, for baselines, the motion predicted by res-sup has large discontinuity between the last observed pose the first predicted one (red box) and loses the movements, which is far from the ground truths. CSM and Traj-GCN suffer from large errors after the 320th ms; all the baselines predict unreasonable poses at the 1000th ms (yellow box); but the predictions from DMGNN could complete the action reasonably.

2 CMU Mocap Dataset

We then test DMGNN on the two actions of ‘Basketball’ and ‘Washing window’ in CMU Mocap dataset. The baselines are the CSM Li_2018_CVPR and Traj-GCN Mao_2019_ICCV.

For the action of ‘Basketball’, the main challenge of motion prediction is the running legs and swaying arms. We illustrate the generated samples of three models in Figure 13.

We see that the errors of the predictions from CSM and Traj-GCN rise after the 320th ms (blue box); two baselines give unreasonable postures at the 1000th ms in long-term (yellow box); that is, CSM has wrong tilt orientation of the body and the left leg (purple) of the pose predicted by Traj-GCN has inaccurate position; DMGNN could predict motions with smaller errors in both short-term and long-term.

For the action of ‘Washing window’, we also predict the future poses in 1000 ms and illustrate them in Figure 13.

We see that the prediction of CSM has large discontinuity between the last observed pose the first predicted one (red box); Traj-GCN tends to have large errors after the 400th ms, since the pose does not raise the left arm (blue box); two baselines give poses at the 1000th ms with large errors (yellow box); but DMGNN could predict motions with smaller errors in both short-term and long-term.