A Spatio-temporal Transformer for 3D Human Motion Prediction

Emre Aksan, Manuel Kaufmann, Peng Cao, Otmar Hilliges

Introduction

3D human motion modelling is typically formulated as the prediction of future poses given a past horizon. Humans are able to effortlessly forecast the complex dynamics of motion in a plausible fashion due to our strong structural and temporal priors. From a learning perspective this problem can be seen as a generative modelling task: A network learns to synthesize a sequence of human poses, where the model is conditioned on the seed sequence. The task requires learning of pose priors for natural articulation and of underlying dynamics to yield plausible motion predictions. Since these factors are highly latent and entangled, introducing inductive biases and tailoring architectures for the task is essential for modelling of 3D human motion data.

Given the temporal nature of human motion, it is not surprising that recurrent neural networks (RNNs) are the most popular choice . RNNs model short and long-term dependencies by propagating information through their hidden state. Convolutional neural networks (CNN) in a sequence-to-sequence framework have also been proposed . Such approaches focus on modelling the temporal aspect of the problem following an auto-regressive approach, but neglect structural priors. Instead, vectorized poses are passed as inputs at every step and the spatial dependencies are assumed to be learned implicitly. However, considering the skeletal structure in the architectural level is shown to be an effective inductive bias in .

Since the auto-regressive approach factorizes the predictions into step-wise conditionals based on previous predictions, these models tend to accumulate error over time and eventually the predictions collapse to a non-plausible pose. This issue can be associated with the exposure bias problem due to discrepancies between data and model distributions. Previous work has applied various strategies to work around this problem, such as using model predictions during training , applying noise to the inputs , or using adversarial losses .

Recent works model the temporal aspects of 3D human motion by encoding every joint’s trajectory with the discrete cosine transformation (DCT). Both the observations and the predicted future frames are represented as a set of DCT coefficients which are then used to model inter-joint dependencies. Such an implicit modelling of the temporal information inherently mitigates the failure cases of the auto-regressive models. DCT appears to be an effective non-learning based representation.

In this work, we present a novel architecture for 3D human motion modelling, which attempts to learn a spatio-temporal representation explicitly without relying on the propagation of a hidden state as in RNNs or fixed temporal encodings such as DCT coefficients. Our approach is motivated by the recent success of the Transformer model in tasks such as NLP , music , or images . While the vanilla Transformer is designed for 1D sequences with a self-attention mechanism , we note that the task of 3D motion prediction is inherently spatio-temporal and we propose a novel representation that decouples the temporal and spatial dimensions.

The proposed spatio-temporal attention mechanism is trained to identify useful information from a known sequence to construct the next output pose. For every joint we define temporal attention over the same joint in the past and spatial attention over the other joints at the same time step (see Fig. 1). The spatial attention block draws information from the joint features at the current time step whereas the temporal block focuses on distilling information from the previous time steps of individual joints. A prediction is then made by summarizing current joint information and previous time steps as a weighted combination.

Our model learns to construct temporally coherent poses from individual joints by considering the temporal and spatial representations learned from the data. The dual self-attention over the sequence allows the model to access past information directly and hence capture the dependencies explicitly , mitigating error accumulation over time. It also enables interpretability since attention weights indicate informative sequence parts that led to the prediction. Our experiments show that a naive application of 1D self-attention still suffers from the collapsing pose problem, whereas our proposed model, the ST-Transformer, is able to outperform the state-of-the-art models in short-term horizons and also produce convincing long-term predictions (up to 20 seconds for periodic motions).

Related Work

Decomposition of the attention mechanism across different dimensions has been shown to be effective in other domains. In , attention is applied separately to the height and width dimensions of images for semantic segmentation. proposes a similar model for skeleton-based action recognition where attention is applied sequentially first to the joints and then to the temporal dimension. introduces axial attention applying self-attention on different dimensions in parallel to reduce computational complexity. Our work is conceptually similar to those but crucially differs in the task domain. In the remainder of this section, we provide a summary of the related work on 3D motion modelling.

Recurrent Models RNNs are the dominant architecture for 3D motion modelling tasks . Fragkiadaki et al. propose the Encoder-Recurrent-Decoder (ERD) model where an LSTM cell operates in latent space. Jain et al. build a skeleton-like st-graph with RNNs as nodes. Aksan et al. replace the dense output layer of a RNN architecture with a structured prediction layer that follows the kinematic chain. The authors furthermore introduce a large-scale human motion dataset, AMASS , to the task of motion prediction. The error accumulation problem is typically combated by exposing the model to dropout or Gaussian noise during training. Ghosh et al. more explicitly train a separate de-noising autoencoder that refines the noisy RNN predictions.

Martinez et al. introduce a sequence-to-sequence (seq2seq) architecture with a skip connection from the in- to the output on the decoder to address the transition problems between the seed and predictions. They also propose training the model with the predictions to alleviate the exposure bias problem. Similarly, Pavllo et al. suggest to use a teacher-forcing ratio to gradually expose the model to its own predictions. In , the seq2seq framework is modified to explicitly model different time scales using a hierarchy of RNNs. Gui et al. propose a geodesic loss and adversarial training and Wang et al. replace the likelihood objective with a policy gradient method via imitation learning. The acRNN of Zhou et al. uses an augmented conditioning scheme which allows for long-term motion synthesis. However their model is trained on specific motion types, whereas in this work we present a single model encompassing multiple motion types.

While the use of domain-specific priors and objectives can improve short term accuracy, the inherent problems of the underlying RNN architectures still exist. Maintaining long-term dependencies is an issue due to the need of summarizing the entire history in a hidden state of fixed size. Inspired by similar observations in the field of NLP , we introduce a spatio-temporal self-attention mechanism to mitigate this problem. In doing so we let the network explicitly reason about past frames without the need to compress the past into a single hidden vector.

Non-recurrent Models Bütepage et al. use dense layers on sliding windows of the motion sequences. In convolutional models are introduced for motion synthesis conditioned on trajectories. More recently, Hernandez et al. propose to treat motion prediction as an image inpainting task and use a convolutional model with adversarial losses. Joints are represented as 3D positions, often requiring auxiliary losses such as bone length and joint limits to ensure anatomical consistency. Li et al. use CNNs instead of RNNs in the sequence-to-sequence framework to improve long-term dependencies. Similarly, Kaufmann et al. propose a convolutional autoencoder for the 3D motion infilling task to fill in large gaps between given sequences.

Implicit Temporal Models Mao et al. represent sequences of joints via discrete cosine transform (DCT) coefficients and train a graph convolutional network (GCN) to learn inter-joint dependencies. Since the GCN operates on temporal windows of poses and produces the entire output in one go, the predictions are limited to a pre-determined length. In follow-up work , DCT coefficients are instead extracted from shorter sub-sequences in an overlapping sliding window fashion which are then aggregated via a 1D attention block. Similarly, Cai et al. leverage a Transformer architecture on the DCT coefficients extracted from the seed sequence and make joint predictions progressively by following the kinematic tree.

Our model is related to these approaches, but differs in three aspects. First, the DCT requires windowed inputs and produces the entire output in one go. This limits full generative modelling of arbitrarily long sequences with sufficient diversity. Instead, we aim to learn spatio-temporal representations directly from the data. Second, we follow a fully auto-regressive approach and model the temporal dependencies explicitly by leveraging the recursive nature of human motion. Third, temporal and spatial modelling is interleaved in our design, whereas in previous work the temporal information is modeled first via the DCT and aggregated with an attention mechanism , and then the spatial structure is captured by a GCN or a Transformer. In contrast, our model stacks several computation blocks each of which aggregates temporal and spatial information and passes it to the subsequent layer in a message passing fashion.

In summary, existing 3D motion modelling works have introduced regularization, structural priors, frequency transformations, or auxiliary and adversarial loss terms to address the inherent problems of the underlying architectures. We show that the self-attention concept itself is very effective in learning motion dynamics and allows for the design of a versatile mechanism that is effective and easy to train.

Method

We now explain the architecture of the proposed spatio-temporal transformer (ST-Transfomer) in detail. For an overview please refer to Fig. 2. Our method uses the building blocks of the Transformer , but with two main differences: (1) a decoupled spatio-temporal attention mechanism and (2) a fully auto-regressive model.

2 Spatio-temporal Transformer

We use the scaled dot-product attention proposed by , requiring query Q\boldsymbol{Q}, key K\boldsymbol{K}, and value V\boldsymbol{V} representations. Intuitively, the value corresponds to the set of past representations that are indexed by the keys. For the joint of interest, we compare its query representation with all keys w.r.t. the dot-product similarity. If the query and the key are similar (i.e., high attention weight), then the corresponding value is assumed relevant. The attention operation yields a weighted sum of values V\boldsymbol{V}:

Spatial Attention In the vanilla Transformer, the attention block operates on the entire input vector xt\boldsymbol{x}_{t} and the relation between the elements are implicitly captured. We introduce an additional spatial attention block to learn dynamics and inter-joint dependencies from the data explicitly. The spatial attention mechanism considers all joints of the same timestep. Moreover, the projections we use to calculate the key and value are shared across joints. Since we aim to identify the most relevant joints, we project them into the same embedding space and compare with the joint of interest.

Joint Predictions Finally, the joint prediction j^t+1(n)\boldsymbol{\hat{j}}_{t+1}^{(n)} is obtained by projecting the corresponding DD-dimensional embedding et(n)\boldsymbol{e}_{t}^{(n)} from the LL-th attention layer back to the MM-dimensional joint angle space. Like , we apply a residual connection between the previous pose and the prediction.

3 Training and Inference

At test time, we compute the prediction in an auto-regressive manner. That is, given a pose sequence {x1,…,xT}\{\boldsymbol{x}_{1},\dotsc,\boldsymbol{x}_{T}\}, we get the prediction x^T+1\boldsymbol{\hat{x}}_{T+1}. Due to memory limitations, we apply the temporal attention over a sliding window of poses which we set as the length of the seed sequence. In other words, to produce x^T+2\boldsymbol{\hat{x}}_{T+2} we condition on the sequence {x2,…,x^T+1}\{\boldsymbol{x}_{2},\dotsc,\boldsymbol{\hat{x}}_{T+1}\}.

Experiments

We evaluate the ST-Transformer on AMASS and H3.6M in Sec. 4.1 following the standard protocols for short-term predictions and adopting distribution-based metrics for long-term predictions. We note that both benchmarks focus on modeling 3D joint angles unlike the 3D position-based benchmarks presented in the previous work . We argue that modeling the 3D position representation of human pose is prone to errors as the models are free to violate the skeletal configuration. In other words, the outputs may contain artifacts such as inconsistent bone lengths across the frames . In contrast, the joint angle representation implicitly preserves the skeletal structure which is important in many downstream tasks.

Sec. 4.2 and Sec. 4.4 show qualitative results and attention weights, thus providing insights into how the model forms predictions. We validate design choices through ablation studies in Sec. 4.3. We run our models and the baselines on various joint angle representations including rotation matrix, quaternion and angle-axis, and report the best performance. Implementation details for our model and the baselines are provided in the supplementary material.

AMASS We follow Aksan et al. and evaluate our model on the large-scale motion dataset AMASS . Table 1 summarizes the results with pairwise angle- and position metrics up to 400400 ms. For longer time horizons, direct comparison to the ground-truth via MSE becomes increasingly problematic particularly for auto-regressive models . Hence, in addition to the standard metrics in , we conduct further analysis by using complementary metrics in the frequency domain allowing benchmarking up to 1515 seconds (Fig. 4).

In the short-term evaluations, we compare our ST-Transformer with the vanilla Transformer, previously reported RNN-based architectures and two DCT-based architectures . We could not compare to as no implementation is publicly available. The vanilla Transformer follows an auto-regressive approach similar to our ST-Transformer but applies 1D attention on the pose vectors. Furthermore, we include results of a variant of our model that does not use softmax σ\sigma in the attention (cf. Eq. 1) as the softmax may lead to gradient instabilities. Instead, we normalize by the sum of all attention scores similar to . Our ST-Transformers achieve state-of-the-art in all metrics while LTD-Attention remains competitive at 400 ms.

Long-Term Our approach is fully generative, so while it maintains local consistency, the global positions may deviate from the ground-truth under natural variation. Hence, we propose to use distribution-based metrics in the power spectrum (PS) space proposed by Hernandez et al. instead of a direct comparison with the ground-truth frames. We report two metrics, (i) PS KLD which measures the discrepancy between the prediction and the test distributions via the KL divergence, and (ii) PS Entropy capturing the entropy of the prediction distribution in the power spectrum. The latter measures how diverse the predictions are. Models collapsing to static pose predictions end up with lower entropy values. However, PS Entropy can be deceived by random predictions and it should thus be interpreted in conjunction with the PS KLD where only a set of predictions that are similar to the real data samples can achieve a lower score.

In Figure 4, we compare our ST-Transformer, the vanilla Transformer, RNN-SPL and LTD-Attention with DCT representations on these PS metrics for predictions up to 1515 seconds. To produce long sequences with LTD-Attention, we run it autoregressively on its own predictions and in a sliding window fashion. The reference values (dotted, purple) from the training and test samples are calculated on randomly extracted 1-second windows (60 frames). Similarly, we compute statistics over the predictions of the respective model by shifting a 1-second window. Thus, we compare every second of the prediction with real 1-second clips. As is expected, the prediction statistics do deviate from the ground-truth statistics with increasing prediction horizon. However, our model remains much closer to the real data statistics than any of the baselines. This indicates that our model does generate plausible poses that are similar to the data distribution without memorizing the exact sequences.

The PS Entropy plot in Fig. 4 shows that the ST-Transformer has a higher entropy than the baselines, thus indicating its power to mitigate the collapse to a static pose. It furthermore indicates that none of the baselines can alleviate this problem as much as the ST-Transformer. The difference is more pronounced with horizon length. These observations are additionally corroborated by our visualizations in the supplementary video and Fig. 6 where we clearly see that the walking motion produced by the baselines phase out earlier compared to our ST-Transformer’s output.

H3.6M Traditionally, motion prediction has been benchmarked on H3.6M . Tab. 2 compares our model in this setting, where we are competitive and often achieve state-of-the-art. H3.6M is roughly 14 times smaller than AMASS and its test split consists only of a few sequences, which has been reported to cause high variance . Furthermore, as is evident from Tab. 2, improvements are often marginal and recent works seem to converge to the same error for most actions. For these reasons we argue that the AMASS benchmark introduced by carries more weight.

Discussion It is evident that the spatio-temporal decoupling of attention is indeed beneficial when compared to the vanilla Transformer. In all settings our ST-Transformer significantly outperforms the vanilla counterpart. Compared to the RNN-based autoregressive baselines such as RNN-SPL, Seq2seq or AGED, our model makes more accurate predictions in short-term horizon as well as plausible longer-term generations by mitigating the error accumulation problem.

The DCT-based baselines are the most competitive and also conceptually more relevant to our work. We argue that the task favors the DCT-based representations as a temporal window is encoded and decoded in one go. In other words, the entire prediction horizon is predicted at once in contrast to our frame-by-frame predictions. Hence, we observe that our model shows strong performance in the shortest prediction horizon with respect to the pairwise comparisons with the ground-truth on both AMASS and H3.6M (cf. Tab. 1, 2). Our model’s error with respect to the ground-truths seems to increase with longer prediction horizons, which is expected for an auto-regressive model due to error accumulation. Yet, it remains very competitive and the distribution-based metrics in Figure 4 also highlight that our model is statistically closer to the real data in very long-term predictions.

2 Qualitative Evaluation

Here, we evaluate the generative capabilities up to 2020 seconds. We feed the model with a particular motion sequence of 22 seconds and auto-regressively predict beyond its training horizon (i.e., 400400 ms or 11 sec).

We qualitatively compare our model with the vanilla Transformer, RNN-SPL and LTD-Attention on a walking sample from AMASS in Fig. 6. With RNN-SPL any variation quickly disappears within 5 seconds. The vanilla Transformer does not reach a static pose for longer, which shows the benefits of the attention mechanism over recurrent networks. However, it still collapses around second 1515, whereas our ST-Transformer maintains the walking motion over the entire duration of 2020 seconds. Also the LTD-Attention model produces walking motion longer than RNN-SPL, but still converges before 10 seconds. More samples can be found in the appendix and video.

While our model performs well on periodic motions for long horizons, the prediction horizon is limited to a few seconds for aperiodic motion types as the motion cycle is completed. This still exceeds previously reported horizons significantly. Also, it is not unexpected since the model is unlikely to be exposed to transition patterns as it is trained on rather short 22-second windows.

3 Ablations

2D Attention To unpack our contribution more clearly we implement and compare to a 2D transformer architecture. Here, every joint attends to all other joints across all frames (Fig. 5). The complexity of the attention thus becomes O(N×T)O(N\times T). With this design, memory requirements increase drastically. Hence, for training we either reduce the model complexity or decrease TT or the batch size. The performance of the best configuration we found is summarized in Tab. 6. It clearly falls behind the ST-Transformer. Our approach reduces the complexity from O(N×T)O(N\times T) to O(N+T)O(N+T), allowing to attend to a longer history and to use a larger model, highlighting our contribution on an architectural level.

Number of Layers We train our model with varying number of attention layers LL. Fig. 7 shows that competitive performance on AMASS is reached with only 33 layers. As the number of layer increases, the representations learned by our model are tuned better. It can be considered as the number of message passing steps to update the available representation.

4 What does Attention Look Like?

First, we observe that there is a diverse set of attention patterns across heads, enriching the representation through gathering information from multiple sources. Second, the temporal attention weights reveal that the model is able to attend over a long horizon. While some heads focus on near frames, some look to the very beginning which would be difficult for an RNN. Finally, in the spatial attention, we observe joint-dependencies not only on the kinematic chain but also across the left and right parts of the skeleton. For example, while predicting the left knee joint, the model attends to the spine, left collar and right hip, knee and collar joints. Similarly, for the right elbow, the most informative joints are spine, the left collar, the hips and knee joints. To see how attention weights change over time, please refer to Fig. 7 in the appendix and the video.

Conclusion

We introduce a novel spatio-temporal transformer (ST-Transformer) network for generative modeling of 3D human motion. Our proposed architecture learns intra- and inter-joint dependencies explicitly via its decoupled temporal and spatial attention blocks. We show that the self-attention concept can be very effective in learning representations for both short- and long-term motion predictions compared to the DCT-based motion representations. Furthermore, it mitigates the long-term dependency issue observed in auto-regressive architectures and is able to synthesize motion sequences up to 2020 seconds conditioned on periodic motion types such as locomotion.

This project has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme grant agreement No 717054.

References

Appendix A Experimental Details

We implemented our models in TensorFlow . Hyper-parameters we used for the experiments are listed in Tab. 4. Due to the limited data size of H3.6M, we achieved better results by using a smaller network as larger models usually suffered from overfitting.

As suggested by Vaswani et al. , the Transformer architecture is sensitive to the learning rate. We apply the same learning rate schedule as proposed in . The learning rate is calculated as a function of the training step as follows:

where DD is the joint embedding size. Warmup is set to 1000010000 for our AMASS and H3.6M datasets. We use a batch size of 3232 and the Adam optimizer with its default parameters. Before updating the parameters, gradients are clipped with respect to the global gradient norm (i.e., clip by global norm) with a maximum norm value of 1.01.0.

Following the training protocol in , we apply early stopping with respect to the joint angle metric. Since our approach does not fall into the category of sequence-to-sequence (seq2seq) models, we use the entire sequence (i.e., seed and target) for training. On H3.6M, the temporal attention window size is set to 7575 frames (i.e., 22-sec seed and 11-sec target at 2525 fps). On AMASS, we fed the model with sequences of 120120 frames (22-sec seed at 6060 fps) due to memory limitations. We followed an auto-regressive approach and train our model by predicting the next pose given the frames so far. In other words, the input sequence is shifted by 11 step to obtain the target frames.

Each attention block contains a feed forward network after the temporal and spatial attention layers. This feed forward network consists of two dense layers where the first one maps the D=128D=128 dimensional joint embeddings into 256256-dimensional space for AMASS (128128-dimensional space for H3.6M), followed by a ReLU activation function. The second dense layer always projects back into the DD-dimensional joint embedding space. We use the same dropout rate of 0.10.1 for all dropout layers in our network.

Data represenations 33D joint angles can be represented with various representations such as angle-axis , quaternion or rotation matrix . In our experiments both on the AMASS and H3.6M datasets, we trained and evaluated our model on all three joint angle representations. Similarly, we also evaluated the baseline models LTD and LTD-Attention (cf. Sec. F and Sec. G) by using all the joint angle representations and reported the best performance.

We found that all models achieved their best performances with the rotation matrix representation (see Tab. 5 for our model’s results). Since the model’s predictions might not be valid rotation matrices, we project them to the nearest valid rotation matrix in SO(3) using the singular value decomposition. The evaluation metrics are then computed on the projected rotation matrices.

Weight sharing In our initial design both temporal and spatial attention were alike (i.e., following Eq. (3)). We experimentally found that our current design, i.e., sharing the key and value weights across joints but keeping the query separate, achieves better performance. We keep the temporal attention as is and compare our current design (i.e., Eq. (5)) with two alternatives: projection weights WQW^{Q}, WKW^{K}, WVW^{V} are (A) separate (i.e., joint-wise as in Eq. (3)) or (B) shared across joints. On AMASS at 400400ms, (A) achieves 0.5110.511 Euler error and (B) 0.5040.504 (vs 0.4900.490 with the current design).

Data augmentation Finally, we observed a benefit of data augmentation on the H3.6M dataset. With a random chance of 0.5, we reverted or mirrored a sample sequence. Mirroring the joints was previously reported to be useful in . For the former one, we hypothesize a reverted sequence still posses valid human poses and backward motion dynamics, which may help the model to minimize the null space.

Appendix B Multi-head Attention

In the temporal attention blocks, we use separate query, key, and value weight matrices for different joints. In the spatial attention blocks, while the key and value weight matrices are shared across joints, the query weight matrices are not. Having obtained all the query, key, and value embeddings, we can get the output of each head according to Eq. (4) in the main submission. Note that the attention is over TT time steps in the temporal attention blocks and NN joints in the spatial attention blocks. Finally, the outputs of all heads are concatenated and then fed to a feed forward network consisting of two dense layers and computing the updated embedding of Eˉ(n)\boldsymbol{\bar{E}}^{(n)}.

Appendix C What does Attention Look Like?

In order to adapt to changing spatio-temporal patterns, our model calculates the attention weights at every step. In Fig. 8, we visualize the change in the attention weights over time. The focus of the network changes as expected. Although the attention window shifts in time, we observe that for quasi-static joints like the hips or spine, the model maintains the focus on the same time-step in this particular attention head. When we calculate the inter-joint dependencies via spatial attention, a large number of joints attend to the left and right collars. This could indicate that such mostly static joints are used as reference.

Appendix D Hyper-parameters

We experiment with varying number of attention heads HH. As shown in Fig. 10(a), the best performance on AMASS is achieved with 88 attention heads. A model with 22 attention heads also yields reasonable performance. Comparison between the performance of multi- and single-head attention mechanism suggests that the model benefits from using more than one spatio-temporal configuration.

Fig. 10(b) plots the performance of our model when trained with seed sequences of varying length, showing the performance w.r.t. the temporal attention window. The decreasing trend suggests that our model benefits from longer sequences. This hypothesis is supported via the temporal attention masks showing that our model accesses poses from the beginning of the sequence (cf. Fig. 8).

We train our model with varying number of layers. Fig. 10(c) shows that reasonable performance on AMASS is reached with only 33 layers. However, as the number of layer increases, the representations learned by our model is tuned better. It can also be considered as the number of message passing steps to update the available representation.

Appendix E Additional Ablation on 2D Attention

To provide more insights into the computational efficiency of our ST-attention compared to the naive 2D attention discussed in Sec. 4.3, we present several additional comparisons in Tab. 6.

In rows 1-3 of Tab. 6, we maximize one of the three hyper-parameters for the 2D attention where row 1 exhausts the batch size and row 2 the window size at the cost of the batch size. We observe that prioritizing one of the three parameters is usually detrimental for the performance of the 2D attention model and we get the best result with a trade-off between the three (cf. rows 4, 6). Furthermore, we also observe that those best 2D configurations are always outperformed when switching to our decoupled ST attention (cf. rows 5, 7). Row 8 reports the best configuration we found with our ST-attention. These results show that our decoupled attention mechanism is superior to the plain 2D attention even when controlling for computational efficiency.

Appendix F Evaluation of LTD on AMASS

In Tab. 1 of the main paper we report the performance of the LTD model (cf. LTD-10-10 entry) on the AMASS dataset. The results correspond to the best results we obtained after hyper-parameter-tuning, explained in more detail in the following.

To train and evaluate LTD on AMASS we use the code provided by Mao et al. but swap out the data pipeline to load AMASS instead of H3.6M. In the inputs to the network and the outputs are 400 milliseconds (10 frames) worth of data. As AMASS is sample at 60 Hz, we hence pass 24 frames as input and let the model predict 24 frames.

We fine-tune the learning rate as well as the number of DCT coefficients. For the number of DCT coefficients we use the original 35, but also the maximum number of coefficients 48. As reported by we find this to make little difference, but 35 coefficients led to slightly better results.

The remaining hyper-parameters are as follows, which mostly corresponds the original setting. We are employing the Adam optimizer with a learning rate of 0.001 and batch size of 16. We train for maximum 100 epochs with early stopping. The learning rate is decayed by a factor of 0.96 every other epoch. The input window size is 48. The model parameters are kept as proposed in the original paper resulting in a model size of roughly 2.232.23 Mio. parameters.

Appendix G Evaluation of LTD-Attention on AMASS

To evaluate the LTD-Attention model on AMASS, we again use the publicly available code and plug in our AMASS data loading pipeline. We adjust the number of input and output frames to reflect the different framerate on AMASS, i.e. we feed seeds of length 120120 (22 seconds) and predict 6060 frames (11 second). We have found that this resulted in better performance than predicting 2424 frames (400400 milliseconds) directly.

We keep the Adam optimizer with the originally proposed learning rate decay, but fine-tuned the learning rate to 0.0050.005 with a batch size of 128. Similarly, we found 4545 DCT coefficients to work best and we kept the kernel size at the original value of 1010. The model is trained for 100100 epochs and all other model parameters are kept the same with the exception of the size of the inputs and outputs. Like for our model, we also tried out different joint angle representations, of which rotation matrices performed best.

Appendix H Power Spectrum Metrics

On AMASS, we show our model’s capability in making very long predictions (i.e., up to 15−2015-20 sec). Such long prediction horizons prevent us from using pairwise metrics such as the MSE because the ground-truth targets are often much shorter than the prediction horizon. Hence, we use Power Spectrum (PS) metrics originally proposed by Hernandez et al. .

In order to adapt the metrics into our new setup, we slightly modify the evaluation protocol. We use 3D joint positions instead of angles as it is straightforward to convert any angle-based representation into positions by applying forward kinematics. This also allow us to compare models operating on arbitrary rotation representations. Given a sequence X\boldsymbol{X}, we treat every coordinate of every joint over time as a feature sequence xf\boldsymbol{x}_{f} following . The power spectrum PS is then equal to PS(xf)=∣∣FFT(xf)∣∣2PS(\boldsymbol{x}_{f})=||FFT(\boldsymbol{x}_{f})||^{2} where FFTFFT denotes the Fast Fourier Transform.

where X\mathcal{X} is either the ground-truth test or training dataset, or the predictions made of a respective model on the corresponding test dataset. ff and ee correspond to a feature and frequency, respectively.

PS KLD

To compute the PS KLD metric, we use the following approach. Instead of using the variable-length ground-truth targets, we randomly get 20′00020^{\prime}000 sequences of length 11 sec (i.e., 6060 frames) from the test dataset and calculate the power spectrum distribution GG. Then, we get non-overlapping windows of 11 sec from the predictions to get PtP_{t} where tt stands for the corresponding prediction window of length 1 second. For example, P5P_{5} is the power spectrum distribution for the predictions between 55 and 66 seconds. This enables us to measure the quality of arbitrarily long predictions by comparing every second of the predictions with the real reference data.

The symmetric PS KLD metric is then defined as

We use publicly available implementations of the metrics (, Github link). It is worthwhile to mention that the PS KLD metric does not make pairwise comparisons between ground-truth and predictions. Instead, it measures the discrepancy between the real and predicted data distributions.