Convolutional Autoencoders for Human Motion Infilling

Manuel Kaufmann, Emre Aksan, Jie Song, Fabrizio Pece, Remo Ziegler, Otmar Hilliges

Introduction

Modeling 3D human motion has seen increased attention from the computer vision community as it bears the potential to benefit many downstream tasks in robotics, autonomous driving, or human-computer interaction. Previously, a large body of work has focused on predicting how a given seed sequence evolves in the future . While this kind of open-ended motion prediction from past observations has its use cases, there is a significant interest in predicting motion when additionally a sequence in the future is available. For example, although commercial motion capture systems boast high accuracy, optical tracking systems still struggle with occlusions, especially when multiple people in close contact or objects are involved. In other cases, due to system failure, damaged tracking sensors, or trackers falling off the actors, valuable data might be lost during a capture session. Hence, a method to recover missing frames, to complete partially observed poses, or predict missing joints over an entire sequence, is an invaluable tool for post-processing of motion capture data. Also, keyframe interpolation, a problem that has long been studied in computer graphics, could benefit from such a method. Automatically filling in large gap sizes is of great interest to animators and artists alike, as it can drastically reduce the amount of required keyframes and hence manual artistic interventions.

This problem of filling in motion between known sequences is sometimes referred to as motion infilling and can be defined as completing the gap between an available start sequence in the past and an end sequence in the future (cf. Figure 1 or 2 for a visual depiction). Ideally, the gap in between is completed in such a way that the resulting motion is plausible, natural, and the transitions at either end of the gap are smooth. Motion infilling can also refer to the case where a few joints are missing over the entirety of a sequence or when only partial poses are observed.

In this paper we present a simple yet effective convolutional autoencoder (CAE) trained to complete long gaps in human motion sequences producing compelling and smooth transitions without further post-processing steps. Our proposed model is fully convolutional and able to synthesize motion for varying gap sizes and multiple gaps occurring in the same sequence. Furthermore, it can also be used to blend together partially observed poses within the gap, or single joints missing over the entire sequence. The proposed method lends itself naturally to the removal of other types of noise as well, for example Gaussian noise, and as such can be viewed as a general de-noising framework.

Motion infilling is challenging as it requires learning smooth transitions between possibly different types of motion. Also, recent success in open-ended motion prediction does not immediately translate to the infilling task, as the dominant method of choice has been recurrent neural networks (RNNs), which are notoriously difficult to condition to future sequences. For example, Berglund et al. show how a bi-directional structure can be used for the task, but the supported gap size is fixed and short. Very recently, Harvey et al. introduced time-to-arrival embeddings to a recurrent structure to support variable gap sizes. However, the authors also report the use of inverse kinematics to post-process the model’s outputs.

To avoid the limitations of recurrent models in motion modeling, we suggest to cast motion infilling as an image inpainting problem. Here, the “image” corresponds to a matrix-based representation of the motion data, where each column contains a single pose and each row stands for a joint’s information over time. Treating motion sequences as matrices has been studied before, e.g. . The problem arising when using such representations is that neighboring “pixels” are not necessarily neighboring joints in the skeletal hierarchy. Hence, capturing the intricate spatial dependencies is not straightforward, especially for convolutional architectures that usually excel at exploiting spatial relationships from images. Previous work addresses this by using dense layers spanning the entire height of the input matrix , using large kernel sizes , or resorting to 1D convolutions over the temporal domain only .

In this work we offer a different take on this problem and propose that a model with a sufficiently large receptive field is able to capture the spatial relationships. We achieve this by designing a convolutional auto-encoder architecture inspired by the VGG model . This way, the model covers a large number of joints in the upper layers and is furthermore capable of learning temporal dependencies to produce smooth motions without collapsing to a safe mean pose. We train our model by using a reconstruction objective only and do not require any additional regularization terms or adversarial training. To enable variable-length infilling, we employ a curriculum learning scheme in which the inputs are perturbed with increasingly large gaps during training. This results in a model that can easily fill in up to 120 frames (2 seconds) between short given sequences (i.e., typically 30–40 frames but also as little as 5).

Our proposed model allows us to train a single model encompassing a large variety of motion types like walking, jumping, kicking and more, while still fulfilling the constraints defined by the known sequences as closely as possible and without requiring post-processing. In doing so we show that a simple but effective convolutional model can indeed be used to unravel the complex spatio-temporal dependencies in 3D human motion data.

In summary, this paper contributes an easy-to-train and efficient CAE for the task of 3D human motion infilling. The proposed model is able to synthesize convincing looking motion to fill in variable-length gaps between sparsely distributed known sequences of various motion types. Unlike previous work we push the boundaries of generating smooth and natural looking transitions using the network’s outputs directly and thus foregoing post-processing steps or regularization terms. The resulting model can be applied to a variety of infilling and other noise removal tasks as demonstrated by our evaluations.

Related Work

Modeling human motion has been of interest to the machine learning community as an instance of time-series modeling, with early works by Taylor et al. employing conditional restricted Boltzmann machines. In more recent years, the computer vision community has often adopted RNNs to address the task of predicting motion into the future given a seed sequence . To avoid a collapse to a mean pose, researchers have suggested to perturb the input at test time with Gaussian noise , removing joints entirely , training the model on its own outputs , or employing a specialized output layer that follows the kinematic chain of the skeleton . Martinez et al. propose a residual connection in a sequence-to-sequence architecture to address the discontinuity between the last known and first predicted pose. Wang et al. further improve upon smoothness and realism of the predictions by employing adversarial and geodesic-inspired losses.

While recurrent models lend themselves well to open-ended motion prediction, they are less suitable for the task of motion infilling studied in this paper. An example of a recurrent structure used for infilling of joint trajectories is presented by Berglund et al. in the form of a bidirectional RNN. However, the presented method uses a fixed gap size, is only demonstrated for small gaps, and is rather expensive. Only recently Harvey et al. have shown the application of a recurrent structure to motion infilling for larger gap sizes. However, in the authors mention that post-optimization is used to close discontinuities in the transition from the last predicted frame to the first known frame of the target sequence. is a recent follow-up work. While demonstrating impressive results, it mentions the use of inverse kinematics as a post-processing step. In this work, we focus on exploring the capabilities of a neural network alone to produce smooth and natural motion for the task of infilling, i.e. without relying on regularization or post-processing.

Non-recurrent Models

While non-recurrent models have prominently been used for time-series modeling of audio or natural language (e.g., ), their adoption for human motion data is not straightforward. This can be explained by the fact that in addition to temporal dependencies, also highly complex spatial dependencies related to the dynamics of human motion must be captured.

Mao et al. were the first to apply a graph convolutional network (GCN) to the task of motion prediction. The GCN operates on discrete cosine transforms of the input data, which makes it more difficult to support variable length inputs. In , Bütepage et al. also represent the data as matrices similar to our work. They then use dense layers in an autoencoder framework to learn meaningful representations, which can be leveraged for prediction and infilling. In contrast to our work, their dense layers span the entire spatial domain of the input matrix, whereas we use small convolutional filters. This allows us to harness the efficiency and computational advantage of CNNs, while still capturing the spatio-temporal dependencies of the data.

Convolutional models have been applied to motion prediction by Hernandez et al. and Li et al. . Both works represent motion sequences in matrices, like we do in this work. To apply convolutions to this data representation, all approaches must deal with the fact that neighboring values in the data matrix are not necessarily neighboring joints in the human skeleton . do so by employing dense layers that span the entire height of the input matrix. While we agree with that the lack of spatial proximity in the inputs is an additional challenge for any CNN, we show in this work that with a sufficiently big receptive field, this is no hindrance for successful motion modeling with CNNs. Given a deep enough model, a CNN with 2D filters spanning both spatial and temporal dimensions is indeed able to unfold the underlying spatio-temporal relationships.

Like in this work, Li et al. use convolutions directly on the input matrix. However, they use bigger filter sizes and do not apply their architecture to motion infilling. Also, we refrain from using any adversarial losses or other kinds of regularizers (e.g. bone length constraints in ) and instead demonstrate that our model discovers desirable properties such as bone length consistency and smoothness from the data alone without external guidance other than the reconstruction loss.

Manifold Learning and De-noising

CNNs have been used to learn the projection of human motion onto a low-dimensional manifold. Drawing samples from this manifold then generates plausible human poses. Holden et al. use such an approach to recover full poses from partial inputs such as corrupted motion-capture data and to fill gaps of fixed sizes (i.e., max 15 frames). A custom un-pooling operation requires careful layer-by-layer training. In follow-up work a similar but shallower architecture is integrated into a hierarchical system that maps from foot-fall patterns and trajectories to pose sequences. However, a separate model has to be trained for each activity, thus this approach is not well suited to interpolate between sequences of different activities. We also leverage convolutional auto-encoders but train a single model for all motion types, predict both poses and rotational and translational velocities and can handle transitions between different activities.

As our method cannot only be used to fill in missing frames, but also to clean up other types of noise such as Gaussian, it is related to the work of Holden . Although Holden’s architecture operates on marker level and not joints, our method is in theory directly applicable. In contrast to our work, uses a frame-wise ResNet-based architecture and requires post-optimization to induce smoothness and improved pose reconstruction.

Motion Modeling in Computer Graphics

A variant of motion infilling, known as keyframing, is a long-standing problem in the computer graphics community with the latest addition provided by . Keyframing aims at reducing the amount of tedious manual labor for artists to create life-like animations without relinquishing control over the resulting animation. Early works date back to the 1980s and many other approaches with the same goal have been proposed since. Such approaches fall into the field of physics-based character animation , space-time optimization or motion graphs . While our method can be understood as a keyframing or general character animation tool, in this work we do not intend to provide a production-ready tool. Instead, we focus on the more basic building block of synthesizing convincing human motion from neural networks alone.

Image Inpainting

Method

In this section we detail the architecture, training, and inference procedure for our model. We propose a deep convolutional de-noising auto-encoder M\mathcal{M}, trained to fill in missing 3D pose information represented as image-like joint position maps (cf. Figure 3). That is, given the constraints for joint positions at the beginning and end of a motion sequence, our model reconstructs the missing frames in between in a visually coherent way.

We cast the infilling problem as an image inpainting task, where the image corresponds to the joint position map. To do so, we must address the problem that joints that are close in image space are not necessarily spatially close in the skeleton hierarchy, which seems to go against the basic assumption of CNNs. Nonetheless, we show that given a sufficiently deep model and sufficiently large receptive field, a convolutional autoencoder with small filter sizes in the spatial and temporal domain is indeed able to capture the underlying spatio-temporal structure of the data. We achieve this by following well-established design rules, i.e. stacking several convolutional layers (with 3×33\times 3 filters) and pooling operations inbetween. This is in contrast to previous work that either resorts to dense layers to encode the spatial relationships , uses larger filter sizes or temporal convolutions only .

Furthermore, in order to forgo an explicit modelling of the skeletal structure we propose a curriculum learning training scheme that aids the model in learning the spatial dependencies between joints, via randomized removal of joint information. In consequence, the model is robust against different types of perturbations in the inputs and thus it can be used for different tasks besides motion infilling, such as noise reduction. We show various applications of our method in Section 4 and in the accompanying video.

Here, we briefly introduce the data representation used as input to our model. For an overview please refer to Figure 3.

2 Model Architecture

Our proposed architecture is inspired by and the VGG model design . An overview is illustrated in Figure 4. The main reasons speaking for such a deep architecture are larger model capacity, more non-linearities to model inherent complexities of the task and increasingly large receptive fields. The latter allows the model to implicitly learn dependencies across large spatial distances (in image space) . In our case this means that the model can implicitly recover the joint dependencies in the skeleton without requiring a pre-determined spatial prior.

As shown in Figure 4(a), the encoder network stacks 5 encoding units, where each unit consists of two 3×33\times 3 convolutional layers whose outputs are activated using a leaky ReLU. Like in the VGG model, the encoding units (cf. Figure 4(b), left) are designed to increase the receptive field while keeping the increase in trainable parameters minimal. Via max-pooling, the output of each encoding unit is halved in both the spatial and temporal dimension before being passed into the subsequent unit. The decoder network mirrors the architecture of the encoder (Figure 4(b), right), but the two are decoupled, i.e. do not share weights. To revert the pooling operation in the decoder, we use strided de-convolutions.

While our model looks similar to the motion manifold network proposed by Holden et al., we would like to highlight the following two key differences. First, the autoencoder layers proposed by use 1D convolutions over the temporal domain, while we employ small 2D filters. Second, propose a custom un-pooling operation, denoted as Ψ†\Psi^{\dagger} in . In our experiments, stacking several of those layers has proven to be difficult to train in an end-to-end fashion. In layer-wise training is employed to train the autoencoder. Still, reports that, even if trained successfully, stacking multiple layers blurs out the results and hence the authors fall back to a single-layer autoencoder. Our model uses strided de-convolutions to replace Ψ†\Psi^{\dagger}, which enables straight-forward end-to-end training of multi-layer architectures and leads to improved performance (cf. Section 4.1).

For a typical input sample of size 69×24069\times 240, the resulting embedding in the latent space (yellow box in Figure 4(a)) corresponds to a 3×8×2563\times 8\times 256 tensor. Obviously the question whether such a compact representation can accurately represent the subspace of valid poses arises. To this end we refer to our experimental results in Section 4 indicating that our architecture can indeed model the sub-space well. Moreover, due to the deep architecture the effective size in the upper layers is comparatively large and hence contributes to capturing long-range relationships.

3 Training

The task of motion infilling is now cast as the problem to fill in missing gaps within the “images” Xi\boldsymbol{X}_{i}. Note that this is different from the usual application scenario of de-noising autoencoders, where inputs are typically assumed to be corrupted by some stochastic diffusion process. In our case we remove entire blocks of complete columns from the inputs (see Figure 6) so that the model is forced to learn how to generate complete, valid poses.

During training we present the model with pairs of complete and incomplete samples, and we instruct it to reconstruct the data to closely match the complete counterpart. To induce robustness against the amount of perturbation we employ a curriculum learning training scheme . During curriculum learning the model is initially trained with samples that allow the task to be easily completed, and subsequently it is exposed to increasingly harder cases. In our case, this translates into increasing the span of masked ranges along the temporal dimension. In this way the model is exposed to more realistic representations of the problem, and consequently is capable of generating plausible poses given drastically different types of inputs. In particular, the model is robust to the length of the gap, to random noise applied to inputs and to missing joint information.

We parameterize the choice of the mask Mi\boldsymbol{M}_{i} by a pair of scalars (λ,τ)(\lambda,\tau), where λ\lambda and τ\tau determine the length and location of the gap, respectively. In order to increase the variation of the inputs we generate the masks by using normal and uniform distributions such that:

where μe{\mu}_{e} and σ2{\sigma}^{2} parametrize the gap size distribution at training epoch ee. Please note, that while μe{\mu}_{e} is gradually increased during training, σ2{\sigma}^{2} is kept fixed. We initialize μe\mu_{e} to 1010, and increase it until 120120 by steps of 1010 every 55 epochs. Additionally, we randomly alternate between the previously described type of distortion and masking out 1,2,1,2, or 33 random joints over the entire length of the sequence.

We apply zero-mean and unit-variance normalization on each row of Xi\boldsymbol{X}_{i}, where the mean and standard deviation are calculated over the training dataset. We use the Adam optimizer with a fixed learning rate of 0.0010.001 and a batch size of 8080. For all convolutional layers, Xavier initialization is used . We train for 200200 epochs which roughly takes 8 hours on an NVIDIA Quadro M6000 (12 GB).

Evaluation

In this section, we show quantitative comparisons and ablations (Section 4.1) and demonstrate the visual quality of our method (Section 4.2). Please also refer to the video for more visualizations.

To assess our method’s performance quantitatively we compute the mean 3D joint error for both reconstruction and infilling tasks. Table 1 compares our best model to three baselines: Interpolation refers to naive linear interpolation. Vanilla AE is the same as our best model but each encoding/decoding block contains only a single strided 3×33\times 3 convolutional layer, yielding a smaller receptive field. We also compare against an implementation of the motion manifold network presented in .

Finally, we also show the effect of curriculum learning on our best model. While the previous baselines were trained to reconstruct 60 frames (1 second) in clips of size 240, the last entry in Table 1 was trained to reconstruct up to 120 frames (2 seconds) with the curriculum scheme described in Section 3.3.

Table 1 shows that a simple application of is suboptimal. Our architectural changes described in Section 3.2 lead to improved performance (Vanilla AE, Ours). Furthermore, the difference in performance between the Vanilla AE and our model highlights the importance of the receptive field’s size. Lastly, the model trained using the curriculum loses some accuracy in the 60-frame and 0-frame (i.e., pure reconstruction) tasks. Despite this trade-off, Ours (curr.) still ranks second at least.

De-noising

To test our method’s noise reduction capabilities, we conduct two experiments. In the first one we perturb the inputs by adding Gaussian noise ϵ∼N(0,σ)\epsilon\sim\mathcal{N}(0,\sigma) to the joint positions. In a second experiment, we mask a joint with a probability of pp in each frame independently. Table 2 summarizes the 3D joint reconstruction error for varying values of σ\sigma and pp computed over 400 randomly chosen validation samples of length 240 frames. In removing Gaussian noise from the inputs, our best model is out-performed by the baseline. However, our model beats the baseline by a large margin when filling in randomly dropped joints. Note that neither models were trained specifically for these tasks. Furthermore, the 3D joint reconstruction metric does not necessarily correlate well with perceived smoothness or naturalness of the output motion. In the supplementary video we visualize several de-noised sequences and show that our model’s outputs are both smooth and plausible.

Bone Length Consistency

Our method does not use an auxiliary loss to enforce bone length consistency as we have found this to have little effect on the overall results. To investigate how much bone length variation we observe in the model outputs, we compute the bone lengths over the entire validation set for the pure reconstruction task. Figure 5 compares our best model to our implementation of Holden et al. . While our model clearly exhibits variance, the spread of predicted bone lengths is smaller compared to and the median of all predictions is typically closer to the ground-truth value.

2 Qualitative Results

In the following we present our experiments assessing the quality of the synthesized motions. Please note that all figures and the accompanying video show raw model outputs without any additional post-processing. Furthermore, the model predicts the entire sequence, i.e. the known subsequences are replaced with reconstructed poses. This allows the model to slightly alter the given frames if necessary in order to create smooth transitions between motions. In addition, the model also predicts the root trajectory.

To show our model’s performance in synthesizing coherent transitions between different actions we have produced a variety of results that blend several motion clips (see Figure 1 and more in the appendix). To this end, we have selected a few poses from various motion sequences (called seed sequences in the following) and chose a variable length of frames to be filled in by the model between each seed.

In Figure 1 and the video we show the results of our model on blending 1616 seed sequences from a test database of different motion categories. All seed sequences are taken from the holdout validation set. Since our model correctly handles variable-length sequences, we place variable gaps of sizes 5050–100100 between seed sequences. The necessary length of a seed sequence depends on the motion category. Experimentally we have obtained good results with ranges from 55 (one-legged jump) to 9090 (sitting). Typically the fewer examples of an activity are represented in the training data, the more key frames are necessary in order to constrain the model sufficiently. Figure 1 contains a total of 19271927 frames, of which 13001300 (roughly 67%67\%) were filled in by our model. Figure 6 shows an excerpt of the “image” that is fed to the model.

The missing frames are generated efficiently by performing one forward-pass through the network. It takes 2.82.8 seconds to generate 19271927 frames on a low-end GPU (i.e, NVIDIA GeForce GT 730, 2 GB memory), or 1.451.45 ms per frame. Higher end hardware (i.e., NVIDIA Quadro M6000, 12 GB memory), further reduces the inference time to a total of 0.580.58 seconds (0.30.3 ms per frame).

Tertiary Motion Blending

In addition to interpolating between key frames, our model can be used to blend in additional motion, e.g., to re-target an existing motion while constraining an end-effector. Here we blend a grabbing action into a plain walking clip. We introduce a gap of 6060 frames in the middle of a walking sequence from the test dataset, and then place joint positions of a right arm over the length of 1515 frames into the gap (inputs shown in inset, outputs in Figure 7). The joint positions for the arm are taken from another test clip and kept in its original representation, i.e., the local body coordinates. Note that it is necessary to provide such tertiary constraints over a number of frames, otherwise the model will treat the input as noise.

Recovering Joints

In a further experiment we evaluate the task of recovering one or more joints that are missing over the entire sequence. This corresponds to the scenario when a tracker (e.g., an optical marker) is lost during a recording session or certain joints are occluded. We mask different joints in various clips of length 240240 from our test dataset by removing entire rows, thus removing joints entirely over the duration of the motion sequence (see inset). Figure 8 and the accompanying video visualize the results. Despite significant loss of information, the model produces convincing looking motions, even if the missing joints are close in the hierarchy of the skeleton (e.g. elbow and hand).

Variable Gap Size and Amount of Context

A key feature of our approach is the capability to generate motion in between seed sequences with variable spacing. Here we evaluate the limits in terms of gap parameters. First, we take two seed motions and increase the gap size between the seeds incrementally. Second, instead of varying gap sizes we vary the number of given frames in the seed sequences.

For both experiments we choose walking and punching motions as seed sequences. Gap sizes 5, 20, 80, and 250 are applied with fixed seed sequences. Note that 5 and 250 are extreme values, with the latter being roughly twice the gap size seen during training. The model produces smooth and plausible motion for gap sizes 20 and 80. With gap size 250 the motion remains smooth, however towards the middle of the sequence, the model starts to converge into a mean pose before starting to prepare for the ensuing punch (see video, minute 02:47). Likewise extremely small gaps of only 5 frames in combination with drastically different source and target motions produces (overly) smooth motion and intermediate fine details are lost.

In terms of number of frames required per seed sequence we experimented with values of 1, 5, 10, 25, and 50 on either side of the gap and fixed the gap size to 8080 in line with the previous experiment. Again the model performs well under most of the configurations but the most extreme: with only a single seed frame on either end the model produces overly smooth motions and simply repositions the end-effectors. Not surprisingly, the more context we introduce, the better the predicted motion. In this particular case there was no more discernible difference in terms of motion details after 25 seed frames. Both experiments point to the same effect that not absolute numbers count, but the ratio between gap size and context provided by the seed sequences. We found this ratio should be approx. 1/61/6 for simple walking motions, 1/31/3 for transitions between different motion types and 2/32/3 for those containing rare activities such as ballet.

Further Results

We show results from the de-noising experiments in the appendix and the video. Also, please refer to the appendix for more infilling results between various motion types, including types that are rare in the data.

Limitations

While we see good pose predictions across a wide range of motions and task variations there are certainly failure cases and limitations. It is possible to successfully recover full motion cycles even if they are masked from start to end, such as a gait cycle, but only if the motion is recurring. For example, motions that are more complex and only occur once such as a full turn in-air that is masked entirely, will not be reconstructed since there is no information that would constrain the model to do so. Similarly, the reconstruction of motions that are very rare in the training data are prone to appear jittery.

Conclusion

In this paper we have proposed a deep convolutional autoencoder that learns to fill in large gaps in 3D human motion data. The method is capable of creating smooth transitions between drastically different motion sequences and generates plausible looking motion overall. The key idea in this approach is to cast the infilling task as an inpainting problem. We train the neural network on image-like representations of human poses where each column represents one time step in the sequence. During training we mask increasingly large blocks of the input data so that the model is forced to learn how to generate plausible pose data to fill in the gaps. A curriculum learning scheme is employed to enable prediction over variable gap length and to achieve robustness against various forms of noise in the inputs. We have evaluated our method in a number of experiments to illustrate its capabilities but also to identify its limits.

While the method produces plausible and natural motion without any further smoothing there are various areas for future work. In particular, we currently predict complete motion sequences including relative translational and rotational root velocities. Integrating these velocities over time results in a global root trajectory. Clearly for the method to be applicable in a production settings it would be necessary to provide control over the global trajectories. In the future, it would be interesting to decouple the root trajectory control from the pose generation process.

Similarly, although in our visualizations this does not seem to majorly degrade visual quality, our method cannot guarantee that bone lengths are always consistent or no foot skating artifacts occur. Future work could look into robustifying the method in this regard, especially so in the context of multi-person datasets.

We thank Janick Cardinale for his support. This project has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme grant agreement No 717054.

References

Appendix

Figure 12 illustrates a number of sequences that we have generated using our method. Note that all results were produced using the same model, which shows that the model produces plausible and smooth transitions for a variety of different motion types. Here short sequences selected form the test data where used to force transitions between different motion styles. Furthermore, we show examples from rare motions that are not well represented in the dataset such as dancing (12(c)) and lying on the floor (12(d)). Clearly the model does produce physically infeasible poses such as some floating above the ground before getting up. However, all results shown are raw predictions without any further processing that would enforce physical plausibility. Figure 12(f) shows a sequence composed out of a total of 55 motion sequences, where two sequences are used to force a grabbing and punching motion respectively. Please also refer to the video.

De-noising

To demonstrate the de-noising capabilities of the model we perform two sets of experiments (results best seen in the video). In the first experiment we perturb the joints of the input clip by adding Gaussian noise with zero-mean and unit-variance (see Figure 9(a)). In a second experiment, we mask a joint with a probability of 0.30.3 in each frame independently (see Figure 9(b)).

Note that even though our model is not trained specifically for this task, in both experiments it is able to recover a plausible motion out of the corrupted inputs (see Figure 10 and 11). However, we also note that this comes hardly as a surprise since auto-encoder architectures are routinely used for de-noising.