UniCon: Universal Neural Controller For Physics-based Character Motion
Tingwu Wang, Yunrong Guo, Maria Shugrina, Sanja Fidler
Introduction
Physics-based animation can provide more realistic motions and richer interactions with the environment compared to traditional keyframe-based animation methods. However, film and video game industries have not yet adopted physics-based animation since the available controllers are still very limited in the number of supported motions, animation quality, interactiveness and efficiency.
The field of physics-based animation has recently seen a plethora of data-driven techniques using deep reinforcement learning (RL). Data-driven techniques promise scalability to a wide variety of motions by learning directly from human demonstrations. Most of the RL physics-based animation algorithms fall under the umbrella of imitation learning (Liu et al., 2010; Liu et al., 2015; Merel et al., 2017; Peng et al., 2018a; Chentanez et al., 2018; Bergamin et al., 2019), where the reward signals are given based on the distance or similarity between the generated and the target motions. A policy network, which maps the current state of the character to the torques applied to the joints, is then optimized to minimize this distance. Policies trained with imitation learning have shown success in driving a virtual character to naturally follow the reference target motions in a physically plausible way.
However, existing physics-based RL controllers are extremely inefficient to train, requiring hours or days of training to reproduce even a single motion. Furthermore, most of the existing controllers have shown little diversity in motions, and usually lack robustness to perturbations in the environment. To increase the range of supported motions and allow for better user interaction, new methods are proposed by, for example, combining the controller with a powerful reference motion generator (Liu et al., 2016; Liu and Hodgins, 2018; Bergamin et al., 2019), or utilizing a hierarchical control system (Merel et al., 2018b, a; Peng et al., 2019; Park et al., 2019). However, these methods usually specialize towards certain applications and support limited number of motions. These controllers are typically trained and tested on the same motion dataset and have not demonstrated generalization to unseen motions.
In this paper, we propose a physics-based universal neural controller (UniCon) which greatly improves training efficiency, robustness, motion capacity and generalizability, allowing for a wide range of real-time interactive applications. UniCon consists of two components: 1) a low-level motion executor that generates physics-based control signal which drives the character to follow a target reference motion, and 2) a high-level motion scheduler which converts various high-level inputs (for example, keyboard commands) into a target reference motion. A powerful and robust motion executor is the key innovation of our work. We introduce several components that allow us to train our motion executor on large-scale motion datasets of diverse motion styles using reinforcement learning. In particular, we utilize a constrained multi-objective reward optimization, a motion balancer and a policy variance controller, which enable efficient and robust policy learning. Once the low-level motion executor is trained, UniCon can utilize different motion schedulers for real-time interactive applications. UniCon can be used to perform keyboard-driven control, compose user-specified motion sequences, and supports teleporting a person captured on video to a physics-based virtual avatar.
Our experiments demonstrate key improvements of UniCon over previous work:
Generalization: UniCon can be used to imitate motions which are unseen during training. Natural transition skills between motions are learnt automatically without the need of recording specific training samples in the dataset.
Robustness: UniCon produces robust control even when the source motion is of poor quality. UniCon demonstrates zero-shot robustness to environment obstacles that are not seen during training, such as projectiles. Our model can also adapt to characters with widely varying mass, or slower or faster motion than the motions seen during training.
Interactive Applications: Generalization and robustness allow many modes of interactive control, ranging from keyboard commands, noisy pose tracking from video capture, and user-specified sequences of locomotion or acrobatic motions, without having to retrain or fine-tune separate models for each application.
Learning Efficiency: Our learning algorithm has better sample efficiency and asymptotic performance compared to existing baselines.
Related Work
The problem of generating plausible human motion on the fly based on user input or a target goal is a long standing problem. We first discuss methods of motion generation which do not adhere to the laws of physics in Section 2.1, then cover physics-aware methods in both non-interactive and interactive settings in Section 2.2 and 2.3. We do not discuss all methods in character animation and refer readers to (Geijtenbeek and Pronost, 2012) for a comprehensive overview.
Keyframe-based animation systems or kinematic systems can be used to interactively animate virtual characters by exploiting pre-recorded motions or human authored data. While often enabling better interactive control compared to their physics-based counterparts, these systems result in motions that are not physically accurate and brittle to perturbations in the environment.
Motion graph (Kovar, 2002; Arikan and Forsyth, 2002; Lee et al., 2002), as a pioneer in interactive animation, connects motions from the dataset based on the character’s states and user specifications. However, the discretized connection in motion graph can result in non-smooth transitions, and the generated motion is not very responsive to changes in direction or perturbations. Follow up works improve motion graph by introducing parameterization, planning or semantic analysis (Kovar and Gleicher, 2004; Safonova and Hodgins, 2007; Min and Chai, 2012; Agrawal and van de Panne, 2016). In motion field (Lee et al., 2010), the authors use reinforcement learning to choose the most appropriate next motion. Motion matching (Buttner, 2015; Clavet, 2016) on the other hand, utilizes unstructured motion data to search for the most appropriate future frames given the character’s current states and user control input. Motion matching is considered by many as the state-of-the-art keyframe based animation technique due to its flexibility and ability to support a large variety of motions (Buttner, 2019; Harrower, 2018).
In (Grochow et al., 2004; Levine et al., 2012), more focus is given to motion synthesis by learning a latent variable model. With advances in data-driven modeling via deep neural networks, powerful generative motion models, such as phase-functioned neural network (PFNN) (Holden et al., 2017) and auto-conditioned recurrent neural network (Zhou et al., 2018) were proposed. PFNN trains a neural network on a labeled dataset, which predicts future character states based on the trajectory control signal and the character’s current states. The idea can be extended to quadrupeds with different walking phases such as dogs (Zhang et al., 2018). Furthermore, data-driven auto-regressive method is capable of animating character-scene interactions with impressive quality (Starke et al., 2019). In (Lee et al., 2018), the authors use a recurrent neural network to enable complex combinatorial basketball motion generation. In (Ling et al., 2020), an auto-regressive conditional variational auto-encoder is used to generate complex soccer motions. In (Kwon et al., 2020), model-predictive control is applied to generate locomotion step plans. Motion retargeting across different skeletons can be obtained via data-driven training with skeleton-aware operators, as shown in (Aberman et al., 2020).
While not physically plausible, keyframe-based algorithms are capable of producing a wide variety of complex motions, and can all be theoretically used as a high-level scheduler for our algorithm, which we showcase in our experiments.
2. Non-interactive Physics-based Methods
We first consider physics-based methods that aim to imitate a pre-recorded motion in a physically valid way without consideration of interactive control. One common approach is to formulate animation as an optimal control problem, or a reinforcement learning problem, which aims to minimize the distance between the reproduced and the recorded motions.
Some of the early attempts include (Liu et al., 2010), where the authors reconstruct and animate a diverse set of captured motions with randomized sampling, which is mathematically similar to the random shooting algorithms in model-based reinforcement learning (Richards, 2005; Rao, 2009; Wang et al., 2019). Other methods include (Tassa et al., 2012), which uses model-predictive control to obtain natural walking motions by specifying a reward function for walking. These methods all assume knowledge of the forward dynamics and can therefore be categorized as model-based reinforcement learning algorithms. The performance can be further improved by designing better sampling schemes as shown in (Liu et al., 2015; Hämäläinen et al., 2015). However, these model-based RL methods are usually time consuming to train. They are either not real-time during test time (Tassa et al., 2012), or require several hours of training time per motion through an iterative planning process (Liu et al., 2015). These methods are also more prone to fall into local minimas during complex motion planning and are less resistant to external perturbations, which makes them difficult to adapt to new environments.
In contrast, model-free reinforcement learning allows more capacity for turbulence and test-time efficiency. In generative adversarial imitation learning (GAIL) (Ho and Ermon, 2016; Merel et al., 2017), a neural controller is optimized with generative adversarial training (Goodfellow et al., 2014), where a discriminator is used to distinguish the target motion from the generated ones. The neural controller is trained iteratively to compete with the discriminator by generating motions that are closer to the target ones. In (Wang et al., 2017), GAIL is further extended with a variational auto-encoder, which encodes global motion information, such as style. Results demonstrate that the controller can reproduce walking motions with several different styles. In DeepMimic (Peng et al., 2018a), the authors show that deep model-free reinforcement learning can be used to train animation controllers for a wide range of skills. Their observation function includes the current states of the humanoid agents and a phase variable to encode time information of the frame. A policy network generates control signal to minimize the distance between the resulting next frame and the corresponding target frame. In (Peng et al., 2018b), the authors further show that DeepMimic can be used to reproduce motions captured by pose estimators. Similar to the original DeepMimic method, their method requires training with RL on the motion frame data and is not real-time. In (Merel et al., 2020), the idea is further extended such that the agent can use first-person perception as input observation. Model-predictive control is used to model realistic eye and head movements in (Eom et al., 2019).
The work most related to ours is (Chentanez et al., 2018), which feeds a set of future frames from a dataset into the network to produce control signal that tracks these future frames. Being one of the earlier attempts to obtain a multi-motion animation controller, this algorithm is not used for interactive control due to its performance limitations on motion imitation and lack of robustness. In our paper, we study and identify several key issues that prevent successful training of an animation system on large-scale motion datasets, and propose a technique to significantly improve performance.
Recently, several works on constructing complex task-driven long motion sequences using a hierarchical RL framework have been proposed (Heess et al., 2016; Merel et al., 2018b, a; Peng et al., 2019). These methods usually focus on solving one or two specific tasks, and the number of supporting motions required is usually quite limited.
3. Interactive Physics-based Methods
Interactive control is typically done via a keyboard or a gamepad input. Earlier attempts such as (Wu and Popović, 2010; Peng et al., 2017) aimed to solve bipedal locomotion for humanoids using a high-level footstep planner that takes control and terrain information as input. In recent work (Bergamin et al., 2019) (DReCon), complex interactive locomotion skills of the full humanoid are enabled, where motion matching is used as a high-level planner that sends future kinematics states to a proportional-derivative (PD) controller. A neural network is used to further generate corrective control on top of PD control targets generated by the motion matching system. This method produces impressive and realistic locomotion results and can be regarded as the state-of-the-art interactive physics-based method for locomotion. Similarly, the idea can also be extended to quadruped locomotion problems as shown in (Luo et al., 2020). In (Liu et al., 2016), the authors demonstrate interactive control of several complex acrobatic skills by designing a high-level control graph that schedules the reasonable motion fragments. In (Liu and Hodgins, 2017), deep Q-learning was combined into the fragment scheduler to further improve the scheduling performance. The authors also extend the proposed method to handle complex basketball controls in (Liu and Hodgins, 2018). However, as mentioned in their papers and later shown in our experiments, the above methods cannot generalize to unseen motions due to severe overfitting. The number of motions supported by these methods is also limited. Our method, on the other hand, can be applied to unseen motions and is capable of generalizing to a larger number of motions of drastically different styles.
Notable recent, concurrent work (Won et al., 2020) also considers learning large number of motions. In (Won et al., 2020), the authors first divide the motions into separate clusters, then use different networks for different clusters. UniCon instead focuses on transferability, zero-shot robustness, and exploits capacity by systematically rethinking the training framework and proposing new techniques.
Besides using the keyboard control to generate future states, interactiveness can also be established by directly encoding control information into the observation function, such as using a user-specified one-hot clip-selection vector (Peng et al., 2018a). In (Park et al., 2019), a recurrent neural network is used to predict task-driven future frames. A motion or intention vector can be interactively controlled by the user. The amount of motions supported in this encoding scheme is usually quite limited. In our experiments, we show that the performance drops catastrophically with increasing number of motions and the learnt controller exhibits virtually no ability to transfer to new motions.
Our proposed UniCon can also be used for physics-based interactive video control, which has not been widely studied in existing research. In real-time interactive video control, we directly teleport a person captured by camera into a virtual physics-based avatar.
Overview
The UniCon framework consists of two levels: a high level motion scheduler which takes as input interactive high-level control such as keyboard command or video and generates target motion frames, and a low-level motion executor which produces physics-based animation based on the target motion frames. The low-level motion executor presents the key innovation of our work, and can be used in conjunction with a variety of different high-level motion schedulers. Figure 2 shows the overview of our model.
We first introduce the low-level executor in section 4, and then separately discuss important design decisions that enable its robust training in a multi-task setting in section 5. In section 6, we describe how different high level motion schedulers can interact with our low-level executor, resulting in various interactive applications. Section 7 provides implementation details, including information about the humanoid model and physics engine used in our work. In section 8, we showcase our method both qualitatively and quantitatively, and provide benchmarks against existing work.
Low Level Motion Executor
The low-level motion executor assumes input in the form of target character states. These states come, for example, from mocap data, or can be produced by a high level motion scheduler. The low-level controller is implemented as a policy neural network that outputs a physics-based control signal, which drives the character to closely follow the target states.
We denote the state of the character as . To describe the state of a character, we consider the following information: the root position , the root rotation quaternion , the joint position , and the joint rotation quaternion where is the number of joints. We note that this information is redundant in that, for example, the joint position can be inferred from . We also consider the first order information from the four states, i.e., the root translation velocity , the root angular velocity , the joint translation velocity , and the joint angular velocity . The state thus takes the following form:
By combining both the local state information and the relative information between states, the observation vector takes the following form:
We avoid the use of absolute information in world space, which helps the agent to learn generalized features for control.
2. Controller
The choice of a controller is an important one. In existing RL-based animation algorithms, the proportional-derivative (PD) controller has been commonly used instead of the alternative torque-based controller that decides how much force to apply on each joint. We here argue in favour of torque-based controllers for our setting. In particular, in the experimental section we show empirically that, while PD controller usually has good sample efficiency and performance in training, it also tends to severely overfit. Although we do not present a mathematical explanation for this phenomena, a general rule of thumb we consider here is to avoid providing the controller with additional observation features, pre-processing or post-processing. In (Greydanus et al., 2017), the authors show that the Atari RL agents actually strongly focus on the scoreboard or the timer in their policies instead of paying attention to the actual game screen, suggesting the risk of RL agents using unexpected features from observation function. We also refer to (Peng and van de Panne, 2017) for comparison study.
Therefore, we use a torque based controller for UniCon, which we denote as , where is the concatenation of torques applied to each joint. We use a fully-connected neural network (Multilayer Perceptron, MLP), and denote the network weights as . Since our controller is tasked to master a much larger number of skills compared to (Peng et al., 2018a; Bergamin et al., 2019), we also use a much larger neural network. Specifically, in this paper we use three hidden layers of 1024 units, and ablate this choice in section 8.4.
3. Constrained Multiobjective Reward Optimization
Similar to (Peng et al., 2018a; Bergamin et al., 2019), we define the reward as a sum of several terms that measure the difference of the target state and the actual state on various statistics, i.e.,
Therefore we propose the following constrained optimization objective:
where is the tolerance coefficient ensuring that no reward term is dominated by other reward terms.
For a practical algorithm, it is impossible to directly optimize equation 6 using existing RL algorithms off-the-shelf. However, we can maintain a soft version of the constraint by enforcing early termination separately for each term, i.e., we terminate the episode if any reward term drops below the tolerance threshold. Empirically, we find that works well.
Training
We use proximal policy optimization (PPO) (Schulman et al., 2017) to optimize UniCon. PPO has been commonly used in recent physics-based research including (Peng et al., 2018a; Bergamin et al., 2019). We refer readers to the original paper for the detailed PPO algorithm (Schulman et al., 2017). We here use the following surrogate objective denoted as from PPOPPO is commonly referred to as a Policy Gradient (PG) method in current research. While PPO shares a lot of similarities with the original PG (Sutton et al., 2000), it is considered as a trust region method by its authors and optimizes a slightly different loss function (Schulman et al., 2017).:
where is the estimated advantage function, represents the policy weights which are fixed during the update, and is the weight for the Kullback-Leibler divergence penalty that discourages over-confident updates, as proposed in (Schulman et al., 2017). The iterative update in equation 7 is sample-based, and we describe how we sample the motion and initial state in section 5.1 and 5.2. During training, we simultaneously use 4096 workers to generate training samples. The number of samples per iteration per worker is 64. We also point out that other potentially more powerful RL algorithms can be applied as well, such as for example (Haarnoja et al., 2018; Fujimoto et al., 2018).
As later described in detail in section 6.1, the training motion dataset contains classes with an imbalanced number of samples. Random sampling during training can lead to a policy that is dominated by one specific skill, for example walking, which makes for around 35.4% of the dataset. It is difficult to control the balance of the dataset if, for example, we wish to extend UniCon to train on a dataset generated from vastly available Youtube videos in the future. We do not want to discard unbalanced data samples which may contain useful information that can improve generalization to a broader set of skills. Furthermore, the class labels for motions are usually a mixture of rough and fine-grained labels, posing additional challenges in maintaining the balance of training samples.
We thus propose to use a motion balancer. In particular, we first build a hierarchical tree structure for class labels, similar to what was done in ImageNet (Deng et al., 2009). For every motion, starting from class root node, we label their high-level class (major style such as walking), and then move downwards in the tree structure to the low-level classes (minor style such as forward or backward for walking). We do not limit the depth of this class hierarchy, such that for complicated motion we are able to generate more fine-grained class labels. For example, we can label a zombie-walking motion as root-walking-forward-specialStyle-zombie.
During training, we sample the motion by going down the hierarchical tree and uniformly sampling from all sub-node (child node) classes of the current node. We denote each node as and its sub-nodes or child-nodes as . The sampling process can be described as
We note that the sampling probability for each motion can be calculated off-line once for the entire dataset. We also point out that motion balancer can be extended with other adaptive sampling schemes in the future. Later in section 6.1, we explain the collection of hierarchical dataset by using human-labor and existing pose estimators. In the experimental section 8 we show that with the motion balancer our controller is able to obtain significantly better performance for a variety of tasks.
2. Reactive State Initialization Scheme
In DeepMimic, reference state initialization (RSI) was proposed to sample the state of a particular frame as the initial state. In our work, we further propose reactive state initialization scheme (RSIS), where the agent is initialized with a state of the frame time-steps away from the actual target frame it is supposed to track. RSIS also includes a much larger noise added to the velocity and translation of the initialized stateas mentioned in later experiment sections. With RSIS, the agent learns how to self-adjust and catch up with the target states. In UniCon, we show that recovering skills can be learnt automatically without the need to train a separate recovery network as in (Chentanez et al., 2018), or adding recovering motions in the dataset. When used with a motion stitching scheduler our agent naturally transitions between stitched motions that may have large discontinuities. The robustness of our agent is also greatly increased, demonstrating extreme recovering skills from perturbations as shown in the accompanied videos. Empirically, we choose a frame offset with 5-10 time-steps, which does not lead to early termination of episodes while encouraging the character to learn recovering skills.
3. Policy Variance Controller
where is the loss defined in equation 7, is the learning rate, and is the current PPO training iteration. We define a control iteration range from to , during which we linearly anneal the target average of from the hyperparameters and . is the operation where we linearly increase or decrease the log of each component of by the same amount, so that the average value matches the linear annealed target value. By doing this we preserve the learnt variance structure, which we show in section 8.4.
High Level Motion Scheduler
The high-level motion scheduler in our framework outputs target reference states from interactive control signal or from a motion dataset. We here discuss several different high level motion schedulers, and note that other existing schedulers can work in conjunction with our low-level executor. We denote the high-level motion scheduler as .
The high-level motion scheduler outputs states of future frames as
By unifying input observation to our low-level motion executor in UniCon, we establish a universal and transferable framework for different sources of control input. In experiments, we show that the low-level executor trained with a large-scale motion dataset can be used directly with a scheduler for specific applications without the need of retraining.
The scheduler terminates and reschedules a new motion only when the motion has reached the end, or the target and actual states have deviated significantly, as discussed in section 4.3.
We utilize the widely available CMU Graphics Lab Motion Capture Database (CMU Mocap dataset) (Hodgins, 2015). While motion dataset is commonly used in a physics-based system, the choice of train and test set splits has been generally overlooked. In our work, we carefully study the overfitting and transferability of the algorithms. We clean up the CMU Mocap dataset, and divide the data samples into training and testing sets based on several categories as presented in Table 1. We create several smaller datasets by grouping via style and tasks. For example, Casual dataset contains casual motions such as sweeping the floor, cleaning dishes, and the majority of the motions are standing motions. Animal dataset contains motions where the actors imitate animals such as cats and dogs. The All dataset contains all remaining motions after filtering out infeasible motions or motions that are highly dependent on external objects such as stairs.
We split the dataset into a training set that has 80% of the frames and a test set that has the remaining 20%. One issue with the CMU Mocap dataset is that the motions are extremely unbalanced. We find that for all of the motions available, 35.4% of them are walking motions, and 25.5% of all motions are walking-forward motions. On the other hand, there are lots of classes with very few motion clips. Unbalanced data is a very common (Lemaître et al., 2017), which may bias the neural network and detriment performance. To build the training and test sets, we split motion classes individually, starting with least frequent categories. If there is a class with only one motion, we always place it in the test set. We then allocate large classes such as walking-forward to fill the remaining training and testing sets. By doing this, we make sure that the test set has a good variability of motions. While the training set is unbalanced across classes, our motion balancer discussed in section 5.1 helps in training our motion executor effectively. For UniCon we manually label the motions. While we used human labeled dataset in our experiments, we point out crawled unlabeled videos can be used to generate future datasets. Fine-grained hierarchical labels can be automatically generated with video action classifiers, which also produce hierarchical labels as shown in (Shao et al., 2020; Feichtenhofer, 2020).
2. Video Stream Scheduler
Video can be used as an interactive control input for UniCon. Specifically, we consider a human who is captured by a camera and wants to teleport her/his motion onto a virtual avatar. We use a real-time pose estimator to estimate the 3D pose from video as shown in figure 4. We use (Iqbal et al., 2020) in our work but any other 3D pose estimation approach can be utilized, such as (Kanazawa et al., 2018; Güler et al., 2018). Without the loss of generality, we assume that the pose estimator is parameterized by a convolutional neural network with weights , and we represent the pose estimator as . We denote the video frame at time step as . The estimator generates the following prediction:
Note that the accuracy of pose estimation approaches is not perfect. Our experiments show that our low-level executor can handle very noisy estimated poses with reasonable accuracy.
3. Keyboard Driven Interactive Control Scheduler
For Unicon, instead of training a hierachical interactive scheduler from scratch, we use Phase-functioned neural networks (PFNN) (Holden et al., 2017) to process the keyboard commands and generate future states, as shown in figure 5. We note that the future states generated by PFNN is not physics-based. In PFNN, one can control the walking direction of the agent, and choose the walking style from walking, jogging, crouching, etc. We refer readers to (Holden et al., 2017) for details regarding PFNN. We write the target state generation of PFNN as
We also point out that besides PFNN, our algorithm can also be extended to use other keyframe based animation systems such as (Starke et al., 2019).
4. Motion Stitching Scheduler
Users can also interactively specify the motion for the character using a scheduler we call motion stitching scheduler. Motion stitching refers to directly stitching motions from database consecutively without worrying about proper transitions, as shown in figure 6. We maintain a stitched target motion buffer . Before the current buffer finishes, one can interactively add another motion with frames into the buffer, i. e.
We use spherical linear interpolation to add several transition target frames. At every time-step, motion stitching scheduler generates the next target state by popping a state like a FIFO buffer:
We note that the stitching scheduler can be regarded as the simplest version of motion graph, where motion is animated one-by-one. However, we show that our controller can animate a wide variety of highly difficult acrobatic motions on-the-fly in a physically plausible way, some of which are unseen in training. Our character automatically makes smooth transitions between motions even though we do not have transition skills recorded in the training set. Empirical results suggest that by plugging in a more powerful interactive motion graph system, our controller can generate an even more diverse set of physically plausible motion sequences.
Environment
Our experiments are performed with a reinforcement learning simulator similar to OpenAI Gym (Brockman et al., 2016). The simulator is powered by the GPU-accelerated Flex physics engine as the core backend for physics simulation. We refer the reader to (Liang et al., 2018) for more implementation details of the simulator and (Macklin et al., 2019) for the Flex physics engine solvers.
In our experiments, we use the CUDA-based Newton Preconditioned Conjugate Residual Method (PCR) solver for rigid-bodies provided by the Flex physics engine, with a simulation timestep of and a simulation substep value of . We set the number of iterations taken by the solver per simulation step as , and the number of inner loop iterations taken by the solver per simulation step as . For scene simulation parameters, we set gravity to downwards and use a value of for coefficient of static and dynamic friction.
2. Humanoid Model
We design our humanoid model with the basic topology of a rigid-body representation modeled after the human body. Our humanoid includes 20 rigid-bodies and 35 degrees-of-freedom. Each degree-of-freedom is assigned an effort factor within the range of 50 to 600, to simulate the difference in strength of joints in the human body. This effort factor is taken into consideration when torque control is applied. The height and mass of the humanoid model resembles a realistic proportion of the human body, at and respectively. The mass of each rigid-body in the humanoid model is proportionally distributed based on a rough estimate of the human body mass distribution. We use a fully symmetric humanoid model with respect to the left and right rigid-bodies and joints. We do not tune the parameters of the humanoid, and the learned policy generalizes across different models, which we refer to section 8.5.
Experiments
In this section we study the performance of the low-level motion executor (section 4), the backbone of our algorithm, on our motion dataset. More specifically, we numerically compare our low-level executor with existing algorithms by modifying them to be trainable on a large-scale motion dataset. Results show that our low-level executor performs better in almost every metric (section 8.2), and we analyze the key factors behind its success in section 8.4. In addition, we demonstrate that, unlike prior work, our low-level executor exhibits zero-shot robustness to environment perturbations not seen during training (section 8.5). These qualities of the low-level executor enable interactive applications outlined in section 9. The effectiveness of motion balancer and variance controller is shown with numerical ablation study. We also show visual ablation study on constrained multi-objective reward optimization and RSIS in the attached demo video.
We first introduce the baselines we use for comparisons here. We emphasize that constrained multi-objective reward optimization and the policy variance controller are crucial to training, without which, neither of our nor the baseline algorithms can be trained successfully. Therefore we apply these techniques equally to all baselines, and focus only on the controller design. We use the same network structure (1024x3) for every algorithm. We also tried the original structures specified in the papers, but the performance is worse than or similar to the ones with 1024x3.
PD-based Method: PD-based methods utilize a PD controller instead of a torque controller. During training, a neural policy network takes as input the current state and the target states from the dataset, and outputs the corrective offsets to the PD-control targets. This method is most similar to (Bergamin et al., 2019), except for the fact that (Bergamin et al., 2019) uses an online motion matching system to generate target states, whereas we directly use the target states from the dataset. We also use similar hyper-parameters from (Bergamin et al., 2019). We show another simple variant of removing the actual state from the observation function and performing open-loop control, which we name as Kinematic-State baseline.
(Chentanez et al., 2018): (Chentanez et al., 2018) is similar to PD-based methods, but the observation function of the policy network contains additional long-term information. A concatenation of future frames with time-step offsets are fed as observations into the tracking network, which outputs PD targets to control the humanoid. We do not include a separate recovery agent as in the original paper, since it can be applied to any algorithm being compared here. We later show that a separate recovery network is not necessary with the existence of a powerful low-level controller.
DeepMimic-Onehot and DeepMimic-Variable: The original DeepMimic algorithm supports multi-motion training through the use of a one-hot vector to encode motion information into the observation function, which we name as DeepMimic-Onehot. Since the number of motions during training can be quite large, we also include another variant called DeepMimic-Variable, where we feed the ID (normalized between 0 and 1 for all motions in the train and test sets) of the motion to the observation function.
2. Training Performance
In figure 7, we show the performance of our executor and the baseline methods on 6 datasets introduced in section 6.1. Our executor consistently obtains better sample efficiency and performance across the datasets. Results with the PD-based method and (Chentanez et al., 2018) show that while PD-controllers are very good at reproducing single or few motions, they perform worse than torque-based controllers on large motion datasets. We note that the PD-based method, which only uses a small number of future frames in the observation functions, actually performs better than (Chentanez et al., 2018). This could potentially be a cause of (Chentanez et al., 2018) feeding too much information into the network, some of which can contain information too far into the future that confuses the network and slows down the training process. We see poor performance from the Kinematic-State baseline, indicating that open-loop control is not enough for complex multi-motion animation tasks. We also see that both DeepMimic-Onehot and DeepMimic-Variable obtain poor performance, indicating their lack of model capacity.
In this section, we study the generalization of algorithms to unseen motions. In figure 8, we show the performance of our method and the baseline algorithms on the test set. The PD-based method and (Chentanez et al., 2018) reach their performance plateau very quickly on the test set, while training performance is still increasing. This over-fitting behaviour is most obvious in the two datasets shown in the figure. We suspect that this is due to the use of PD-controllers. Over-fitting in deep reinforcement learning is still an under-explored topic, and we do not explore further into this direction.
On the other hand, the test performance of Kinematic-State, DeepMimic-Onehot and DeepMimic-Variable shows almost no improvement throughout the training process, demonstrating serious over-fitting as well.
3. Transfer Learning with Fine-tuning
Since the test set was not seen during training, for academic purposes only, we can reuse the test-set to study how the controllers’ learnt features can be transferred to a dataset. In figure 10, dashed lines represent models trained from scratch, and solid lines represent models with pre-trained knowledge. We show that our low-level controller is highly transferable, suggesting potential use cases where we can first train a model on an enormous dataset, then fine-tune the model on smaller datasets for better specialization. The PD-based method and (Chentanez et al., 2018) also have slightly worse but reasonable transferability, while Kinematic-State, DeepMimic-Onehot and DeepMimic-Variable show no or even negative transferability.
4. Ablation Study on Low-level Executor
In figure 9, we study how the choice of target states and network structures affects performance of the network. We design two variants, the lookahead and stack. In lookahead- variant, we use the target state from the th future frame, instead of the next frame. In stack-, we include all target states in future frames into the observation function. As we can see from figure 9 (a), (b), the number of future frames from the high-level scheduler (i. e. ), does not visibly affect the performance curve of the low-level executor. We also observe that including information too far into the future can be detrimental to both training and testing performance. Furthermore, for applications such as real-time video streams, it is not possible to generate future frames for . Therefore, in section 9, we always set , which is essentially an inverse dynamics controller. In figure 9 (c), we also note that in contrast with single-motion or few motion trainings such as (Peng et al., 2018a; Bergamin et al., 2019), where narrow (layer width 256, 512 for example) neural networks with 2 hidden-layers are used, our algorithm requires much wider and deeper neural network structures. Empirically, we use 3 layers of dimension 1024, trained with 4096 agents, due to limitations of computational resources.
In figure 11, we show that the motion balancer and the variance controller are essential components to the success of training. Removing any of the two modules from our algorithm will cause a big performance gap.
5. Zero-Shot Robustness
Previously in (Peng et al., 2018a; Bergamin et al., 2019), results are shown where the agent can resist perturbation from projectiles, or retarget to models with different weight distributions. However, we argue that this leaves a gap between actual applications and what is demonstrated. More specifically, the retargeting is done by retraining a new model, and the robustness to projectiles, as we show later in the experiments, is still very limited.
In this project, we introduce zero-shot robustness, where the agent never sees the perturbation or retargeting information during training, and is asked to perform tasks under perturbations or using different humanoid models with varying masses. We argue that a robust controller with the ability to combat unseen perturbations and retargeting problems will have a much broader potential for real-life applications. More specifically, we have the following zero-shot robustness tasks:
Zero-shot Pertubation Robustness: In this case, the agents are trained without projectiles being thrown at them. During testing, the agents are required to perform the tasks under projectiles. We include experiments with a full range of projectile density and frequencies, and use the reward as the metric to numerically evaluate the performance.
Zero-shot Speed Robustness: Traditionally the motion data samples have a certain fixed speed. In DeepMimic, a phase variable is used to encode the time information of the motion. In this experiment, the agents are trained with motions at the original speed, but during testing, the agents are required to reproduce motions at different speeds.
Zero-shot Model Retargetting Robustness: In this case, instead of retraining on the new humanoid models as in DeepMimic, we ask the agent to reproduce the motion using a different humanoid model, which it has never seen during training. We use humanoid models that are 25% lighter and 25% heavier.
To numerically analyze the performance, we use the relative performance compared to the original performance without perturbation, projectiles or model-mismatch. As we can see from table 2, our method performs significantly better than DeepMimic. DeepMimic demonstrates very limited zero-shot robustness, while our method can still obtain almost performance in the majority of experiments. We did not include PD-based method and (Chentanez et al., 2018) in this comparison, as their performance on the evaluated cartwheel motion has a large gap compared to DeepMimic and our method, where they fail to reproduce the specific cartwheel motion when training with a large dataset. We also refer to the attached videos for more details.
Interactive Applications
In this section, we combine different high-level motion schedulers with our trained low-level executor (section 4), showcasing a number of interactive applications. We emphasize that all of the presented applications are real-time interactive and do not require any additional training or fine tuning of our low-level executor, which can be used on-the-fly in all of these settings. Since snapshots cannot fully demonstrate motion, we refer readers to the demo video.
Here we show the application using the high-level planner described in section 6.3. Figure 13 shows the snapshots of our agents controlled by the keyboard command (inset), with our agent shown in yellow and PFNN reference motion shown in white on the left. Note that we do not use inverse kinematics to force the agent’s feet to touch the ground in PFNN. This increases the engineering efficiency of PFNN module, and yet we show that the yellow agent generated by UniCon can still demonstrate realistic physics-based motions in a real-time fashion. We are almost able to perfectly follow the target states generated by PFNN.
Due to varying engines and humanoid models, we cannot fully reproduce the original PFNN with more motion gaits and uneven terrain. However, our universal framework’s efficiency and simplicity indicates that the motion quality and variety generated by our algorithm is only bottlenecked by the quality of the high-level scheduler used. Our algorithm has the potential to adapt to different schedulers with varying designs and implementations, such as the ones used in basketball and soccer video games.
2. Interactive Motion Stitching
In this section, we discuss the results where we randomly select a motion from a motion dataset and our algorithm will respond to that real time. In the cover image figure 1, the agent demonstrates the master of much more interactive complex skills compared to DeepMimic. In the attached video, UniCon also demonstrates strong emergent physics-based transition, where we show that, in contrast with some existing methods which design or interpolate realistic transitions, our system can smoothen sharp transition, where animation principles such as anticipation, ease-in & ease-out, are automatically satisfied with our low-level executor. We also note that the transition skills are not recorded in the dataset (not learnt from motion data), and they are mastered by UniCon by generalizing from other skills.
In figure 15, we further demonstrate the effectiveness and extreme transferability of our algorithm, by forcing our agent to react to motions it has never seen before. Our agent can still generate high-quality physics-based animation as shown in the figure 15.
It is worth mentioning that motion stitching can be viewed as the simplest motion graph system, which completely ignores generating smooth transitions between motions from the scheduler. However, the low-level executor is able to naturally generate smooth physics-based transitions on its own. Since our algorithm can demonstrate realistic motions using the simplest motion graph system, we believe it can also utilize better designed motion graph systems vastly available both in the research and engineering communities. The transferability skills on unseen motions demonstrated by our model also suggest potential use for motion systems with large motion variety.
3. Interactive Video Controlled Animation
In figure 14, we show how our algorithm can be used to teleport the motions captured from a remote host, to its physics-based avatars in the simulated environment real-time. Note that different from (Peng et al., 2018b), our system is real-time and does not require a high-fidelity pose-estimator, or hours of online re-training. In figure 14, indoor behaviors such as walking, turning, waving and jumping can be animated efficiently. Despite the mismatch in frame rate between the pose-estimator and our simulated environment, and the visually obvious pose estimation errors, our agent generates realistic real-time physics-based motions.
We expect the algorithm can be further improved, and combined with the vast available online datasets generated from video websites such as YouTube.
Conclusion and Discussion
In this paper, we proposed a universal neural controller for a variety of real-time interactive control applications. We closely study physics-based motion control on a large-scale dataset where novel techniques including constrained multi-objective reward optimization, motion balancing, and variance control are essential for the success of our framework. The controller we propose obtains much better robustness and generalization compared with existing research, where training and testing are generally performed on the same motion distribution. Once trained, our method does not require further online retraining and can be applied on-the-fly to various applications, such as real-time interactive control from keyboard, videos and motion stitching.
We also identify limitations of our framework and potential research topics for the future. One such topic is exploring how we can further improve the capacity of the neural controller, so that it can master even more skills without worrying about asymptotic performance drop for each motion. Most of the applications we demonstrate are research driven. The high-level schedulers used still have areas for improvement compared with those used in high-quality video games today. It remains to be explored how well our framework can be combined with an AAA video-game motion system. We also did not explore the ability of our method on varying terrain types and physical interactions with objects in the scene. All of these topics present an exciting avenue for future work.