D&D: Learning Human Dynamics from Dynamic Camera
Jiefeng Li, Siyuan Bian, Chao Xu, Gang Liu, Gang Yu, Cewu Lu
Introduction
Recovering 3D human pose and shape from a monocular image is a challenging problem. It has a wide range of applications in activity recognition , character animation, and human-robot interaction. Despite the recent progress, estimating 3D structure from the 2D observation is still an ill-posed and challenging task due to the inherent ambiguity.
A number of works turn to temporal input to incorporate body motion priors. Most state-of-the-art methods are only based on kinematics modeling, i.e., body motion modeling with body part rotations and joint positions. Kinematics modeling directly captures the geometric information of the 3D human body, which is easy to learn by neural networks. However, methods that entirely rely on kinematics information are prone to physical artifacts, such as motion jitter, abnormal root actuation, and implausible body leaning.
Recent works have started modeling human motion dynamics to improve the physical plausibility of the estimated motion. Dynamics modeling considers physical forces such as contact force and joint torque to control human motion. These physical properties can help analyze the body motion and understand human-scene interaction. Compared to widely adopted kinematics, dynamics gains less attention in 3D human pose estimation. The reason is that there are lots of limitations in current dynamics methods. For example, existing methods fail in daily scenes with dynamic camera movements (e.g., the 3DPW dataset ) since they require a static camera view, known ground plane and gravity vector for dynamics modeling. Besides, they are hard to deploy for real-time applications due to the need for highly-complex offline optimization or simulation with physics engines.
In this work, we propose a novel framework, D&D, a 3D human pose estimation approach with learned Human Dynamics from Dynamic Camera. Unlike previous methods that build the dynamics equation in the world frame, we re-devise the dynamics equations in the non-inertial camera frame. Specifically, when the camera is moving, we introduce inertial forces in the dynamics equation to relate physical forces to local pose accelerations. We develop dynamics networks that directly estimate physical properties (forces and contact states). Then we can use the physical properties to compute the pose accelerations and obtain final human motion based on the accelerations. To train the dynamics network with only a limited amount of contact annotations, we propose probabilistic contact torque (PCT) for differentiable contact torque estimation. Concretely, we use a neural network to predict contact probabilities and conduct differentiable sampling to draw contact states from the predicted probabilities. Then we use the sampled contact states to compute the torque of the ground reaction forces and control the human motion. In this way, the contact classifier can be weakly supervised by minimizing the difference between the generated human motion and the ground-truth motion. To further improve the smoothness of the estimated motion, we propose a novel control mechanism called attentive PD controller. The output of the conventional PD controller is proportional to the distance of the current pose state from the target state, which is sensitive to the unstable and jittery target. Instead, our attentive PD controller allows accurate control by globally adjusting the target state and is robust to the jittery target.
We benchmark D&D on the 3DPW dataset captured with moving cameras and the Human3.6M dataset captured with static cameras. D&D is compared against both state-of-the-art kinematics-based and dynamics-based methods and obtains state-of-the-art performance.
The contributions of this paper can be summarized as follows:
We present the idea of inertial force control (IFC) to perform dynamics modeling for 3D human pose estimation from a dynamic camera view.
We propose probabilistic contact torque (PCT) that leverages large-scale motion datasets without contact annotations for weakly-supervised training.
Our proposed attentive PD controller enables smooth and accurate character control against jittery target motion.
Our approach outperforms previous state-of-the-art kinematics-based and dynamics-based methods. It is fully differentiable and runs without offline optimization or simulation in physics engines.
Related Work
Numerous prior works estimate 3D human poses by locating the 3D joint positions . Although these methods obtain impressive performance, they cannot provide physiological and physical constraints of the human pose. Many works adopt parametric statistical human body models to improve physiological plausibility since they provide a well-defined human body structure. Optimization-based approaches automatically fit the SMPL body model to 2D observations, e.g., 2D keypoints and silhouettes. Alternatively, learning-based approaches use a deep neural network to regress the pose and shape parameters directly . Several works combine the optimization-based and learning-based methods to produce pseudo supervision or conduct test-time optimization.
For better temporal consistency, recent works have started to exploit temporal context . Kocabas et al. propose an adversarial framework to leverage motion prior from large-scale motion datasets . Sun et al. model temporal information with a bilinear Transformer. Rempe et al. propose a VAE model to learn motion prior and optimize ground contacts. All the aforementioned methods disregard human dynamics. Although they achieve high accuracy on pose metrics, e.g., Procrustes Aligned MPJPE, the resulting motions are often physically implausible with pronounced physical artifacts such as improper balance and inaccurate body leaning.
0.2 Dynamics-based 3D Human Pose Estimation.
To reduce physical artifacts, a number of works leverage the laws of physics to estimate human motion . Some of them are optimization-based approaches . They use trajectory optimization to obtain the physical forces that induce human motion. Shimada et al. consider a complete human motion dynamics equation for optimization and obtain motion with fewer artifacts. Dabral et al. propose a joint 3D human-object optimization framework for human motion capture and object trajectory estimation. Recent works have started to use regression-based methods to estimate human motion dynamics. Shimada et al. propose a fully-differentiable framework for 3D human motion capture with physics constraints. All previous approaches require a static camera, restricting their applications in real-world scenarios.
On the other hand, deep reinforcement learning and motion imitation are widely used for 3D human motion estimation . These works rely on physics engines to learn the control policies. Peng et al. propose a control policy that allows a simulated character to mimic realistic motion capture data. Yuan et al. present a joint kinematics-dynamics reinforcement learning framework that learns motion policy to reconstruct 3D human motion. Luo et al. propose a dynamics-regulated training procedure for egocentric pose estimation. The work of Yu et al. is most related to us. They propose a policy learning algorithm with a scene fitting process to reconstruct 3D human motion from a dynamic camera. Training their RL model and the fitting process is time-consuming. It takes hours to obtain the human motion of one video clip. Unlike previous methods, our regression-based approach is fully differentiable and does not rely on physics engines and offline fitting. It predicts accurate and physically plausible 3D human motion for in-the-wild scenes with dynamic camera movements.
Method
The overall framework of the proposed D&D (Learning Human Dynamics from Dynamic Camera) is summarized in Fig. 1. The input to D&D is a video with frames. Each frame is fed into the kinematics backbone network to estimate the initial human motion in the local camera frame. The dynamics networks take as input the initial local motion and estimate physical properties (forces and contact states). Then we apply the forward dynamics modules to compute the pose and trajectory accelerations from the estimated physical properties. Finally, we use accelerations to obtain 3D human motion with physical constraints iteratively.
In this section, before introducing our solution, we first review the formulation of current dynamics-based methods in §3.1. In §3.2, we present the formulation of Inertial Force Control (IFC) that introduces inertial forces to explain the human motion in the dynamic camera view. Then we elaborate on the pipeline of D&D: i) learning physical properties with neural networks in §3.3, ii) analytically computing accelerations with forward dynamics in §3.4, iii) obtaining final pose with the constrained update in §3.5. The objective function of training the entire framework is further detailed in §3.6.
2 Inertial Force Control
In this work, to facilitate in-the-wild 3D human pose estimation with physics constraints, we reformulate the dynamics equation to impose the laws of physics in the dynamic-view video. When the camera is moving, the local frame is an inertial frame of reference. In order to satisfy the force equilibrium, we introduce the inertial force in the dynamics system:
The inertial force control (IFC) establishes the relation between the physical properties and the pose acceleration in the local frame. The pose acceleration can be subsequently used to calculate the final motion. In this way, we can estimate physically plausible human motion from forces to accelerations to poses. The generated motion is smooth and natural. Besides, it provides extra physical information to understand human-scene interaction for high-level activity understanding tasks.
The concept of residual force is widely adopted in previous works to explain the direct root actuation in the global static frame. Theoretically, we can adopt a residual term to explain the inertia in the local camera frame implicitly. However, we found explicit inertia modeling obtains better estimation results than implicit modeling with a residual term. Detailed comparisons are provided in §4.4.1.
3 Learning Physical Properties
In this subsection, we elaborate on the neural networks for physical properties estimation. We first use a kinematics backbone to extract the initial motion . The initial motion is then fed to a dynamics network (DyNet) with probabilistic contact torque for contact, external force, and inertial force estimation and the attentive PD controller for internal joint torque estimation.
The root motion of the human character is dependent on external forces and inertial forces. To explain root motion, we propose DyNet that directly regresses the related physical properties, including the ground reaction forces , the gravity , the direct root actuation , the contact probabilities , the linear camera acceleration , and the angular camera velocity . The detailed network structure of DyNet is provided in the supplementary material.
The inertial force can be calculated following Eqn. 3 with the estimated and . The gravity torque can be calculated as:
When considering gravity, bodyweight will affect human motion. In this paper, we let the shape parameters control the body weight. We assume the standard weight is kg when , and there is a linear correlation between the body weight and the bone length. We obtain the corresponding bodyweight based on the bone-length ratio of to .
Probabilistic Contact Torque: For the resultant torque of the ground reaction forces, previous methods compute it with the discrete contact states of joints:
where for contact and for non-contact. Note that the output probabilities are continuous. We need to discretize with a threshold of 0.5 to obtain . However, the discretization process is not differentiable. Thus the supervision signals for the contact classifier only come from a limited amount of data with contact annotations.
To leverage the large-scale motion dataset without contact annotations, we propose probabilistic contact torque (PCT) for weakly-supervised learning. During training, PCT conducts differentiable sampling to draw a sample that follows the predicted contact probabilities and computes the corresponding ground reaction torques:
where are i.i.d samples drawn from the Gumbel distribution. When conducting forward dynamics, we use the sampled torque instead of the torque from the discrete contact states . To generate accurate motion, DyNet is encouraged to predict higher probabilities for the correct contact states so that PCT can sample the correct states as much as possible. Since PCT is differentiable, the supervision signals for the physical force and contact can be provided by minimizing the motion error. More details of differentiable sampling are provided in the supplementary material.
3.2 Internal Joint Torque Estimation.
Another key process to generate human motions is internal joint torque estimation. PD controller is widely adopted for physics-based human motion control . It controls the motion by outputting the joint torque in proportion to the difference between the current state and the target state. However, the target pose states estimated by the kinematics backbone are noisy and contain physical artifacts. Previous works adjust the gain parameters dynamically for smooth motion control. However, we find that this local adjustment is still challenging for the model and the output motion is still vulnerable to the jittery and incorrect input motion.
Attentive PD Controller: To address this problem, we propose the attentive PD controller, a method that allows global adjustment of the target pose states. The attentive PD controller is fed with initial motion and dynamically predicts the proportional parameters , derivative parameters , offset torques , and attention weights . The attention weights denotes how the initial motion contributes to the target pose state at the time step and . We first compute the attentive target pose state as:
where is the initial kinematic pose at the time step . Then the internal joint torque at the time step can be computed following the PD controller rule with the compensation term :
where denotes Hadamard matrix product and represents the sum of centripetal and Coriolis forces at the time step . This attention mechanism allows the PD controller to leverage the temporal information to refine the target state and obtain a smooth motion. Details of the network structure are provided in the supplementary material.
4 Forward Dynamics
To compute the accelerations analytically from physical properties, we build two forward dynamics modules: inertial forward dynamics for the local pose acceleration and trajectory forward dynamics for the global trajectory acceleration.
Prior works adopt a proxy model to simulate human motion in physics engines or simplify the optimization process. In this work, to seamlessly cooperate with the kinematics-based backbone, we directly build the dynamics equation for the SMPL model . The pose acceleration can be derived by rewriting Eqn. 2 with PCT:
4.2 Trajectory Forward Dynamics.
To train DyNet without ground-truth force annotations, we leverage a key observation: the gravity and ground reaction forces should explain the global root trajectory. We devise a trajectory forward dynamics module that controls the global root motion with external forces. It plays a central role in the success of weakly supervised learning.
Let denote the root translation in the world frame. The dynamics equation can be written as:
where is the mass of the root joint, denotes the camera orientation computed from the estimated angular velocity , denotes the direct root actuation, and and denote the first three entries of and , respectively.
5 Constrained Update
After obtaining the pose and trajectory accelerations via forward dynamics modules, we can control the human motion and global trajectory by discrete simulation. Given the frame rate of the input video, we can obtain the kinematic 3D pose using the finite differences:
Similarly, we can obtain the global root trajectory:
In practice, since we predict the local and global motions simultaneously, we can impose contact constraints to prevent foot sliding. Therefore, instead of using Eqn. 12 and 14 to update and directly, we first refine the velocities and with contact constraints. For joints in contact with the ground at the time step , we expect they have zero velocity in the world frame. The velocities of non-contact joints should stay close to the original velocities computed from the accelerations. We adopt the differentiable optimization layer following the formulation of Agrawal et al. . This custom layer can obtain the solution to the optimization problem and supports backward propagation. However, the optimization problem with zero velocity constraints does not satisfy the DPP rules (Disciplined Parametrized Programming), which means that the custom layer cannot be directly applied. Here, we use soft velocity constraints to follow the DPP rules:
6 Network Training
The 3D loss includes the joint error and the pose error:
where denotes the projection function. The loss is added for the supervision of the root translation and provides weak supervision signals for external force and contact estimation:
The contact loss is added for the data with contact annotations:
The regularization loss is defined as:
where the first term minimizes the direct root actuation, and the second term minimizes the entropy of the contact probability to encourage confident contact predictions.
Experiment
We perform experiments on two large-scale human motion datasets. The first dataset is 3DPW . 3DPW is a challenging outdoor benchmark for 3D human motion estimation. It contains video sequences obtained from a hand-held moving camera. The second dataset we use is Human3.6M . Human3.6M is an indoor benchmark for 3D human motion estimation. It includes subjects, and the videos are captured at 50Hz. Following previous works , we use subjects (S1, S5, S6, S7, S8) for training and subjects (S9, S11) for evaluation. The videos are subsampled to 25Hz for both training and testing. We further use the AMASS dataset to obtain annotations of foot contact and root translation for training.
2 Implementation Details
We adopt HybrIK as the kinematics backbone to provide the initial motion. The original HybrIK network only predicts 2.5D keypoints and requires a separate RootNet to obtain the final 3D pose in the camera frame. Here, for integrity and simplicity, we implement an extended version of HybrIK as our kinematics backbone that can directly predict the 3D pose in the camera frame by estimating the camera parameters. The network structures are detailed in the supplementary material. The learning rate is set to at first and reduced by a factor of at the th and th epochs. We use the Adam solver and train for epochs with a mini-batch size of . Implementation is in PyTorch. During training on the Human3.6M dataset, we simulate a moving camera by cropping the input video with bounding boxes.
3 Comparison to state-of-the-art methods
We first compare D&D against state-of-the-art methods on 3DPW, an in-the-wild dataset captured with the hand-held moving camera. Since previous dynamics-based methods are not applicable in the moving camera, prior arts on the 3DPW dataset are all kinematics-based. Mean per joint position error (MPJPE) and Procrustes-aligned mean per joint position error (PA-MPJPE) are reported to assess the 3D pose accuracy. The acceleration error (ACCEL) is reported to assess the motion smoothness. We also report Per Vertex Error (PVE) to evaluate the entire estimated body mesh.
Tab. 1 summarizes the quantitative results. We can observe that D&D outperforms the most accurate kinematics-based methods, HybrIK and MAED, by and mm on MPJPE, respectively. Besides, D&D improves the motion smoothness significantly by % and % relative improvement on ACCEL, respectively. It shows that D&D retains the benefits of the accurate pose in kinematics modeling and physically plausible motion in dynamics modeling.
3.2 Results on Static Camera
To compare D&D with previous dynamics-based methods, we evaluate D&D on the Human3.6M dataset. Following the previous method , we further report two physics-based metrics, foot sliding (FS) and ground penetration (GP), to measure the physical plausibility. To assess the effectiveness of IFC, we simulate a moving camera by cropping the input video with bounding boxes, i.e., the input to D&D is the video from a moving camera. Tab. 2 shows the quantitative comparison against kinematics-based and dynamics-based methods. D&D outperforms previous kinematics-based and dynamics-based methods in pose accuracy. For physics-based metrics (ACCEL, FS, and GP), D&D shows comparable performance to previous methods that require physics simulation engines.
We further follow GLAMR to evaluate the global MPJPE (G-MPJPE) and global PVE (G-PVE) on the Human3.6M dataset with the simulated moving camera. The root translation is aligned with the GT at the first frame of the video sequence. D&D obtains 785.1mm G-MPJPE and 793.3mm G-PVE. More comparisons are reported in the supplementary material.
4 Ablation Study
In this experiment, we compare the proposed inertial force control (IFC) with residual force control (RFC). To control the human motion with RFC in the local camera frame, we directly estimate the residual force instead of the linear acceleration and angular velocity. Quantitative results are reported in Tab. 3. It shows that explicit modeling of the inertial components can better explain the body movement than implicit modeling with residual force. IFC performs more accurate pose control and significantly reduces the motion jitters, showing a % relative improvement of ACCEL on 3DPW.
4.2 Effectiveness of PCT.
To study the effectiveness of the probabilistic contact torque, we remove PCT in the baseline model. When training the baseline, the output contact probabilities are discretized to or with the threshold of and we compute the discrete contact torque instead of the probabilistic contact torque. Quantitative results in Tab. 3 show that PCT is indispensable to have smooth and accurate 3D human motion.
4.3 Effectiveness of Attentive PD Controller.
To further validate the effectiveness of the attentive mechanism, we report the results of the baseline model without the attentive PD controller. In this baseline, we adopt the meta-PD controller that dynamically predicts the gain parameters based on the state of the character, which only allows local adjustment. Tab. 3 summarizes the quantitative results. The attentive PD controller contributes to a more smooth motion control as indicated by a smaller acceleration error.
5 Qualitative Results
In Fig. 3, we plot the contact forces estimated by D&D of the walking motion from the Human3.6M test set. Note that our approach does not require any ground-truth force annotations for training. The estimated forces fall into a reasonable force range for walking motions . We also provide qualitative comparisons in Fig. 2. It shows that D&D can estimate physically plausible motions with accurate foot-ground contacts and no ground penetration.
Conclusion
In this paper, we propose D&D, a physics-aware framework for 3D human motion capture with dynamic camera movements. To impose the laws of physics in the moving camera, we introduce inertial force control that explains the 3D human motion by taking the inertial forces into consideration. We further develop the probabilistic contact torque for weakly-supervised training and the attentive PD controller for smooth and accurate motion control. We demonstrate the effectiveness of our approach on standard 3D human pose datasets. D&D outperforms state-of-the-art kinematics-based and dynamics-based methods. Besides, it is entirely neural-based and runs without offline optimization or physics simulators. We hope D&D can serve as a solid baseline and provide a new perspective for dynamics modeling in 3D human motion capture.
This work was supported by the National Key R&D Program of China (No. 2021ZD0110700), Shanghai Municipal Science and Technology Major Project (2021SHZDZX0102), Shanghai Qi Zhi Institute, SHEITC (2018-RGZN-02046) and Tencent GY-Lab.
References
Appendix
In the supplemental document, we provide:
A more detailed explanation of differentiable sampling.
Architectures of the kinematics backbone, DyNet and attentive PD controller.
Experiments that compare D&D with more baselines.
Physical-based results on the 3DPW dataset.
Global trajectory results on the Human3.6M dataset.
Appendix 0.A Derivations of the Dynamics Quantities
Following Featherstone et al. , the inertia matrix M is derived as:
where is the parent index of the -th joint and represents the relative angular Jacobian matrix that can be computed from . The linear Jacobian matrix of the -th joint can be computed based on :
where is the skew-symmetric matrix of the body-part vector .
A.0.2 Angular Jacobian Matrix.
The angular velocity of the -th body joint can be computed recursively:
where is the parent index of the -th joint and denotes the relative angular velocity. Notice that the angular Jacobian matrix of the -th body joint should satisfy:
Therefore, we can build the recursive equation for by replacing with in Eqn. 25:
To compute , we now need to compute in each step. Denote as the relative rotation of the -th joint in Euler angles. Then is defined as:
For the root joint that has no parent, we have:
A.0.3 Linear Jacobian Matrix.
The linear velocity of the -th body joint can be computed based on the angular velocity:
where denotes the vector of the -th body part. Notice that the linear Jacobian matrix of the -th body joint should satisfy:
Therefore, we can compute by replacing with and with in Eqn. 32:
For the root joint that has no parent, we have:
where denotes the identity matrix.
A.0.4 Inertia Tensor.
Let denote the inertia tensor of the -th body joint under the rest pose that can be pre-computed. The inertia tensor of the -th body joint under the pose can be computed as:
where denotes the rotation matrix of the -th body joint.
Appendix 0.B Differentiable Sampling
The Gumbel-Max trick provides a simple way to draw samples from a categorical distribution. The contact distribution of the -th joint follows the Bernoulli distribution:
To draw sample with class probability , we can conduct:
where . Then the corresponding ground reaction torque can be computed as:
B.0.2 Differentiable Process.
However, Gumbal-Max is not differentiable. Therefore, we adopt Gumbel-Softmax to conduct differentiable sampling from the predicted distribution. Gumbel-Softmax is a continuous and differentiable approximation to Gumbel-Max by replacing argmax with the softmax function:
Therefore, the corresponding ground reaction torque can be computed as:
Alternatively, we can compute the expectation of the ground reaction torque, which is also differentiable:
However, using the expectation to generate motion cannot encourage well-calibrated probabilities , i.e., DyNet is not encouraged to generate high probabilities for the correct contact states. Consequently, the contact forces might be incorrect since the model is trained without direct supervision. We plot the contact forces of the walking motion using the model trained with the torque expectation in Fig. 4(b). It shows that using the torque expectation makes the contact forces lie in an unreasonable range.
Appendix 0.C Network Architecture
The detailed network architecture of the kinematics backbone is illustrated in Fig. 5. We implement an extended version of HybrIK as the backbone network. The original HybrIK model first predicts 2.5D joints of the body skeleton. Then the RootNet is adopted to predict the distance of the root joint to the camera plane and obtain the 3D joints in the camera frame via back-projection. In our implementation, we design an integrated model that directly estimates the 3D joint positions. Specifically, we add a fully-connected layer to regress the camera parameters . Therefore, we can obtain the 3D joint positions as well as the initial motion within a single model.
C.2 DyNet
The detailed network architecture of the DyNet is illustrated in Fig. 6. For all 1D convolutional layers, we use the kernel of size and channel of size . The output layer consists of a bilinear GRU layer with the hidden dimension of and a fully-connected layer to predict the physical properties.
C.3 Attentive PD Controller
The detailed network architecture of the attentive PD controller is illustrated in Fig. 7. The kernel size and channel number are set to and , respectively. To predict the attention weights that satisfy , we use a softmax layer for normalization.
Appendix 0.D Details for Contact Annotations
To provide supervision signals for contact states, we generate contact annotations on the AMASS dataset . Given a video sequence, we first retrieve the toe joints of the body in each frame and then fit the ground plane using the toe positions. Finally, the joints within cm of the ground are labeled as in contact. Following previous works , we use contact joints, i.e., toes and heels. In daily scenes, the human body might contact the environment with other joints, such as hips and hands. The current model will compensate these undefined contact forces with the residual force, and it is easy for our model to extend to more contact joints.
Appendix 0.E Comparason with Baselines.
To further assess the effectiveness of D&D, we conduct experiments to compare D&D with two baselines: the acceleration network (AccNet) and the velocity network (VelNet). Instead of predicting physical properties like forces and contacts, AccNet and VelNet directly predict the pose acceleration and velocity, respectively. The acceleration and velocity are used to control the human motion. The architectures of AccNet and VelNet are similar to DyNet, with an output layer that predicts the pose acceleration and velocity. Other training settings are the same as the proposed D&D. As shown in Tab. 4, D&D obtains superior performance to AccNet and VelNet. It demonstrates that the improvement of D&D comes from the explicit modeling of the physical properties, not controlling human motion with acceleration or velocity.
Appendix 0.F Physical-based Results on the 3DPW Dataset.
We report two physical-based metrics, foot sliding (FS) and ground penetration (GP), to measure the physical plausibility on the 3DPW dataset. We remove the video sequences that contain stairs since it is inaccurate to measure the ground penetration in these cases. Quantitative results are provided in Tab. 5. It shows that D&D can predict physically plausible motion in daily scenes with dynamic camera movements.
Appendix 0.G Global Trajectory Results on the Human3.6M Dataset.
We compare the predicted global trajectory using the simulated moving camera with the results using the static camera on the Human3.6M dataset. Quantitative results are reported in Tab. 6. When testing on the static camera, we directly use the predicted trajectory from HybrIK in camera coordinates as the global trajectory. We then follow the standard evaluations for open-loop reconstruction (e.g., SLAM and GLAMR ) to remove the effect of the accumulative error. G-MPJPE and G-PVE are computed using a sliding window (10 seconds) and align the root translation with the GT at the start of each window. We can see that using static camera obtains better results than the closed-loop results of D&D with the moving camera. When we follow the open-loop protocol to eliminate the accumulative errors, D&D obtains better results than using the static camera.
Appendix 0.H Computation Complexity.
The whole system is run online by using a sliding window with a length of 16 frames and a stride of 16 frames. The system takes 1349ms for each window (84.3ms for each frame).
Appendix 0.I Pseudocode
The pseudocode of the proposed D&D is given in Alg. 1.
Appendix 0.J Qualitative Results
Additional qualitative results are shown in Fig. 8, Fig. 9 (the Human3.6M dataset), Fig. 10, Fig. 11, Fig. 12, Fig. 13 (the 3DPW dataset), and Fig. 14 (the Internet videos).