Locomotion-Action-Manipulation: Synthesizing Human-Scene Interactions in Complex 3D Environments
Jiye Lee, Hanbyul Joo
Introduction
Synthesizing interactions within real-life 3D environments has been a challenging research problem due to its complexity and diversity. The spatial constraint arising from real-life 3D environments where many objects are cluttered makes motion synthesis highly constrained and complex. Furthermore, the nearly indefinite diversity of possible spatial arrangements of the 3D environment and human interaction behaviors makes generalization in synthesis difficult.
Due to the wide range of technical challenges involved in human-scene interactions, previous approaches have focused on sub-problems, such as (1) modeling static poses or (2) human object interactions with a single target object or interaction type . More recent methods extend to synthesizing dynamic interaction motions in real-world 3D scenes, where they use “scene-paired” motion datasets in which motion is simultaneously captured with the surrounding 3D environment. As such paired dataset is rare and difficult to scale up, the performance of these methods is fundamentally limited in fully covering the complexity and diversity of human interaction in real-world 3D scenes.
In this paper, we present LAMA, Locomotion-Action-MAnipulation, to synthesize natural and plausible long-term human motions in complex indoor environments. The key motivation of LAMA is to build a unified framework covering a series of everyday motions within real-world 3D scenes: locomotion through cluttered areas, interaction with the scene, and manipulation of objects. Unlike previous approaches that use a “scene-paired” motion dataset for supervision, we formulate it as a test-time optimization by utilizing only human motion capture data. Exploiting reinforcement learning (RL) as a tool for optimization, we present an RL-based framework coupled with a motion matching algorithm to synthesize locomotion and scene interaction seamlessly while adapting to complex 3D scenes with collision avoidance handling. The object manipulation in our framework is performed via a motion editing approach on top, by learning an autoencoder-based motion manifold space . As a test-time optimization framework, LAMA is applicable to any 3D scene scenarios (e.g., public datasets or any newly scanned scenes). Through extensive quantitative and qualitative evaluations against existing methods, we demonstrate that our method outperforms in various challenging scenarios.
Our contributions are summarized as follows: (1) The first method to generate realistic long-term motions combined with locomotion, scene interaction, and manipulation in complex 3D scenes without “paired” datasets; (2) A novel test-time optimization framework requiring human motion capture data only by incorporating a reinforcement learning framework coupled with motion matching, equipped with well-designed state and rewards for collision avoidance and scene interactions; (3) the state-of-the-art motion synthesis quality with longer duration (near 10 sec); (4) A newly captured and polished motion capture dataset including locomotion and action (e.g., sitting) suitable for motion matching.
Related Work
Generating Human-Scene Interactions. Generating natural human motion has been a widely researched topic in the computer vision community. Early methods focus on synthesizing or predicting human movements by exploiting neural networks . However, these approaches primarily address the synthesis of human motion itself, without taking into account the surrounding 3D environments. Recent approaches begin to tackle modeling and synthesizing human interactions within 3D scenes, or with objects. Many focus on statically posing humans within the given 3D environment , by generating human scene interaction poses from various types of input including object semantics , images , and text descriptions .
Recently, there have been approaches to synthesize dynamic human-object interactions (e.g., sitting on chairs, carrying boxes). Starke et al. introduce an autoregressive learning framework with object geometry based environmental encodings to synthesize human-object interactions. Although the encoding includes information on multiple objects within a scene, as demonstrated in , no explicit module for navigating through cluttered 3D scenes exists in . Later work extends this by synthesizing motions conditioned with variations of objects and contact points. Other approaches focus on generating natural hand movements for manipulation, which is extended by including full body motions . Physics-based character control to synthesize human-object interactions has been also explored in . Although these methods cover a variety of human-object interactions, most of them focus on a specific interaction type or the relationship between the human and the target object without long-term navigation in cluttered 3D scenes.
More recent approaches include generating natural human scene interactions in cluttered 3D scenes , closely related to ours. These methods are trained using human motion datasets paired with 3D scenes, which require both ground truth motion and simultaneously captured 3D scenes for supervision. Due to difficulties in acquiring such data, some methods exploit synthetic datasets , data fitted from depth videos , or motion snapshots with short duration (1-3 sec) . In previous approaches , navigation in cluttered environments is often performed by a separate module via path planning (e.g., algorithm) by approximating the volume of a human as a cylinder. These path planning based methods approximate the spatial information of the scene and the body and therefore have limitations under highly complex conditions.
Motion Synthesis and Editing. Synthesizing natural human motions by leveraging motion capture data has also been a long-researched topic in computer graphics. Some approaches construct motion graphs, where plausible transitions are inserted as edges, and motion synthesis is done by traversing through the graph. Similar approaches connect motion patches to synthesize interactions in a virtual environment or multi-person interactions. Due to its versatility and simplicity, variations have been made to the graph-based approach, such as motion grammar which enforces traversing rules in the motion graph. Motion matching can also be understood as a special case of motion graph traversal, where the plausible transitions are not precomputed but searched during runtime. Recent advances in deep learning allow to leverage motion capture data for motion manifold learning . Autoregressive approaches based on variational autoencoders (VAE) and recurrent neural networks are also used to forecast future motions based on past frames. These frameworks are generalized to synthesize a diverse set of motions including locomotion on terrains mazes , action-specified motions , and interaction-involved sports . Neural network-based methods are also reported to be successful in various motion editing tasks such as skeleton retargeting , style transfer , and in-betweening .
Reinforcement learning (RL) has also been successful in combination with both data-driven and physics-based approaches for synthesizing human motions. Combined with data-driven approaches, RL serves as a control module that generates corresponding motions to a given user input by traversing motion graphs , latent space , and precomputed transition tables . Deep reinforcement learning (DRL) has been widely used as well to synthesize physically plausible movements with a diverse set of motor skills . The key idea of these methods comes from imitation learning, where the control policy in DRL is optimized to actuate the character based on the character’s physical state, to meet the goal of tracking the given reference motion in a physically simulated environment.
Method
Our system, dubbed as , outputs a sequence of human poses by taking the 3D scene , desired interaction cues , and initial state as inputs:
LAMA is designed via a three-level system composed of action controller , motion synthesizer , followed by a manifold-based motion editor . The locomotion and action parts are seamlessly performed via the action controller and synthesizer . The essential idea in our design is to combine the RL framework with motion matching to synthesize realistic human motions while fulfilling the desired scene interaction tasks. By taking 3D scene , action cue , and initial state as input, the action controller makes the use of RL as a way of test-time optimization to synthesize corresponding motion. A control policy is optimized We use the term “optimized” rather than “learned” for the policy since we perform a test-time optimization. to sample an action at time , , where indicates the plausible next action containing predicted action types and short-term future forecasting. is the state cue to represent the current status of the human character including its body posture, surrounding scene occupancy, and the targeting action cue. Intuitively, action controller is optimized to generate plausible next action by considering the current character-scene state . The generated action signal from the action controller is provided as input to the motion synthesizer , which then determines the posture at the next time step , i. e., . The character’s next state can be computed again from , which is subsequently taken by the action controller as an input for the next time frame.
2 Scene-Aware Action Controller
Unlike previous approaches that use path planning for navigation and a learning-based module trained with a scene-paired motion dataset for interaction, our action controller performs locomotion and desired actions seamlessly by fulfilling the action cue and avoiding collisions in the 3D scene . Importantly, given the scene , action and initial state as inputs, our action controller is directly optimized to choose the most plausible motion clip in our motion database at each state, synthesizing natural human motions while taking the 3D scene into account without any scene-paired motion dataset or training procedure. Intuitively, given the current state , the goal of the action controller is to output the best next action which is used to search the next motion clip in Motion Synthesizer .
Action. Given the current status of the character , the control policy outputs the feasible action . provides the probabilities of the next action type among possible actions (e.g., walk, sit, or stop), determining the transition timing between actions (e.g., from locomotion to sitting). predicts future motion cues such as plausible root position for the next 10, 20, and 30 frames. Posture offset is intended to modify the raw motion data searched from the motion database in motion synthesizer module . Intuitively, our optimized control policy generates a posture offset to alter the closest plausible raw posture chosen in the database. This enables the character to perform more plausible scene-aware human poses only with human motion data. More details are addressed in Sec. 3.3.
3 Motion Synthesizer
By taking the current posture and actions signal from the action controller as inputs, the motion synthesizer produces the next plausible posture: . As the first step, the motion synthesizer searches for motion from a motion database that best matches the closest motion feature, then modifies the searched raw motion to be more suitable to the scene. To this end, the motion synthesizer’s output is in turn fed into the action controller recursively. We exploit a modified version of the motion matching algorithm for the first step of motion synthesis. In motion matching, motion synthesis is performed periodically by searching the most plausible next shot motion segments from a motion database, and compositing them into a long connected sequence.
Motion searching and updating. Given the query motion feature and the motion features in the motion database (where is the index of the clip), motion searching finds the best matches in the motion database by computing the weighted euclidean distances between the query feature and motion database features:
where is a fixed weight vector to control the importance of feature elements. After finding the best match from the motion database, the motion synthesizer updates it with the predicted motion offset from , that is , where is the next plausible character posture and is an update function to update selected joints in . In practice, motion searching is performed periodically (e.g., every N-th frame) to make the synthesized motion temporally more coherent.
4 Optimizing Scene-Aware Action Controller
The objective of our reinforcement learning framework is to optimize the policy by maximizing the discounted cumulative reward. In our method, we design the rewards to guide the character to perform both locomotion and desired actions (e.g., sitting) under common constraints (e.g., smooth transitions, and collision avoidance). Our reward function consists of the following terms:
As reported in , multiplying rewards with consistent goals can enforce all reward terms to be simultaneously met. We also use early termination and limited action transitions to accelerate learning. Details are in supp. mat.
5 Generalizing Action Controller
While our major focus of the use of RL is for a test-time optimization given a single target task, the optimized policy can handle variations of the task to some extent, as an advantage of the nature of RL. As shown in our experiments in Sec. 4.4, we demonstrate that our optimized controller can be directly used for various action cues and initials without further optimization for the same scene .
As an extension of our framework, we can make our controller more generalized by optimizing the policy with random variations of inputs, and per each episode during policy optimization. This procedure is more similar to the usual RL framework, where the policy is “learned” in advance for the target scene , and applied to the provided inputs during inference. We also demonstrate that our controller can handle a wider range of input variations via this augmentation process. This extension of our framework can provide better efficiency for the cases where varying tasks are instructed under a fixed 3D scene . As shown in Sec. 4.4, via the generalization process we can directly use the policy for diverse inputs without further optimization. Or, if necessary, efficiently fine-tuning the policy is also possible. Note that this extension still differs from other learning-based methods in that we do not require any scene-paired motion datasets or other supervision.
6 Task-Adaptive Motion Editing
To cover the diversity in interactions, we include a task-adaptive motion editing module in our motion synthesis framework. In particular, in the case of object manipulation, manipulation cue is provided to enforce an end-effector (e.g., a hand) to follow the desired trajectory expressing the manipulation task on the target object, as in Fig 6. The manipulation cue can be provided via any possible way, and in our experiments we produce it semi-automatically. We compute the desired trajectory by simulating the target articulated object’s motion by considering a contact point on the surface of the object mesh.
Experiments
We evaluate LAMA’s ability on synthesizing long-term motions in real-world 3D scenes with various human-scene and object interactions involved. We exploit an extensive set of quantitative metrics and perceptual studies for evaluation.
Dataset. For constructing the database for the motion synthesizer, we capture a new motion capture dataset involving locomotion and action. Motion is captured with IMU-based system XSens MVN Link . The collected data include high quality human motion with locomotion and interaction in various scenarios, such as walking around at different angles and sitting on a chair with random starting points. Captured motion data are post-processed to be suitable for motion matching. All the data used in this system are motion capture data (in bvh format) with no scene or object related prior information. We use PROX and Matterport3D datasets for 3D scenes and SAPIEN object meshes for manipulation. See supp. mat. for details.
Evaluation metrics. As our system does not rely on supervision for motion synthesis, quantifying synthesized quality is challenging due to the lack of ground-truth data or official evaluation metrics. We try to evaluate in terms of physical plausibility and naturalness.
\mathbin{\vbox{\hbox{\scalebox{0.5}{\bullet}}}} Physical Plausibility: We use contact and penetration metrics to evaluate the physical plausibility of the synthesized motions. Contact penalizes the foot movement when the foot is in contact. Since foot contact is a critical element in dynamics, the contact-based metric is closely related to determining the physical plausibility of motions. Penetration loss (“Penetration” in Table 1) measures implausible cases when the body penetrates the objects in the scene. We compute penetration metric by counting frames where intersection points (Sec. 3.2) go over a certain threshold. 10 for legs and 7 for arms
\mathbin{\vbox{\hbox{\scalebox{0.5}{\bullet}}}} Naturalness: We evaluate the naturalness of the synthesized motion via perception study (A/B test) on Amazon Mechanical Turk. The motions used for testing are rendered with the exact same view and 3D characters, making them indistinguishable from the appearance side. Human observers are asked to choose a more natural motion based on two criteria: (1) the character movement is human-like and (2) the movement is plausible in the given scene. Details of the study setup are in supp. mat.
Baselines. We compare LAMA with the state-of-the-art methods as well as variations of ours.
\mathbin{\vbox{\hbox{\scalebox{0.5}{\bullet}}}} Wang et al. is the state-of-the-art long-term motion synthesis method for human-scene interactions within a given 3D scene. We use the author’s code for evaluation. As Wang et al. post-processes synthesized motion to improve foot contact and reduce collisions which are directly related to our metric, we both compare Wang et al. with and without post-processing.
\mathbin{\vbox{\hbox{\scalebox{0.5}{\bullet}}}} SAMP generates interactions that can be generalized not only for object variations but also random starting points within a given 3D scene. SAMP explicitly exploits path planning to navigate through cluttered 3D scenes.
\mathbin{\vbox{\hbox{\scalebox{0.5}{\bullet}}}} Ablative Baselines We perform ablation studies on the action controller and motion editing module. We perform ablation studies on the scene reward , and action offset to present the contribution of both terms on generating scene-aware motions. We also compare our method without the transition reward and terms (Sec. 3.2) of the action controller. Finally, we demonstrate the strength of our motion editing module to edit motions naturally (Sec. 3.6) by comparing it with inverse kinematics (IK).
2 Comparisons with Previous Work
Evaluation Setup and Details. For comparison with baselines, we generate 50 motion sequences in total with random input and from 4 PROX 3D scenes used in testing for Wang et al. . Since our method is based on test-time optimization without explicit training and testing split, our action controller is optimized per each input, and no prior information on inputs is given before policy optimization. It takes 4 to 20 minutes to optimize a policy and 3 to 4 minutes (500 epochs) for optimization in the motion editing module. We only consider locomotion and action (walk-to-sit) motions and do not include manipulation as the baselines do not tackle manipulation. Contact metric is measured by the position difference of foot in contact, where contact is automatically labeled based on foot velocity. To compute penetration metric in a fair way, SMPL-X outputs of Wang et al. and SAMP are converted to box-shaped skeletons as in ours and intersection points are counted. Table 1 shows the results.
Physical Plausibility. As shown, LAMA outperforms both Wang et al. and SAMP in physical plausibility. Wang et al. post-processes the synthesized motion to ensure contact and reduce penetration, yet LAMA still outperforms. Moreover, our RL-based method with motion matching shows its advantage in collision avoidance in cluttered 3D scenes compared to path-planning based navigation in SAMP.
Naturalness. For perception study, we build two separate sets for comparison with Wang et al. and SAMP, and each testset is done with non-overlapping participants. For 50 motion sequences per set, 5 unique responses are collected per sequence for comparison. With Wang et al., LAMA received 215 votes while Wang et al received 35 (relative ratio 16.27%). With SAMP, LAMA received 176 votes, SAMP received 74 (relative ratio 42.04%). The results demonstrate that our method greatly outperforms baselines in terms of naturalness as well.
3 Ablation Studies
Ablation Studies on Action Controller. For quantitative ablations, we compare the original LAMA and the LAMA without collision reward . Ablation studies are performed in 5 PROX scenes. In original LAMA, penetration occurs in only 1.1% of the frames among the whole motion sequence, while the ratio is 15.7% in LAMA without . The result supports that the enforces the action controller to synthesize motions according to the given 3D scene. Example results are shown in Fig. 8. We also qualitatively compare the contribution of other components in the action controller. As seen in Fig. 9, without action offset the character does not tilt its limbs to avoid penetration with objects or walls, as the raw motion brought from the motion database does not have any information about the scene. This shows that also plays a role in generating detailed scene-aware poses. Moreover, the results without smoothness rewards and are not smooth enough, showing unnatural and abrupt movements.
Ablation Studies on Task-Adaptive Motion Editing. We ablate our motion editing module by replacing it with an alternative approach via IK. Same as , only the trajectory of a joint in contact (e.g., the right hand) is given to the IK solver. As shown in Fig. 10 (left), LAMA with motion editing module shows natural moves such as bending knees and tilting hips to make contact. However, results with IK show awkward poses as such spatiotemporal correlations in natural human motions are not reflected in the IK solver. Furthermore, as seen in Fig. 10 (right), the motion editing module makes the character properly sit in chairs with different shapes.
4 Robustness Test of Action Controller
As described in Sec. 3.5, utilizing RL for test-time optimization allows the optimized policy to handle variations in input. In this experiment, we aim to measure the extent to which a policy optimized for a single task and initial can generalize to varying inputs. To test the robustness with varying initials and tasks, we apply the optimized policy to all possible input variations in the scene and count the number of inputs the policy succeeds in synthesizing. From all possible initials sampled, the colored points in Fig 12 illustrate initial starting locations where the policy can synthesize motions meeting the given action cue . As shown in Table 2, a policy initially optimized for a single set of inputs (red in Fig. 12) can successfully synthesize motions even with distinct set of inputs without any additional optimization. Furthermore, to test the robustness of the generalized policy (described in Sec. 3.5), we perform the same test with the policy trained with our augmentation strategy during optimization. As shown above, it shows even more robustness in variations as expected.
We further demonstrate the generalization ability among unseen scenes with a policy optimized with the augmentation strategy. The generalized policy (Sec. 3.5) optimized in scene (scene in Fig. 12) are tested on two unseen scenes and from PROX shown in Fig. 13. As demonstrated in Table 3, an generalized policy (Sec. 3.5) optimized with scene can be generalized to some extent to scene , as and shares a similar structure (sofa and chairs around a table). However, as expected, the generalization ability decreases when tested on a totally distinct scene .
Note that the inference time here is about 0.2-3 sec per input as no further policy optimization is required. Details of the test setup are in supp. mat.
Discussion
We present a unified framework to synthesize human motions within complex real-world 3D scenes with motion-only datasets. We formulate it as a test-time optimization, leveraging RL with motion matching for realistic motion synthesis, and also utilize motion manifold to further cover the diversity of manipulation behaviors. Our method has been thoroughly evaluated in diverse scenarios, outperforming previous approaches .
Despite RL is used for test-time optimization, a single policy can cover variations in input and can also be generalized for extensive variations. Combining this framework with supervised learning for further efficiency increase can be an interesting future research direction. Furthermore, although we assume a fixed skeleton throughout the system, interaction motions may change depending on the character’s body shapes and sizes. We leave synthesizing motions on varying body shapes as future work.
This work was supported by SNU-Naver Hyperscale AI Center, SNU Creative-Pioneering Researchers Program, NRF grant funded by the Korea government (MSIT) (No. 2022R1A2C2092724), and IITP grant funded by the Korea government (MSIT) (No.2022-0-00156 and No.2021-0-01343). H. Joo is the corresponding author.
Appendix A Supplementary Video
The supplementary video shows the results of our method, LAMA, on various scenarios. In the video, we show our human motion synthesis results on PROX , Matterport3D , and also our own 3D scene scanned by Polycam App with an iPad pro. We use SAPIEN object meshes to semi-automatically produce manipulation cues, which is also shown in our videos. As shown, our method successfully produces plausible and natural human motions in many challenging scenarios.
While our original pipeline is designed for test-time optimization, in our video we also qualitatively demonstrate the strength of our framework in generalized scenarios by using a single optimized policy in handling different inputs without further optimization. In the video, we also show a policy optimized via our augmentation strategy (Sec. 3.5.) can handle more extensive input variations.
Appendix B More Details on Experiments
In this section, we describe further details on our experiments on the Robustness Test of Action Controller (in Sec. 4.4) and the Perception Study (in Sec 4.2).
In Sec. 4.4, Fig 10, and Table 2 of our main paper, we demonstrate that a single policy optimized for a specific input can handle varying target actions and initial . We describe more details on the experiment in Sec 4.4. For the experimental setup, we consider all possible variations for the input to test the generalization ability of the policy trained to a specific input. Specifically, a set of “all” valid initials is automatically chosen via grid sampling of the floor plane for the locations , by excluding points occupied by objects, with a random body orientation for . For the action target , we manually choose multiple plausible locations (e.g., chairs) for the actions. In the test scene we use in Fig. 10, there exist 2635 plausible initial positions and we consider 4 target action cues shown in the white boxes in Fig. 10.
The original policy (Fig. 10 top) is optimized to a specific input and action cue , marked as red in the top left of Fig. 10. The colored points in Fig. 10 show the locations where the policy achieves the goal successfully without any further optimization for the policy. For each input pair and , we perform the motion synthesis with the policy 5 times. In each trial, the initial body orientation is chosen randomly to provide more variations. We determine the policy is successful for the current initial location when no early termination conditions (collision, stall, moving out of the scene) are met while fulfilling at least twice out of the 5 trials. As shown in Fig. 10 and Table 2, our action controller optimized for a specific target can be applicable to many input variations.
We perform the same test for the generalized policy (described in Sec. 3.5) in the bottom of Fig. 10 and Tab. 2. As shown, this policy can cover much more extensive input variations on the same scene.
Comparison of Computation Time. As a test-time optimization without requiring scene-paired motion datasets, our original framework takes time to train a policy from a scratch for a given input pair. However, reusing the same policy that is optimized for the specific input for other inputs can greatly reduce the computation time, because no further optimization is needed for the policy. To compare the time between performing the inference only and optimizing a policy from scratch, we test with 5 input pairs consisting of initial and . Here, the term “inference only” indicates that we use a pre-optimized policy without any further optimization for varying inputs. As the result, the inference-only scenario takes 0.15 seconds on average per input pair for motion synthesis, while optimizing a policy from scratch per pair takes 6.32 minutes (379 seconds) on average. As shown, the capability of the reinforcement learning framework provides the potential to greatly improve the efficiency of our method.
Motion Quality Measurement. We also evaluate the physical plausibility of the synthesized motion in the robustness test in Sec. 4.4. An optimized policy synthesized 15 motion sequences with distinct input pairs (the input pair which the policy is initially optimized to is not included). We also perform the measurement to motions synthesized by the generalized policy optimized with an augmentation strategy. The results are shown in Table 4. This shows while a policy can handle variations in input, there is no performance drop in the synthesized motion quality.
B.2 Perception Study Setup
The videos used for perception study are in the supplementary video. We include 3 videos per set to the supplementary video.
Appendix C More Details on Implementations
The policy and the value network of the action controller module consists of 4 and 2 fully connected layers of 256 nodes, respectively. The control policy is optimized through Proximal Policy Optimization (PPO) algorithm . Adam optimizer is used with Nvidia RTX 3090 GPU. For the action controller and motion synthesizer module , we use the animation library DART . We also use a publicly available PPO implementation , where we remove the variable time-stepping functions stepping in by following the original PPO algorithm. The details of the optimization regarding the policy and value network of the action controller are written in Table 5.
Acceleration Techniques.
As written in the main paper, we use early termination conditions to accelerate policy optimization. The episode is terminated when (1) the character moves out of the scene bounding box; (2) when the collision reward is under a certain threshold; and (3) the root velocity for a specific time duration (50 frames) is under a certain threshold to prevent the character standing still for a overly long time. Also, the action controller first checks in advance whether the action signal is valid when it makes transitions from locomotion to other actions. When the nearest feature distance of Eq. 2 in the motion synthesizer (Sec. 3.3) is over a certain threshold, the action controller discards the transition and continues navigating.
C.2 Motion Synthesizer
Motion is captured by an IMU based system XSens MVN Link and is post-processed via XSens MotionCloud software . The captured motion is then retargeted to a single unified skeleton using Autodesk MotionBuilder and is post-processed to be suitable for motion matching. For action motions we mirror the motion segments for data augmentation. The length (in frames) of motion segments (“Seg. Length” in tables), number of motion segment (“Seg. Count” in tables), and the number of total frames (“Total Frames” in tables) are summarized in Table 6.
Action-Specific Feature Definition.
C.3 Motion Editing via Motion Manifold
The encoder and decoder of the task-adaptive motion editing module consist of three convolutional layers. For the convolutional autoencoder of task-adaptive motion editing, we use PyTorch , FairMotion , and PyTorch3d . The autoencoder is trained with the Adam optimizer with learning rate 0.0001. We use Nvidia RTX 3090 GPU. We use 3 layers of 1D temporal-convolutions with kernel width of 25 and stride 2, and the channel dimension of each output feature is 256. For training the autoencoder module in task-adaptive motion editing we use data in Mixamo , Lafan1 , COUCH , and ours. The training datasets are summarized in Table 7. Note that data used for training the autoencoder also does not include scene related information (in bvh format), and we use different pre-processing steps between the Motion Editing module and the Motion Synthesizer.
Reconstruction Loss.
Motion Editing Loss.
For motion editing, the positional loss and regularization loss are defined as follows.
Generating Manipulation Cues from SAPIEN [69].
While the manipulation cue can be provided via diverse ways depending on the applications, we mainly consider the scenarios of interacting with articulated objects. For this purpose, we semi-automatically produce the manipulation cues by extracting the desired target vertex trajectories of the parts of articulated objects from the SAPIEN dataset . Specifically, we place a target object in our 3D scene, and choose a target vertex of the object where we assume the character’s hand contacts to manipulate the target part (e.g., a vertex in the lid of a trash can object). Then, the trajectory of the vertex can be obtained by varying the parameter for the articulated motion with a fixed interval, where , , are the global orientation and translation of the object and is the parameters for the object articulation (e.g., the hinge angle of the cover of a laptop) at time . represents the 3D location of the chosen vertex given the parameters. The resulting manipulation cue is the target trajectory that a hand joint should follow for the manipulation motion. Note that our system requires only the manipulation cue , and the 3D object mesh is shown only for visualization purposes, where we visualize it with the synced .
Further Implementation Details for Manipulation
Depending on possible applications (e.g, sitting down and opening a laptop), the manipulation motion may need to be “added” in the middle or the end of the synthesized motion . In this case, we simply duplicate the target frame by to build a longer motion , and apply the motion editing to the target motion segment that is a stationary motion produced via the duplication.