Object Motion Guided Human Motion Synthesis

Jiaman Li, Jiajun Wu, C. Karen Liu

Introduction

Capturing and synthesizing human movements in contextual environments is critical to progressing embodied AI, character animation, VR/AR, and robotics. The real world in which humans live is complex and highly dynamic. Humans routinely interact with dynamic objects to accomplish everyday tasks, demonstrating a diverse range of full-body manipulations. For example, humans pull and push a mop to tidy a floor, reposition a floor lamp to illuminate a specific area, drag a chair toward a desk, and place a monitor on a desk. Realistically simulating such complex manipulation behaviors is a fundamental problem in computer graphics with a lot of downstream applications.

Prior works have made significant progress in addressing the contextual human motion synthesis problem for activities such as navigating through a 3D scene or sitting on a chair (Wang et al., 2021a, b; Hassan et al., 2021; Zhang et al., 2022a; Zhao et al., 2023; Mir et al., 2023). They model interactions with static 3D scenes or static objects based on large-scale human motion datasets. In comparison, datasets containing full-body interaction with moving objects are scarce. Prior works rely on reinforcement learning to model such behaviors (Merel et al., 2020; Hassan et al., 2023; Xie et al., 2023), but the learned policies are often limited to manipulating specific types of geometry used for training.

We present a new approach to synthesizing the dynamic interactions between humans and large-sized objects, particularly in manipulation tasks requiring full-body movements and precise coordination between hands and objects. We aim to bridge the gap between current research and real-world manipulation behaviors by introducing a large-scale dataset and developing a robust approach to synthesize full-body motion from object motion.

We present a new framework – Object MOtion guided human MOtion synthesis (OMOMO). We leverage a conditional diffusion formulation to predict plausible full-body poses with a sequence of object geometry as input. One key observation is that hand position is a deciding factor for full-body movement during manipulation. Thus, we devise a two-stage approach to generate hand positions conditioned on object geometry features and then synthesize full-body poses based on the predicted hand positions. The two-stage design enables us to apply contact constraints to our predicted hand joint positions, which significantly enhances the contact realism of the generated results. We demonstrate the effectiveness of our proposed method in our dataset and showcase its generalizability to unseen objects.

Moreover, we introduce an innovative application that generates full-body human poses based on object motion captured by an iPhone. In particular, we mount an iPhone on an unseen object, employ the iPhone ARKit to obtain camera poses and deduce the motion of the object. Subsequently, we apply these object poses to 3D geometry reconstructed using Luma (AI, 2023). Our pipeline takes the sequence of object geometry as input and generates the corresponding full-body human motion. This application demonstrates an affordable and user-friendly method for capturing human interaction motions during everyday tasks.

An additional contribution of this work is a new dataset with paired object motion and human motion to facilitate the learning of full-body human manipulation behaviors. We leverage an advanced 3D reconstruction technique to extract 3D object geometry from a monocular video. We then use motion capture devices to capture human and moving objects simultaneously. To capture motions that resemble real-world scenarios, we provide language descriptions to guide our volunteers to perform meaningful interactions with various objects. Our dataset can be used for different tasks to model full-body human manipulation behaviors.

To summarize, the contributions of this work include:

A novel approach to full-body manipulation synthesis by generating full-body motion from object motion. We introduce an effective framework based on conditional diffusion to synthesize full-body movements from object motion.

A novel application that employs an iPhone to capture object motion from the egocentric view of the object, enabling the synthesis of full-body movements by simply attaching an iPhone to various objects.

A large-scale high-quality dataset consisting of 3D object geometry, object motion, and full-body motion.

Related Work

Human motion modeling has been extensively studied with motion capture datasets (Mahmood et al., 2019). Recently, there has been a surge of interest in human scene interactions. PROX (Hassan et al., 2019) provides paired 3D scenes and human motions extracted from RGB videos. HPS (Guzov et al., 2021) contributes a dataset of paired scenes, egocentric video, and human motion captured with an IMU-based suit. EgoBody (Zhang et al., 2022c) collects a dataset consisting of 3D scenes, egocentric video, eye gaze, and human motions extracted from multi-view RGBD frames with a focus on social interactions. GIMO (Zheng et al., 2022) explores the problem of gaze-guided motion prediction using a similar data modality. Synthetic datasets (Li et al., 2023; Wang et al., 2022) combine scene datasets (Straub et al., 2019; Dai et al., 2017) with motion datasets (Mahmood et al., 2019) to produce paired human motions in 3D environments. CIRCLE (Araujo et al., 2023) integrates VR and MoCap techniques to collect high-quality motion within virtual scenes.

A couple of datasets focus on human-object interactions. For example, SAMP (Hassan et al., 2021) contains sitting and lying down motions while interacting with chairs and sofas. COUCH (Zhang et al., 2022a) is dedicated to data collection for sitting on different chairs. These datasets primarily contain motions interacting with static objects.

Moreover, some datasets collect both human motion and object motion (Taheri et al., 2020; Bhatnagar et al., 2022; Guzov et al., 2023; Fan et al., 2023). GRAB (Taheri et al., 2020) focuses on the interaction between humans and small-sized objects, involving mostly hand motions. BEHAVE (Bhatnagar et al., 2022) records interactions with larger-sized objects, making it closely related to our dataset. However, it relies on multi-view RGBD input to extract human and object motion, which does not yield motion of sufficient quality for motion synthesis tasks. Furthermore, the limited data for each object impedes its capacity for training a motion generative model. In contrast, our work focuses on synthesizing dynamic human interactions with large-sized objects and we introduce a large-scale dataset consisting of high-quality human motion and object motion.

Contextual Human Motion Synthesis.

Motion synthesis is a long-standing problem in computer graphics, and here we survey prior works centered on motion synthesis in 3D environments. Leveraging the dataset with paired scenes and human motions (Hassan et al., 2019), a couple of work (Wang et al., 2021b, a) learn separate modules to predict root trajectory first and generate full-body poses conditioned on both scene and the planned path. However, constrained by the scale and motion quality of the dataset, these methods struggle to synthesize realistic human motions. To improve the motion quality of generation results, SAMP (Hassan et al., 2021) collects a high-quality dataset consisting of walking, sitting and lying down motions. And they present a pipeline that first produces a collision-free path based on A∗ algorithm, generates full-body motion following the path and then synthesizes interaction motions to sit on chairs and sofas. A recent work (Mir et al., 2023) introduces action keypoints as scene abstraction, enabling continual motion synthesis generation across various scenes. In order to produce physically plausible movements, several works employ reinforcement learning techniques to learn interaction policies through meticulously designed task rewards (Hassan et al., 2023; Lee and Joo, 2023; Chao et al., 2021).

Another line of work focuses on reaching motion synthesis within contextual environments. GOAL (Taheri et al., 2022) and SAGA (Wu et al., 2022) generate full-body poses aimed at grasping a specific object. IMoS (Ghosh et al., 2023) further synthesizes human and object motions simultaneously after grasping an object. However, these works only consider the target object and do not involve navigation in cluttered scenes. Meanwhile, CIRCLE (Araujo et al., 2023) incorporates human-scene interaction features and formulates the problem with a scene-aware motion refinement model, enabling reaching synthesis in complex static scenes.

While most existing work that involves interaction with dynamic objects aims to synthesize dexterous hand motions (Ye and Liu, 2012; Li et al., 2007; Zhang et al., 2021), our work diverges from this line of work. Instead, we focus on the synthesis of full-body movements for manipulation without synthesizing detailed hand movements.

Full-body human motion synthesis for manipulation has been explored in both kinematic-based (Starke et al., 2019) and physics-based methods (Hassan et al., 2023; Merel et al., 2020; Xie et al., 2023). NSM (Starke et al., 2019) learns a gating network and a motion prediction network to synthesize interaction movements including sitting and carrying objects. As for physics-based character animation, reinforcement learning has been widely used to learn different skills (Peng et al., 2018, 2021; Xie et al., 2022; Liu and Hodgins, 2018). In terms of manipulation, Merel et al. (2020) devise a hierarchical reinforcement learning framework to synthesize box catching and carrying movements with egocentric observations. More recently, Hassan et al. (2023) propose to learn policies based on the Adversarial Motion Priors framework (Peng et al., 2021) for box manipulation task.

In summary, most prior research has not considered the dynamic interaction between humans and large-sized objects. A few works studied the problem of full-body manipulation but were constrained to interactions with boxes. In contrast, our work examines contextual environments with diverse dynamic objects. Leveraging our large-scale dataset, we develop an approach to synthesize manipulation movements for diverse objects. And inspired by the success of diffusion in motion modeling (Li et al., 2023; Dabral et al., 2023; Tevet et al., 2023; Zhang et al., 2022b; Tseng et al., 2023; Huang et al., 2023), we design our framework based on conditional diffusion.

Method

Object Representation

2. Conditional Diffusion Formulation

The diffusion model consists of a forward diffusion process and a reverse diffusion process. The forward diffusion process is gradually adding noise to the data representation x0\bm{x}^{0} for NN steps formulated using a Markov chain,

The transition of forward diffusion is modeled by a posterior distribution qq. And each step is decided by a fixed variance schedule using βn\beta_{n} and is defined as

where I\bm{I} represents identity matrix.

The reverse diffusion process is to generate desired data representation from random noise xN∼N(0,I)\bm{x}^{N}\sim\mathcal{N}(0,\bm{I}). This is achieved by learning a neural network pθp_{\theta} to denoise recursively. Specifically, at noise level nn, we use c\bm{c} to represent the conditions, and we have the reverse diffusion process represented as follows:

where μθ(xn,n,c)\bm{\mu}_{\theta}(\bm{x}^{n},n,\bm{c}) is the learned mean, σn\sigma_{n} is the fixed variance. μθ(xn,n,c)\bm{\mu}_{\theta}(\bm{x}^{n},n,\bm{c}) (we use μθ\bm{\mu}_{\theta} in the following equation for brevity) can be formulated as,

where x^θ(xn,n,c)\hat{\bm{x}}_{\theta}(\bm{x}^{n},n,\bm{c}) is the prediction of x0x^{0}, αn,αˉn\alpha_{n},\bar{\alpha}_{n} are fixed parameters that satisfy αˉn=∏i=1nαn\bar{\alpha}_{n}=\prod_{i=1}^{n}\alpha_{n}.

Learning the mean can be reparameterized as learning to predict the original data x0\bm{x}^{0}. We use reconstruction loss of x0\bm{x}^{0} during training:

3. Our Pipeline

In the first stage, we employ conditional diffusion to generate hand joint positions H1,H2,...,HT\bm{H}_{1},\bm{H}_{2},...,\bm{H}_{T} from object geometry features O1,O2,...,OT\bm{O}_{1},\bm{O}_{2},...,\bm{O}_{T}. Here, the conditions c\bm{c} are represented by O1,O2,...,OT\bm{O}_{1},\bm{O}_{2},...,\bm{O}_{T}. We adopt a transformer model architecture (Vaswani et al., 2017) as our denoising network which consists of four self-attention blocks. Each self-attention block contains a multi-head attention layer followed by a position-wise feedforward layer. As shown in Figure 3, we introduce an additional step to include noise level embedding as an input to our transformer model.

Apply Hand Contact Constraints.

The hand joint positions generated in the initial stage may not always be precise. They may occasionally deviate from the object, resulting in perceived non-contact at certain time steps. To mitigate this, we propose a post-processing strategy based on the observation that human hands typically maintain consistent relative positions with respect to objects during contact.

Given a sequence of hand joint positions H1,H2,...,HT\bm{H}_{1},\bm{H}_{2},...,\bm{H}_{T}, we begin by computing the minimum distance from the hand joints to the corresponding object mesh V1,V2,...,VT\bm{V}_{1},\bm{V}_{2},...,\bm{V}_{T} at each time step, denoted as d1,d2,...,dTd_{1},d_{2},...,d_{T}. We then traverse the sequence d1,d2,...,dTd_{1},d_{2},...,d_{T} starting from the first frame. We set an empirical contact threshold th=0.03\text{th}=0.03 and record a specific time step kk where dk<thd_{k}<\text{th}.

Generating Full-body Poses from Hand Positions.

In the second stage, we utilize the same denoising network architecture as in stage one to generate full-body poses from the hand joint positions. The conditions in this stage are the hand joint positions (H^1,H^2,...,H^T\hat{\bm{H}}_{1},\hat{\bm{H}}_{2},...,\hat{\bm{H}}_{T}) that have been rectified using the contact constraints. The model is trained using human motion data only.

By integrating these three components, we establish a complete pipeline to generate full-body poses from object motion. This pipeline models the one-to-many mapping from object motion to human poses and ensures that the generated poses maintain realistic contact with the object.

Dataset

We collected a large-scale high-quality dataset consisting of 3D object geometry, human and object motions. In this section, we elaborate on our object geometry acquisition, motion capture, and data processing.

We selected 15 objects commonly used in everyday tasks, which include a vacuum, mop, floor lamp, clothes stand, tripod, suitcase, plastic container, wooden chair, white chair, large table, small table, large box, small box, trashcan, and monitor. For each object, we filmed a video circling the object and employed Luma (AI, 2023) to reconstruct the 3D object geometry from this monocular video. We then utilized Meshlab to manually remove noisy points and downsample object meshes to contain a reasonable number of points for training.

Motion Capture

We utilized a Vicon system comprised of 12 cameras controlled by Vicon Shogun, which record at a rate of 120 FPS. For each object, we attached 5 markers and captured the object and human motion simultaneously. We invited 17 subjects (13 males, 4 females) to participate in our motion capture sessions. During each mocap session, the volunteer was provided with verbal instructions on how to interact with each object to avoid meaningless interactions. We show some examples of our language guidance in Figure 4. Each mocap session lasted approximately 1.5 to 2 hours. The total duration of captured motion for each object is shown in Table 1.

Data Processing

For the object geometry data, we employed a public python library (Kleineberg, 2023) to compute the SDF for objects. In cases where objects contained noisy SDFs, we used SIREN (Sitzmann et al., 2020) to train neural networks and extract the SDF.

In terms of motion data processing, we used Mosh++ (Loper et al., 2014; Mahmood et al., 2019) to process our raw mocap files and extract SMPL-X model (Pavlakos et al., 2019) parameters for each sequence. In order to compute object transformations based on marker positions, we initially manually annotate the marker positions on the reconstructed object mesh. Subsequently, we utilize the analytical solution of the orthogonal Procrustes problem to compute the scale, rotation, and translation needed to align the annotated points with the marker positions. Furthermore, we visualize the object and human meshes, and conduct a manual verification on our collected dataset, discarding any sequences that fail to meet our high-quality standard.

Experiment

We first introduce the dataset and evaluation metrics used for this task. Then we describe the chosen baselines and showcase comparisons against them. Additionally, we conduct an ablation study to investigate the effects of hand positions on overall performance. We encourage readers to watch our supplementary video for more qualitative evaluations.

We conduct all experiments using our collected dataset. This dataset consists of motion capture data from 17 subjects, with 15 subjects used for training and 2 subjects for testing. We adopt two data partitioning for evaluation. In the first setting, we use 15 objects for both training and testing. To further evaluate the model’s generalization ability to new objects, we divide the 15 objects into 10 for training and 5 for testing as shown in Figure 5.

Evaluation Metrics.

We evaluate the synthesized results from two perspectives. Firstly, we compare the generated poses against the ground truth motion data. Additionally, we assess the physical plausibility of the results, considering contact correctness, object penetration, and foot sliding. We detail our evaluation metrics as follows.

HandJPE, MPJPE and MPVPE represent mean hand joint position errors, mean per-joint position errors, and mean per-vertex errors in centimeters (cm)\left(cm\right).

TrootT_{root} and OrootO_{root} represent the root translation error computed using Euclidean distance in centimeters (cm)\left(cm\right) and orientation error defined by the Frobenius norm of the difference between the 3 × 3 rotation matrix ∣∣RpredRgt−1−I∣∣2||R_{pred}R^{-1}_{gt}-I||_{2}.

FS represents foot sliding metric and is computed following previous work (He et al., 2022).

Collision Percentage. At time step tt, for iith vertex on reconstructed human mesh, we query the object SDF and acquire a signed distance value dtid_{t}^{i}. We use a threshold (44cm) to compute collision. If there exists vertices that satisfy dti<0,∣dti∣>4d_{t}^{i}<0,|d_{t}^{i}|>4, we increment the collision count. By traversing the sequence, we can compute the collision percentage.

Contact Metrics. We adopt metrics precision (CprecC_{prec}), recall (CrecC_{rec}), and F1 score from the object detection task to evaluate contact performance. We first compute the distance between hand positions and object meshes. We empirically set a contact threshold (55cm) and use it to extract contact labels for each frame. We perform the same calculation for ground truth hand positions. Then we count true/false positive/negative cases to compute precision, recall, and F1 score.

2. Evaluations

Since no existing work specifically addresses the task of object motion-guided human motion synthesis, we adapt a prior work GOAL (Taheri et al., 2022) on object-reaching motion synthesis as our baseline. GOAL proposed an autoregressive model that predicts future 10 frames conditioned on past 5 poses, hand distance between the current frame and the target goal frame, and the BPS representation which encodes hand-to-object distance at the target frame. In our problem setting, the input is a sequence of object geometry that guide the motion generation instead of a single target frame in GOAL. Thus, we make changes to the input features and use the next frame as the target frame. Specifically, the input features in our modified version consist of the past 5 poses, and the BPS representation that encodes the distance features between the current human mesh and the object mesh at the next frame. We use their default model setting which consists of four learning blocks. Each block contains a set of MLPs. The output dimension for each block is 2048, 1024, 1024, 2048 respectively.

Implementation Details.

Our denoising network in OMOMO-single-stage, stage 1 model of OMOMO, and stage 2 model of OMOMO all consist of 4 self-attention blocks with 4 attention heads. The dimension of key, query, and value in the transformer architecture is 256. The output dimension of each layer is 512. Our implementation uses PyTorch (Paszke et al., 2019). For training stage 1 and stage 2, we both use Adam optimizer (Kingma and Ba, 2015) and start the training with a learning rate 0.0002. The training takes about 18 hours to converge for both stage 1 and stage 2 using a single NVIDIA Titan RTX GPU.

Results.

Since our approach is based on conditional diffusion, there can be multiple plausible generation results given the same object motion. To make a quantitative comparison, we sample 20 times for the same object motion input and select the one with the smallest MPJPE. We show quantitative evaluations in Table 2 and Table 3 for two different data splits. One splits training and testing on all 15 objects. The other one uses 10 objects for training and the other 5 unseen objects for testing. For each configuration, there is only one random seed used to train our model, and the statistics were computed using a single model. Note that our training is not sensitive to random seeds. We outperform baselines in both settings. In particular, OMOMO has superior results in terms of contact evaluations compared to the other two OMOMO variants, which demonstrate the effectiveness of our two-stage design and contact constraints. Note that the reason for the smaller collision percentage in GOAL is that the character in the baseline results often does not attempt to manipulate the object at all, hence the low collision percentage. As for smaller FS scores, we observed that the feet position in the baseline results usually drifts above the floor which will not be counted as foot sliding according to the foot sliding metric. In addition, it is worth mentioning that applying the contact constraints to GOAL is not straightforward. Since GOAL predicts all the joints’ rotations, it requires inverse kinematics to rectify the human pose based on the corrected hand positions.

We also showcase qualitative results in Figure 6. OMOMO contains better contacts compared to the setting without hand joint positions as an intermediate representation, as evidenced by Figure 6 and higher contact F1 scores. For more qualitative comparisons, please watch our supplementary video.

Human Perceptual Study

We further conduct a human perceptual study to complement the evaluations. The goal is to evaluate the motion quality and contact realism. We random sample 100 generated sequences for each approach including OMOMO, OMOMO-single-stage, GOAL, and ground truth, covering all 15 objects. We show some generated results of each approach in Figure 8. We compare OMOMO and the other three settings and totally form 300 pairs for evaluation. For each question, we ask amazon mechanical turk workers which sequence looks more natural and interacts with objects more realistically. Each question is evaluated by 20 different workers ( Figure 7).

We show that our OMOMO clearly outperforms the baseline GOAL and OMOMO-single-stage. And when compared with ground truth, 31% preferred our results (the upper bound would be 50%). It is worth noting that our results of OMOMO are produced via a single forward pass, without any optimization or post-processing for the full-body poses. Therefore, certain artifacts such as penetration may be produced in the generated motion, which results in ground truth motion is preferred in some sequences.

3. Ablation Study

To investigate the effects of hand positions on our overall performance, we compare the full-body human poses generation results that use the predicted hand joint position as input (OMOMO) and ground truth hand positions as input (OMOMO-GT). In Table 4, we show that the synthesis results can be further improved by feeding more accurate hand joint positions.

4. Test on Manually Animated Object Trajectory

We further evaluated our pipeline using manually crafted animations of previously unseen objects. In this process, we began by reconstructing the 3D geometry of the object with the aid of Luma (AI, 2023). Once reconstructed, the 3D object was imported into Blender. Within Blender, we manually established keyframes at 15-frame intervals. Based on these keyframes, Blender then produced a complete object motion sequence. This sequence, exported from Blender, served as the input for our OMOMO. The resulting outputs are shown in Figure 9.

Application

We introduce our novel approach to capturing human motion interacting with objects using a single smartphone attached to the object. Specifically, we mount an iPhone XR on the target object and ask the subject to interact with the object while the iPhone camera is filming the environment. We leverage the API ARWorldTrackingConfiguration provided by iPhone ARKit to extract camera poses. This feature is based on visual-inertial odometry techniques that combine visual information and sensor information to estimate accurate camera pose in the world coordinate system. Since the camera is rigidly mounted on objects, we can derive object motion from camera poses. Similar to the data collection process, we film a video and use Luma (AI, 2023) to reconstruct 3D geometry of the target object. From a sequence of object-moving geometries, we can generate full-body human poses with our proposed pipeline. We showcase some results in Figure 10. Note that these objects are not used during model training.

Conclusion

In summary, we presented a novel approach for synthesizing human motion guided by moving objects. Specifically, we proposed a framework based on a two-stage paradigm to enforce contact constraints, demonstrating its effectiveness in generating realistic human motions in interaction. Moreover, we introduced a novel application that enables capturing human interaction motion using a smartphone only. To facilitate the research on human-object interactions, we also introduced a large-scale dataset consisting of 3D object geometry, high-quality object motion, and human motion.

Our current dataset falls short of accurately representing dexterous hand movements, which often result in implausible hand motions. A promising avenue for future research would be incorporating hand priors and optimization techniques, enhancing the realism of hand motions in our full-body pose generations. Furthermore, the contact constraints in our current framework cannot effectively address scenarios with intermittent contacts with the object as shown in Figure 11. This could be addressed by identifying and predicting contact states to enable the generation of more complex, long-term manipulation with the objects. Lastly, while our methodology is based on kinematics, future efforts could benefit from integrating physics-based components to mitigate the occurrence of artifacts.

References