NIFTY: Neural Object Interaction Fields for Guided Human Motion Synthesis

Nilesh Kulkarni, Davis Rempe, Kyle Genova, Abhijit Kundu, Justin Johnson, David Fouhey, Leonidas Guibas

Introduction

Animating a character to sit in a chair or pick up a box is useful in gaming, character animation, and populating digital twins. Yet, generating realistic 3D human motion trajectories with objects is challenging for two main reasons. One challenge is creating effective models that capture the nuance of human movements, particularly during the final phase of object interaction called the "last mile." Unlike navigation that is primarily collision avoidance, the last mile involves intricate contacts and object affordances, which influence the motion. The second challenge is acquiring paired data that includes high-quality human motions and diverse object shapes, which is essential for training.

Recent approaches to motion modeling can synthesize realistic human movements using state-of-the-art generative models . They are, however, scene-agnostic and cannot produce interactions with specific objects.

To address this, some approaches condition motion synthesis on scene geometry (e.g., a scanned point cloud) . This enables learning object interactions, but motion quality is hindered by the lack of paired full scene-motion data. Other approaches instead focus on a small set of interactions with a single type of object (e.g., sitting in a chair), and can produce high-quality motions in their domain. However, these methods require a high-quality motion capture (mocap) dataset for each action/object separately and may make action-specific modeling assumptions (e.g., affecting the human approach and/or contact with object).

In this work, we tackle both the modeling and data aspects of interaction synthesis to enable generating realistic interactions with a variety of objects, such as sitting on a chair, table, or stool and lifting a suitcase, chair, etc.. We extend a human motion diffusion model to condition on object geometry, and pair it with a learned object interaction field to encourage realistic movements at test time in the last mile of interaction. To train this model and overcome the lack of available mocap data, we develop an automated data pipeline that leverages a powerful pre-trained and scene-agnostic human motion model. As shown in Fig. 1, our interaction field, diffusion model, and motion data pipeline make up a general framework to synthesize human-object interactions for a desired character that is flexible to multiple actions, even when dense mocap data is unavailable. We refer to this framework as NIFTY: Neural Interaction Fields for Trajectory sYnthesis.

To ensure realistic motions in the last mile of interaction, we propose an object-centric interaction field that takes in a human pose and learns to regress the distance to a valid interaction pose (e.g., the final sitting pose). At test time, our object-conditioned diffusion model is guided by this interaction field to encourage high-quality motions. Unlike manually designed guidance objectives that encourage contact and discourage penetration , our interaction field is data-driven and implicitly captures notions of contact, object penetration, and any other factors learned from data.

We propose using synthetic data generation to enable learning interactions from limited mocap data. In particular, we leverage a pre-trained motion model that produces high-quality motions but is unaware of object geometry. Starting from an anchor pose that captures the core of a desired interaction (e.g., the final sitting pose in Fig. 1, right), the pre-trained model is used to sample a large variety of motions that end in the anchor pose. This approach generates a diverse set of plausible interactions from only a handful of anchor poses, which are readily available from existing small datasets or are relatively easy to capture.

We evaluate NIFTY on sitting and lifting interactions for a variety of objects, demonstrating the superior quality of synthesized human motions compared to alternative approaches. Overall, this work contributes (1) a novel object interaction field approach that guides an object-conditioned human motion diffusion model to synthesize realistic interactions, (2) an automated synthetic data generation technique to produce large numbers of plausible interactions from limited pose data, and (3) high-quality motion synthesis results for human interactions with several objects.

Related Work

Synthesizing Human Motion and Interactions. While various methods have been successful in generating human motion in isolation , our work is primarily focused on incorporating environmental context . Some approaches condition motion generation on scanned scene geometry that encompasses multiple objects , but these methods typically offer limited control over the specific objects for interaction. Object-centric models are trained to generate motions for a single character and limited actions , such as sitting on a chair. These models heavily rely on high-quality motion capture datasets but still exhibit issues like floating, skating, and penetration. Our work also focuses on individual objects but utilizes diffusion guidance and a learned interaction field to minimize undesired artifacts. In contrast to prior work, we train our models using a novel data generation pipeline to learn interactions from limited data. Our focus is on macroscopic interactions like sitting and lifting with objects, distinguishing us from other works that generate full-body motions for grasping and manipulation .

Motion Modeling with Diffusion. Following success for image and video generation, diffusion models have shown promise in modeling motion for robots and pedestrians . Recently, diffusion models have been successful in generating full-body 3D human motion . SceneDiffuser generates human motion conditioned on a point cloud from a scanned scene. It employs gradient-based guidance and analytic objectives to ensure collision-free, contact-driven, and smooth motion during the denoising process. On the contrary, our approach is object-centric and does not rely on noisy motions for training. Our data-driven interaction field guides denoising by implicitly capturing plausible interactions and obviating the need for hand-designed objectives.

Neural Distance Fields for Pose and Interaction. Neural networks have been used to learn a parametric function that outputs a distance given a query coordinate . Grasping Fields parameterize hand-object grasping through a spatial field that outputs distances to valid hand-object grasps. Pose-NDF learns an object-unaware distance field in the full-body pose space for human poses. NGDF and SE(3)-DiffusionFields learn a field in the robot gripper pose space to define a manifold of valid object grasps. Our object interaction field extends this idea to full-body human-object interactions by learning to predict the distance between a human pose and the interaction pose manifold. Unlike prior works, we use this field to guide denoising.

Human Interaction Data. Though large-scale mocap data is available to train scene-agnostic human motion models , learning human-object interactions is hampered by the challenge of capturing humans in scenes. Datasets that contain full scene scans paired with human motion are relatively small and often noisy due to capture difficulties. Other datasets contain single-object interactions with a small set of objects . These are better quality due to simpler capture conditions, but are small with limited scope. Recent approaches circumvent the data issue through automated synthetic data generation. For example, 3D scenes can be inferred from pre-recorded human motions to get plausible paired scene-motion data . However, motions from these methods are limited to available pre-recorded data. Our data generation pipeline requires only a small set of interaction anchor poses and generates novel motions not contained in prior datasets using tree-based rollouts from a pre-trained generative model .

Method

In this section, we detail our NIFTY pipeline for learning to synthesize realistic human-object interaction motions. §3.1 introduces a conditional diffusion model to generate human motions given the geometry of an object. §3.2 details the object-centric interaction field, which guides the denoising process of the diffusion model to capture the nuances of interactions in a data-driven way. In §4, we discuss the synthetic data generation using a pre-trained motion model that is seeded with anchor poses from a smaller dataset. This data is used to train the diffusion model and interaction field.

Motion Representation. Motion generation is formulated as predicting a sequence of 3D human pose states that capture a person’s motion over time. The pose state representation is based on the SMPL body model and is similar to prior successful human motion diffusion models . The human pose state XiX_{i} at frame ii in a motion sequence is:

Model Formulation. The diffusion model simultaneously generates all human poses in a motion sequence to achieve a desired interaction. Intuitively, diffusion is a noising process that converts clean data into noise. We want our motion model to learn the reverse of this process so that realistic motions can be generated from randomly sampled noise. Mathematically, forward diffusion is a Markov process with a transition probability distribution:

where the diffusion step kk is also given as input. Instead of predicting the noise ϵk\bm{{\epsilon}}^{k} added at each step of the diffusion process , our model directly predicts the final clean signal . Mathematically, the motion model MθM_{\theta} with parameters θ\theta predicts a clean trajectory τ^0=Mθ(τk,k,C)\hat{\bm{\tau}}^{0}=M_{\theta}(\bm{\tau}^{k},k,C) from which the mean μθ(τk,k,C)\mu_{\theta}(\tau^{k},k,C) is easily computed . This formulation has the benefit that physically grounded objectives can be easily computed on τ^0\hat{\bm{\tau}}^{0} in the pose space, which is useful for guidance as discussed below.

While training the diffusion model, a ground truth clean trajectory τ0\bm{\tau}^{0} is noised and given as input, then the model is trained to minimize the objective ∥τ^0−τ0∥22\lVert\hat{\bm{\tau}}^{0}-\bm{\tau}^{0}\rVert^{2}_{2}. To enable using classifier-free guidance at sampling time, the conditioning CC is randomly masked out with 10% probability during the training process so that the model can operate in both conditional and unconditional modes.

Sampling and Guidance. At test time, samples are generated from the model given random noise and interaction conditioning CC as input. We find that leveraging classifier-free guidance tends to generate higher-quality samples. This amounts to generating one conditional and one unconditional sample from the model and then combining them as τ^0=Mθ(τk,k)+s(Mθ(τk,k,C)−Mθ(τk,k))\hat{\bm{\tau}}^{0}=M_{\theta}(\bm{\tau}^{k},k)+s(M_{\theta}(\bm{\tau}^{k},k,C)-M_{\theta}(\bm{\tau}^{k},k)), where the strength of the conditioning is controlled by the scalar ss.

Ensuring that the sampled motions adhere to the geometric and semantic constraints of the object is key to plausible interactions. Diffusion models are well-suited for this, since guidance can encourage samples to meet desired objectives at test time . The core of guidance is a differentiable function G(τ0)G(\bm{\tau}^{0}) that evaluates how well a trajectory meets a desired objective; this could be a learned or an analytic function. In our case, we want G(τ0)G(\bm{\tau}^{0}) to evaluate how plausible an interaction motion is w.r.t. the object, and in §3.2 we show that this can be done with a learned object interaction field. Throughout denoising during sampling, the gradient of the objective function will be used to nudge trajectory samples in the correct direction. We use a formulation of guidance that perturbs the clean trajectory output from the model τ^0\hat{\bm{\tau}}^{0} at every denoising step kk as follows :

Architecture. As shown in Fig. 2 (right), the motion model MθM_{\theta} is based on a transformer encoder-only architecture . The model takes as input the current noisy trajectory τk\bm{\tau}^{k}, the denoising step kk, and the conditioning CC. Each human pose in the trajectory is a token, while each conditioning becomes a separate token. Of note, the object point cloud PoP_{o} is encoded with a PointNet , the rigid pose RoR_{o} is encoded with a three-layer MLP, and kk is encoded using a positional embedding . Our noise levels kk vary between 0 to 1000 diffusion steps. The transformer handles variable-length sequence inputs and outputs the clean motion prediction τ^0\hat{\bm{\tau}}^{0}. Full details are available in the supplementary material.

2 Object Interaction Fields

After training on human-object interactions, the diffusion model can generate reasonable motion sequences but fails to fully comply with constraints in the last mile of interaction , even when conditioned on the object. This causes undesirable artifacts such as penetration with the object.

To alleviate this issue, we propose to guide motion samples from the diffusion model (i.e., use Eq. 4) with a learned objective GG that captures realistic interactions for a specific object.

We take inspiration from recent work that uses neural distance fields to learn valid human pose manifolds and robotic grasping manifolds . For our purposes, the field must take in an arbitrary human pose and output how far the query pose is from being a “valid” object interaction pose. We define an interaction pose to be an anchor frame in a motion sequence that captures the core of the interaction, e.g., the moment a person settles in a chair during sitting (as in Fig. 1) or contacts an object before lifting.

Architecture. As shown in Fig. 2 (left), the interaction field architecture is an encoder-only transformer that operates on the input pose as a token. In practice, it also takes in the canonical object point cloud as a conditioning token to allow training a single field for multiple objects.

3 Automatic Synthetic Data Generation

Training the diffusion model requires a large, realistic, and diverse dataset of motions for each human-object interaction we wish to synthesize. Unfortunately, this data exists only for specific interactions and is difficult and expensive to collect from scratch. Therefore, we propose an automated pipeline to generate synthetic interaction motion data. In short, we first select anchor pose frames from an existing small dataset that are indicative of an interaction we want to learn. Our key insight is to use a pre-trained scene-unaware motion model to sample a diverse set of motions that end at a selected anchor pose, and therefore demonstrate the desired interaction. We provide the key details in this section and a full description appears in the supplementary material.

Anchor Pose Selection. We require a small set of anchor poses that capture the core frame of an interaction motion. As described in §3.2, for sitting on a chair this is the sitting pose when the person first becomes settled in the chair (see Fig. 1). In generating motion data, these anchor poses will be the final frame of each synthesized motion sequence. For the experiments in §4, these anchor frames are chosen manually from a small dataset that contains a variety of interactions .

Generating Motions in Reverse. The goal is to generate human motions that end in the chosen anchor poses and reflect realistic object interactions. We leverage HuMoR , which is a conditional motion VAE trained on the AMASS mocap dataset. It generates realistic human motions through autoregressive rollout, but is scene-unaware. To force rollouts from HuMoR to match the final anchor pose, we could use online sampling or latent motion optimization, but these are expensive and not guaranteed to exactly converge. Instead, we re-train HuMoR as a time-reversed motion model that predicts the past instead of the future motion given a current input pose. Starting from a desired interaction anchor pose XNX_{N}, our reverse HuMoR will generate XN−1,XN−2,⋯ ,X1X_{N-1},X_{N-2},\cdots,X_{1} forming a full interaction motion that, by construction, ends in the desired pose.

Tree-Based Rollout & Filtering. To ensure sufficient diversity and realism in motions from HuMoR, we devise a branching rollout strategy that is amenable to frequent filtering and results in a tree of plausible interactions. Starting from the anchor pose, we first sample 3030 frames (1 sec) of motion. Then, multiple branches are instantiated and random rollout continues for another 3030 frames on these branches independently. Continuing in this branching fashion allows growing the motion dataset exponentially while also filtering to ensure branches are sufficiently diverse and do not contain undesireable motions. Filtering involves heuristically pruning branches with motions that collide with the object, float above the ground plane, result in unnatural pose configurations, and become stationary. For the experiments in §4, we rollout to a tree depth of 7 and sample many motion trees starting from each anchor pose. Individual paths are extracted from the tree to give interaction motions, and we post-filter out sequences that start within 1 meter of the object.

Generated Datasets. We use this scalable strategy to generate data for training our motion model for sitting and lifting interactions. Fig. 4 demonstrates the diversity of our generated datasets by visualizing top-down trajectories and example motions from a single tree of sitting motions. For the sitting interaction dataset, we choose 174 anchor pose frames across 7 subjects in the BEHAVE dataset. This results in a dataset of 200K motion sequences that include sitting on chairs, stools, tables, and exercise balls. Each motion sequence in this dataset ends at a sitting anchor pose. For lifting interactions, 72 anchor poses from 7 subjects produces 110K motion sequences. Each sequence ends at a lifting anchor pose when the person initially contacts the object.

Experiments

We evaluate our NIFTY method after training on the sitting and lifting datasets introduced in §4. Implementation details are given in §4.1, followed by a discussion of evaluation metrics in §4.2 and baselines in §4.3. Experimental results are presented in §4.4 along with an ablation study in §4.5.

We train our diffusion model MθM_{\theta} for 600K iterations with a batch size of 32 using the AdamW optimizer with a learning rate of 10−410^{-4}. A separate model is trained for sitting and lifting. We use K=1000K{=}1000 diffusion steps in our model and sample the diffusion step kk from a uniform distribution at each training iteration. The object interaction field FϕF_{\phi} is trained on the data described in §3.2 for 300K iterations using AdamW with a maximum learning rate of 5×10−55\times 10^{-5} and a one cycle LR schedule . When sampling from the diffusion model, 10 samples are generated in parallel and all are guided using the object interaction field; the sample with the best guidance objective score is used as the output. We apply interaction field guidance on the last frame of motion (i.e., the interaction anchor pose in our datasets). Our models are trained using PyTorch on NVIDIA A40 GPUs, and take about 2 days to train. Visualizations use the PyRender engine .

2 Evaluation Setting and Metrics

To ensure we properly evaluate the generalization capability of methods trained on our synthetic interaction datasets, we do not create a test set using the procedure described in Fig. 4, which may result in a very similar distribution to training data. Instead, we create a set of 500 test scenes for each action, where objects are randomly placed in the scene and the human starts from a random pose generated by HuMoR. All methods are tested on these same scenes during evaluation.

Evaluating human motion coupled with object interactions is challenging and has no standardized protocol. Hence, we evaluate using a diverse set of metrics including a user perceptual study. We briefly describe the metrics next and include full details in the supplementary material.

User Study. No single metric can capture all the nuances of human-object interactions, so we employ a perceptual study . For each method, we create videos from generated motions on the test scenes. To compare two methods, users are presented with two videos on the same test scene and must choose which they prefer (full user directions are in the supplement). We perform independent user studies for lifting and sitting actions using hive.ai . Responses are collected from 5 users for every comparison video, giving 2500 total responses in each comparison study.

Penetration Score (% Pen). To evaluate realism as the human approaches an object for interaction, we measure how much they penetrate the object. Based on our synthetic data, we define the first NAN_{A} frames of motion to be the approach for each action type (see supplement).

Skeleton Distance & Contact IoU. These evaluate how well generated interaction poses align with ground truth poses and their human-object contacts. We start by finding the minimum distance between the final pose of a generated sequence and the anchor poses in the synthetic training data. The distance to this nearest neighbor pose is reported as the skeleton distance. To measure how well contacts from the generated motion match the data, we compute the IoU between contacting vertices (those that penetrate the object) on the predicted body mesh and those on the nearest neighbor mesh.

3 Baselines

Cond. VAE . Closest to our problem definition, this model comes from recent work HUMANISE , which learns plausible human motions conditioned on scene and language for four actions (lie, sit, stand, walk). This state-of-the-art model is a conditional VAE with a GRU motion encoder and sequence-level transformer decoder. Since we evaluate on sitting and lifting actions separately, we modify their approach to remove language conditioning. The model is trained on our synthetic data for 600K iterations with the recommended hyperparameters and learning rate of 10−410^{-4}.

Cond. MDM . This baseline is the motion diffusion model (MDM) with added conditioning CC, i.e., our object-conditioned diffusion model without interaction field guidance.

4 Experimental Results

User Study. Fig. 5 shows how often users prefer our method (NIFTY) over baselines and Synthetic Data (Syn. Data) for both sitting and lifting. We perform separate studies for each comparison. Users prefer NIFTY over baselines a vast majority of the time. Averaged over both actions, NIFTY is preferred over the Cond. VAE baseline 89.4% of the time. Similarly, NIFTY is preferred over Cond. MDM 86.3% of the time, highlighting the importance of using guidance with our interaction field during sampling. Compared to held out motions from synthetic data, NIFTY is preferred 47.2%47.2\% of the time, which indicates that the motions are nearly indistinguishable from those of our data generation pipeline.

We further extract robust consensus across users by majority vote over the 5 responses for each video. In this case, motions generated by our method are preferred 94.4% (sitting) and 97.8% (lifting) of the time over Cond. VAE motions, making the improvement gap even more apparent.

Additionally, we also conduct an user study on a Likert scale of scores between 1 (unrealistic) to 5 (very realistic). We report that motions from our synthetic dataset achieve a score of 4.39 vs. 4.87 for motions in the AMASS . Further details are available in Supp. § A.1.

Quantiative Results. In Tab. 1, NIFTY is compared to baselines for both sitting and lifting interactions. NIFTY generates motions that reach the target object and approach realistically, as indicated by distance-to-object (D2O) and penetration metrics. Although Cond. MDM produces realistic motion with low foot skating, it struggles to properly approach the object since it does not use guidance from the learned interaction field. We see that interaction poses and the resulting object contacts generated by our method do reflect the synthetic dataset, resulting in low skeleton distance and high contact IoU, unlike Cond. VAE which is worse across all metrics.

Qualitative Results. Fig. 6 shows a qualitative comparison between motions generated by our method and baselines. NIFTY synthesizes realistic sitting and lifting with a variety of objects. Examples show that the baselines struggle to generalize to unseen object poses, and have no mechanism to correct for this at test time. Our learned interactions field helps to avoid this through diffusion guidance. Please see the videos provided in the supplement to best appreciate the results.

5 Ablation Study

Conclusion and Limitations

We introduced NIFTY, a framework for learning to synthesize realistic human motions involving 3D object interactions. Results demonstrate that our object-conditioned diffusion model gives improved motions over prior work when guided by a learned object interaction field and trained on automatically synthesized motion data. Our current approach is limited to the body shapes present in the training data (e.g., the 7 subjects from BEHAVE ), so future work should explore data augmentation strategies to generalize to novel humans. Moreover, we have shown results on sitting and lifting, but we would like to widen the scope to handle additional interactions by collecting new anchor poses, synthesizing data, and training our diffusion model and interaction field.

Acknowledgements.. We express our gratitude to our colleagues for the fantastic project discussions and feedback provided at different stages. We have organized them by institution (in alphabetical order) – Google: Matthew Brown, Frank Dellaert, Thomas A. Funkhouser, Varun Jampani – Google (co-interns): Songyou Peng, Mikaela Uy, Guandao Yang, Xiaoshuai Zhang – University of Michigan: Mohamed El Banani, Ang Cao, Karan Desai, Richard Higgins, Sarah Jabbour, Linyi Jin, Jeongsoo Park, Chris Rockwell, Dandan Shan

This work was partly done when NK was interning at Google Research. DR was supported by the NVIDIA Graduate Fellowship.

References

Appendix A Automated Synthetic Training Data Generation

All models in the paper train on synthetic human-object interaction motion data generated using this pipeline. To evaluate the quality of generated data compared to other data, in § A.1 we perform a large scale user-study with 10K user responses. In § A.2 we describe the complete details of data generation including pseudo-code for the algorithm.

Our synthetic data generation pipeline helps us collect high-quality motion data corresponding to different interaction anchor poses. We show that this generated data is high-quality by conducting a user study on a five-point Likert scale as in prior work . Our results show that the generated synthetic training data is on par with data collected using a real mocap setup.

User Study Setup. We created a user-study dataset of 20002000 videos, consisting of 500500 motions from the AMASS subset of HUMANISE sitting data (i.e., real-world motion captured data), 500500 motions from our data generation pipeline, 500500 predicted motions from our NIFTY sitting model, and 500500 from Cond. VAE predictions. For each motion, we rendered a video without an object present in the scene to make the source of the video indistinguishable. All motions had a random number of frames uniformly sampled from 6060 to 120120, where the last motion frame always corresponded to the sitting interaction pose. We only show results on sitting as the HUMNISE does not have lifting interaction AMASS subset in their data.

We ask the users to rate the video on its realism. Users are asked to rate on a scale of 11 to 55 corresponding to “Strongly Disagree", “Disagree", “Neutral", “Agree", “Strongly Agree". We set up the study on hive.ai . Instructions to the user are shown in Fig. 7.

User Study Results. Results are shown in Fig. 8. As expected, AMASS has a high realism score of 4.87 since it is actual mocap data. Training data generated using our algorithm has an average user rating of 4.39, implying the quality is comparable motion collected using an expensive mocap setup. We also report the performance of NIFTY and Cond. VAE methods on the same study for completeness. NIFTY achieves a strong score of 4.11 (between “agree” and “strongly agree”), which is close to score of the Syn. Data. The Cond. VAE performance remains low at 2.33 (between “disagree” and “neutral”).

Filtering Unreliable Users. Note that every user is required to pass a qualification test containing easy examples to label. User accuracy is computed and users with accuracy > 60% are admitted. To ensure that we collect valid responses and that users completely understand the task during the actual study, they are occasionally tested on “obvious" data called “honey pots" during the labeling process. To this end, we add motions with objective “Strongly Agree" labels (motions from AMASS) and some with "Strongly Disagree" labels (low-quality motions generated by cVAE). This is common practice while conducting such a study, and we also do this for the user study in the main paper as detailed in §C.1. The honeypot accuracy for this task is set at 82%: drops in performance below this thresholds removes a user from continuing the study any further.

A.2 Training Data Synthesis Algorithm

Our generation process revolves around utilizing a pretrained motion model, specifically the HuMoR generative model , to produce motion trajectories that end in a specific anchor pose. However, we train this model on reverse-time sequences, enabling us to generate reverse-time sequences that start from the provided anchor seed pose. Then, when we convert these rollouts into forward motions (i.e., play them backwards), the final generated pose in the rollout aligns with the anchor pose by design.

Our full algorithm for generating a single motion tree is shown in Algorithm 1. This algorithm constructs a tree of a specified depth, where each node corresponds to a 1 sec motion clip. Each node is connected to several possible branches to continue the motion (based on a branching factor BB). The algorithm begins by creating a root node starting at an input anchor pose. It then repeatedly constructs the tree by generating motion sequences using the RollOut function and checking their validity using the PruneCheck function. If a valid motion sequence is obtained, a child node is created and added to the tree. The process continues until the desired depth is reached or the tree is fully explored (no more branches left to explore)

The algorithm maintains a queue of nodes to be processed, allowing for breadth-first construction of the tree. If a node reaches the maximum depth, it is skipped to ensure the tree is constructed as per the specified depth. The algorithm outputs the resulting tree, which contains valid motion sequences as paths from the root to the leaf nodes.

RollOut Function. The RollOut function takes an start pose and utilizes the pre-trained motion model to generate a short 1 sec (30 frame) motion sequence. It iteratively runs the motion model until a valid sequence is obtained or a specified maximum number of attempts is reached. If a valid sequence is found, it is returned as the generated motion.

PruneCheck Function. The PruneCheck function examines a given motion sequence to determine its validity. It algorithmically checks if the motion collides with the object, has unnatural human poses, if the human is floating in the air, or intersecting with the floor etc.. It returns a boolean value indicating whether the motion sequence is valid or not.

Implementation. In our implementation, we set BB as 66 for the nodes at depths 11 and 22, while B=2B=2 for nodes at higher depths. We also set NTries as 2020 to secure a good rollout sequence. We then convert all the motion nodes in these trees into individual motion sequences for a particular interaction.

Appendix B Implementation Details

Recovering Motion from τ0\bm{\tau}^{0}. Our trajectory representation is over-parameterized and this allows using the model outputs in multiple ways. To recover the generated motion we extract the per-frame joint angles jir{\bm{j}}^{r}_{i} for the SMPL model. We integrate the velocity tiv{\bm{t}}_{i}^{v} along the XZ plane to recover the XZ translation for the root joint and extract the corresponding Y component (upward) from tip{\bm{t}}_{i}^{p}. This strategy of extracting motion from the output parameterization is motivated by our use of guidance with the diffusion model, which only operates on the last frame of a motion sequence. By integrating velocity predictions over time, applying the guidance objective at the last frame will still have a strong effect on earlier frames in the sequence.

Variable Length Input. Our model takes input motion trajectories with up to 150 frames. For training, we pad motion sequences of lengths shorter than this with the last interaction frame from the sequence.

SMPL model. Our SMPL model does not have hand articulation, so we use the SMPL model with only 2222 articulated joints.

Pre-trained Motion Model for Data Generation. We train the motion model on a subset of the AMASS dataset that does not contain extreme sporting actions like jumping, dancing, etc.We do this by removing sequences from AMASS based on the labels from the BABEL dataset . We use the HuMoR-Qual variant of the model to get high-quality motions, which uses the joint positions computed through the SMPL parametric model as input to future roll-out time steps (as opposed to using its own joint position predictions).

Transfomer Encoder.. We use a transformer encoder implemented using torch.nn.TransformerEncoder from PyTorch . Our each transformer layer consists of 4 heads and a latent dim on 512512. We have 8 such layers in our transformer.

Appendix C Experimental Details

This section provides additional details on the implementation of our user study and metrics from the main paper in §4.

We conduct a user study to qualitatively evaluate the performance of two methods. We design a study such that, given a pair of motions, a user must choose one that is the most realistic. Specifically, we ask the user “Which motion among the both is more realistic?" when we show them two videos (each containing a motion generated by a different method) “LEFT VIDEO" & “RIGHT VIDEO". Fig. 9 shows the instructions and user interface from the study. We conduct 3 such studies using hive.ai , the results of which are in Fig 5 of the main paper.

Filtering Unreliable Users. We require users to understand instructions given in English. User selection for the study is conditioned on the performance of a qualification test. Users with an accuracy of ≥\geq 80% on this test are allowed to take the study. To ensure continued reliability during the labeling process we randomly mix the real task data with “obvious" honeypot data where the labels are objective. We require users to have a performance of ≥\geq 89% on these honeypot tasks. A drop in performance below this results in the user being disqualified from taking the study further.

C.2 Metrics

Apart from performing the user study described in §C.1 we also evaluate all our models and baselines on several quantitative metrics. We detail these metrics below (apart from the details already described in Sec 4.2 of the main paper).

Penetration Score. To assess the realism of human motion when interacting with an object, we calculate the penetration score during the approach phase. We define the approach phase as the initial NAN_{A} motion frames from a sequence of 150 frames (5 sec). Our rationale for selecting NAN_{A} is that during the approach phase, there should be minimal penetration of the human motion into the object geometry. However, during the interaction, there should be increasing contact with the object. These contacts typically result in zero or positive values in the signed distance function (SDF), indicating penetration of points on the object surface into the human SMPL mesh.

We compute NAN_{A} for sitting and lifting separately based on our synthetic dataset. In particular, we determine the first frame index of motion where object penetration distance continues to only increase thereafter. We assume that after this point, the person is actually interacting with the object and not just approaching it. For sitting, the typical onset of motion interaction occurs after the initial 117 frames of approach, based on the median NAN_{A}. Likewise, lifting has a 15th percentile NAN_{A} of 124 frames. We use the 15th percentile instead of the median (148 frames) to make this metric more meaningful as 148 frames is almost the end of the complete motion and we wish to evaluate the approach. This difference between sit and lift action is due to the difference in their inherent interaction with the object.

For completeness, we also report this performance as a function of different NAN_{A} values in Fig. 10 (sit) and Fig. 11 (lift).

Skeleton Distance. This metric uses the anchor poses from our human-object interaction data to evaluate whether generated motions faithfully reflect interactions from data. We compute a sum over the per-joint location error (2222 joints in our case) between the final generated interaction pose and the nearest neighbor anchor pose from the training dataset in the joint locations space. We report the average of this metric across generated motions.

Appendix D Supplemental Results

In this section, we include supplemental analyses to support the evaluations in the main paper that were not included due to space constraints. First, we evaluate the effect of having a parametric vs a non-parametric guidance field in §D.1. In §D.2, D.3, and D.4 we evaluate the impact of hyperparameters like the number of samples at inference, number of anchor poses at training, and a variant of our Object Interaction Field that guides a motion sequence instead of just the final interaction frame. We also evaluate the difference in performance across different objects.

We conducted a comparison between our method and a variant where we replaced the object interaction field with a non-parametric field implemented using the nearest neighbor measure. Specifically, during the guidance phase, we identified the nearest anchor pose of the object from the training set and used the difference between this pose and the predicted final pose as the correction. This correction was then utilized to define our distance field and guide the diffusion model accordingly.

Tab. 3 presents the comparison between this baseline and our method. The skeleton distance metric can be sensitive to outliers (e.g., a few generations that are far from the object), so we additionally report % Skel. Dist. ≤25cm\leq 25cm to get a more robust metric. The results demonstrate that our learning approach offers a significant improvement of at least 20% in terms of Skeleton Distance ≤25\leq 25 cm, as well as an additional 10% in terms of Contact IoU. The main paper reports results on the Parametric approach as our primary model.

D.2 Effect of Number of Samples

In the main paper, we generate 10 guided samples from the diffusion model and use the one with the best guidance score. We investigate the impact of varying these number of samples in Tab. 4. We observe that increasing the number of samples leads to improved performance. Particular improvements occur when transitioning from 1 sample to 5 samples. Since guidance does not always result in perfect samples, drawing a diverse set gives better chance for a high-quality output. Note that drawing additional samples can be done efficiently in parallel.

D.3 Effect of Number of Anchors Poses

We also train our Interaction Field (IF) using subsets of motion that yield a limited number of anchor poses. Specifically, we train the IF using 10%, 25%, and 50% of the available seed anchor poses and report results in Tab. 5. It is worth noting that Contact IoU and Skeleton Dist metrics are calculated using all anchor poses in the training set. However, methods trained with only XX% of the anchor data will not be able to generate the complete range of seed poses. Therefore, when comparing methods trained with different percentages of seed anchor poses, we primarily assess them based on other metrics, but Contact IoU and Skeleton Dist are still included for completeness.

NIFTY’s performance remains stable even with the limited availability of anchor poses. Looking at Foot Skating, D2O, and Penetration metrics, there is not a significant decline in performance. The main paper reports results on 100% data for NIFTY.

D.4 Effect of Number of Input Frames on Interaction Field

D.5 Effect of training Interaction Field in the Local Human Frame

Our interaction field is object-centric since it takes in a canonical object point cloud as input. To test this design choice, we implement the object interaction field in the local frame of the human requiring it to understand the spatial positioning of the object w.r.t to the human motion. As shown in Tab. 7, this leads to a subpar performance across the board on sit and lift actions.

D.6 Performance Breakdown Per-Object

We analyze if the performance of our method is biased towards certain objects by computing the metrics for about 100 interaction motion samples per object instance. We show the results of this in Tab. 8. Results indicate that the performance of our method is not dependent on the kind of the object. For instance, in the case of sitting, the performance for sitting on a “Armchair" vs “Chair" are close. This demonstrates the flexibility of the NIFTY pipeline to a diverse set of objects.

Appendix E Qualitative Results

Motion generation results are best seen as videos on the attached webpage. We also include static visualizations here in Fig. 12 and Fig. 13. The webpage additionally also shows visualizations ( 10 motions) from our method for every object in our dataset.

Appendix F Limitations

Our proposed pipeline demonstrates the ability to achieve human-object interaction results with a diverse sets of objects while only relying on a limited number of anchor poses. One of the key factors contributing to the performance of NIFTY is the utilization of a pretrained motion model trained on the AMASS repository . Our data generation pipeline has the capability to generate motions and interpolate between existing data in this dataset. However, in cases where a completely novel and extreme seed anchor pose is provided, such as a headstand, HuMoR would struggle to generate reasonable and high-quality motion sequences. Developing more robust motion models which can handle such poses, would be beneficial.

Furthermore, during the inference stage, it is necessary to draw multiple samples from the diffusion model and guide them. This approach yields significantly better performance compared to guiding only a single sample. Exploring research directions that can enhance the stability of the guidance process would be valuable in consistently generating high-quality interaction motions.