Controllable Human-Object Interaction Synthesis
Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu, Xavier Puig, C. Karen Liu
Introduction
Synthesizing human behaviors in 3D environments is critical for various applications in computer graphics, embodied AI, and robotics. Humans effortlessly navigate and engage within their surroundings, performing a plethora of tasks routinely. For example, drawing a chair closer to a desk to create a workspace, adjusting a floor lamp to cast the perfect glow, or neatly storing a suitcase. Each of these tasks requires precise coordination between the human, the object, and the surroundings. These tasks are also deeply rooted in purpose. Language serves as a powerful tool to articulate and convey these intentions. Synthesizing realistic human and object motion guided by language and scene context is the cornerstone of building an advanced AI systems that simulate continuous human behaviors in diverse 3D environments.
While some existing works study the problem of human-scene interaction , they are constrained to scenarios with static objects such as sitting on a chair, neglecting the highly dynamic interactions happening frequently in daily life. Recent advancements have been made in modeling dynamic human-object interactions, yet these approaches focus solely on smaller objects or lack the ability to manipulate diverse objects . Manipulating diverse objects of larger size has been explored in recent work . However, these approaches rely on sequences of past interaction states or complete sequences of object motion, and are incapable of synthesizing both object motion and human motion from initial states alone. In this work, we aim to advance the field by focusing on synthesizing realistic interactions involving diverse objects of larger size from language and initial states.
Generating continuous human-object interactions from language descriptions within 3D scenes poses several challenges. First, we need to generate object and human motion which is realistic and synchronized. The human hands should maintain appropriate contact with objects during interaction and object motion should maintain a causal relationship to human actions. Second, 3D scenes are often cluttered with numerous objects, constraining the space of feasible motion trajectories. Thus, it is essential for interaction synthesis to accommodate for environment clutter, rather than operating under the assumption of an empty scene.
In this work, we focus on the key problem of synthesizing human-object interactions in 3D environments from natural language commands, generating object motion and human motion guided by language and sparse object waypoints. Starting with a language description outlining the desired human actions, a set of waypoints extracted from the environment, and an initial object and human state, our goal is to generate motions for both humans and objects. These motions should align with the directives specified in the language input, while also conforming to the environmental constraints defined by waypoint conditions derived from 3D scene geometry.
To achieve this, we employ a conditional diffusion model to generate synchronized object and human motion simultaneously, conditioned on language descriptions, initial states, and sparse object waypoints. To improve the accuracy of the predicted object motion, we incorporate an object geometry loss during training. In addition, we devise guidance terms applied during the sampling process to improve the realism of the generated interaction. Furthermore, we demonstrate the effectiveness of our learned interaction synthesis module within a system that produces continuous realistic and context-aware interactions given language descriptions and 3D scenes.
To summarize, our work makes the following contributions. First, we identify that the combination of language and object waypoints provides precise and expressive information for human-object interaction synthesis. We show that object waypoints do not need to be dense or precise, which allows us to utilize existing path planning algorithms to generate sparse waypoints which represent long-horizon interactions in complex scenarios. Second, based on this finding, we devise a method that synthesizes human-object interaction guided by language and sparse waypoints of the object, using a conditional diffusion model. We demonstrate that our approach synthesizes realistic interactions on FullBodyManipulation dataset . In addition, the learned model can generalize to novel objects in 3D-FUTURE dataset . Third, we integrate our method into a pipeline that synthesizes long-horizon environment-aware human-object interactions from 3D scenes and language input.
Related Work
With the development of large-scale high-quality motion capture datasets like AMASS , there has been a growing interest in generative human motion modeling. BABEL and HumanML3D further introduce action labels and language descriptions to enrich the mocap dataset, enabling the development of action-conditioned motion synthesis and text-conditioned motion synthesis . Prior work has shown that VAE formulation is effective in generating diverse human motion from text . Recently, with the success of the diffusion model in this domain , extensive work has explored generating motion from text using conditioning . In this work, we also take language descriptions as input to guide our generation. Instead of synthesizing human motion alone, we generate both object motion and human motion conditioned on the text.
Motion Synthesis in 3D Scenes.
With the advent of paired scene-motion data and paired object-motion data , approaches have been developed to generate human interactions such as sitting on a chair and reaching a target position in 3D scenes. To populate human-object interactions without training on paired scene-motion data, path planning algorithms have been deployed to generate collision-free paths which then guide the human motion generation . Another line of work leverages reinforcement learning frameworks to train scene-aware policies for synthesizing navigation and interaction motions in static 3D scenes . In this work, instead of focusing on static scenes or objects, we synthesize interactions with dynamic objects. Also, inspired by approaches that decompose scene-aware motion generation into path planning and goal-guided generation phases, we design an interaction synthesis module conditioned on sparse object waypoints that can be effectively integrated into a scene-aware synthesis pipeline.
Interaction Synthesis.
The field of modeling dynamic human-object interactions has largely focused on hand motion synthesis . Recently, with the advent of full-body motion datasets with hand-object interactions , models have been developed to synthesize full-body motions preceding object grasping. Some recent studies predict object motion based on human movements , and others have taken this further by synthesizing both body and hand motion, subsequently applying optimization to predict object motion. However, these approaches focus on smaller objects where hand motion is the primary focus. In terms of manipulating larger objects, some methods train reinforcement learning policies to synthesize box lifting and moving behaviors , yet these models struggle to generalize to manipulation of diverse objects. Based on paired human-object motion data , recent works predict interactions from a sequence of past interaction states or an object motion sequence , incapable of synthesizing interactions in 3D scenes solely from initial states. In this work, we generate synchronized object and human motion conditioned on sparse object waypoints, serving to ground the resulting trajectories in 3D scenes.
Method
Our goal is to generate synchronized object and human motion, conditioned on a language description, object geometry, initial object and human states, and sparse object waypoints. Two primary challenges arise in this context: first, modeling the complexity of synchronized object and human motion while also respecting the sparse condition signals; and second, ensuring the realism of contact between the human and object. To tackle the generation problem of complex interactions, we employ a conditional diffusion model to generate object motion and human motion at the same time. However, naively learning a conditional diffusion model to generate both object motion and human motion cannot ensure the precise contact between hand and object and the realism of the interaction. Thus, we incorporate several constraints as guidance during the sampling process of our trained diffusion model. We illustrate our approach in Figure 2.
Object Geometry Representation.
Input Condition Representation.
2 Interaction Synthesis Model
The conditional signals of our model, denoted as , include initial states, sparse object waypoints, the object BPS representation, and language descriptions. The diffusion model consists of a forward diffusion process that progressively adds noise to the clean data and a reverse diffusion process which is trained to reverse this process. The forward diffusion process introduces noise for steps formulated using a Markov chain,
where represents a fixed variance schedule and is an identity matrix. Our goal is to learn a model to reverse the forward diffusion process,
where denotes the predicted mean and is a fixed variance. Learning the mean can be re-parameterized as learning to predict the clean data representation . The objective is defined as
Model Architecture.
We employ a transformer architecture as our denoising network. Our input consists of object geometry conditions , masked motion conditions , and noisy data representation at noise level . The input is projected to a sequence of feature vectors using a linear layer. We employ an MLP to embed the noise level . Then we combine the noise level embedding and the language embedding to form a single embedding vector denoted as . The embedding vector has the same dimension as these feature vectors and is fed to the transformer along with these vectors. The final prediction is made by projecting the updated feature vectors of the transformer excluding the time step corresponding to the embedding . The interaction synthesis model is illustrated in Figure 2.
Object Geometry Loss.
At each time step in our model, the predicted object rotation (converted to relative rotation with respect to the object geometry in rest pose) and position are employed to calculate the corresponding positions of these selected vertices. This is represented by the following equation, where and denote the predicted rotation and translation of the object, and refers to the ground truth vertices at time step . The object geometry loss is computed as
This loss function plays a critical role in guiding the model to accurately predict the transformation of the object.
3 Guidance
During the training phase of our interaction synthesis model, there are no explicit contact constraints enforced in the losses. Incorporating loss terms such as hand-object contact loss, and object-floor penetration loss poses a challenge for training. First, these types of loss terms are computationally expensive and would slow down training significantly. Second, introducing more loss terms requires meticulously balancing different losses which usually necessitates re-training models with different settings. Instead, enforcing these constraints during test time is more flexible and makes it easier to select appropriate weights for different terms. Thus, to refine our generated interactions, we propose the application of guidance during the sampling process.
In this work, we leverage reconstruction guidance in the sampling process as we empirically found it to be more stable. We define multiple analytical functions as guidance terms which we will introduce in the following sections.
We have implemented a specialized contact guidance function to improve the hand-object contact accuracy for frames generated by our model. This function is specifically designed to address cases where a noticeable distance exists between the hands and the object, thereby improving the realism and precision of the interaction. The contact guidance function is defined as follows:
Feet-Floor Contact Guidance.
When generating joint positions and rotations, our model operates without awareness of the body’s shape. Consequently, using the SMPL-X model with predicted root positions, joint rotations, and a test subject’s specific body shape parameters to reconstruct the human mesh can sometimes lead to scenarios where the feet do not touch the floor. To rectify this, we implement a guidance term that encourages realistic feet-floor contact.
The joint positions of the left and right toes are represented as and , respectively. We identify the supporting foot in each frame by comparing the z components of these two joints at each frame. We also introduce a threshold height meters, which is determined from the analysis of foot height in the ground truth motion. The guidance term is defined as follows:
This function computes the norm of the vertical difference between the lowest point of either toe and the threshold height .
Object-Floor Penetration Guidance.
To address the issue of generated object states potentially penetrating the floor, we integrate an additional guidance function into the sampling process. Given that our floor is positioned at the plane where , we define the guidance term as follows:
where represents the z-coordinate of the object vertices.
During the inference phase, we apply multiple guidance concurrently defined as follows,
where , , denote the loss weights for different terms. We apply the guidance in the last 10 denoising steps only since the prediction in the early steps is extremely noisy.
Experiments
We first introduce the datasets and evaluation metrics. Then we show comparisons of our proposed approach against the baselines. We further conduct a human perceptual study to complement our evaluation and ablation study to verify the effectiveness of our proposed guidance terms. Moreover, we demonstrate an application that generates long-term interactions conditioned on object waypoints extracted from 3D scenes.
The FullBodyManipulation dataset consists of 10 hours of high-quality, paired object and human motion, including interaction with 15 different objects. However, our study does not encompass the generation of motion for articulated objects, leading us to exclude sequences related to two such objects (vacuum and mop). We employ this dataset both for training our interaction model and for evaluating the generated results. The training set comprises 15 subjects, with an additional 2 subjects designated for testing, adhering to the dataset partitioning used in OMOMO .
The 3D-FUTURE dataset includes 3D models of various furniture items. From this dataset, we select 17 objects representing diverse types (such as chairs, tables, floor lamps, and boxes). This dataset serves to test our model’s ability to generalize to objects it has not previously encountered. Given that the 3D-FUTURE dataset only includes 3D models, we integrate object position data from the testing set of the FullBodyManipulation dataset for evaluation.
2 Evaluation Metrics
Condition Matching Metric: This metric calculates the Euclidean distance between the predicted and input object waypoints. It includes the start object position error , end object position error , and waypoint errors , all measured in centimeters (cm).
Human Motion Quality Metric: This metric encompasses the foot sliding score (FS) and foot heights . FS is the weighted average of accumulated translation in the xy plane, following prior work , measured in centimeters (cm). assesses the height of the feet, also in centimeters.
Interaction Quality Metric: This metric assesses the accuracy of hand-object interactions, encompassing both contacts and penetrations. For contact accuracy, it employs precision , recall , and F1 score metrics following prior work . Additionally, it includes contact percentage , determined by the proportion of frames where contact is detected. To compute the penetration score , each vertex of the hand is used to query the precomputed object’s Signed Distance Field (SDF). This process yields a corresponding distance value for each vertex. The penetration score is then derived by computing the average of the negative distance values (representing penetration), formalized as , measured in centimeters (cm).
Ground Truth (GT) Difference Metric: This metric measures the deviation of generated results from the ground truth motion. It comprises the mean per-joint position error (MPJPE), translation error of the root joint , and object position error , all computed using the Euclidean distance between the predicted and actual ground truth positions in centimeters (cm). Additionally, this metric includes the root joint orientation error and the object orientation error . These errors are calculated with the Frobenius norm of the rotational difference, formulated as where and represent the predicted and ground truth rotation matrices respectively.
3 Results
As there is no prior work presenting a solution for our task, we adapt the most related work, OMOMO , to our problem setting in order to evaluate against a baseline. This method was designed for synthesizing human motion from a provided object motion trajectory. Since OMOMO requires a sequence of object states to generate full-body human poses, we implement a linear interpolation strategy for the object positions. This interpolation is based on the given start and end positions of the object, as well as predefined waypoints in the xy-plane. We also maintain a consistent object rotation, using the orientation from the initial frame throughout the entire sequence. Additionally, we evaluate our approach CHOIS against two ablations: CHOIS w/o and CHOIS w/o . CHOIS w/o is trained as a conditional diffusion model but does not include an additional object geometry loss. This variant allows us to understand the baseline performance of the diffusion model in a straightforward setup. In contrast, CHOIS w/o incorporates the object geometry loss in its training process but operates without guidance during inference. This approach lets us explore the effectiveness of object geometry loss during training while assessing the model’s capability in the absence of guidance.
Results on the FullBodyManipulation Dataset.
We evaluate our approach using objects from the FullBodyManipulation dataset as shown in Table 1. Introducing object geometry loss notably improves the condition matching metric. Furthermore, adding guidance during inference leads to better contact accuracy, reduced hand-object penetration, and less foot floating. Note that OMOMO has zero deviation from the input object trajectory since it only predicts human motion and does not change the object motion input. We also showcase qualitative comparisons in Figure 3. Note that OMOMO’s object motion is updated via linear interpolation and thus cannot follow the text prompt (row 2, lifting the table above the head).
Results on the 3D-FUTURE Dataset.
To test our model’s ability to generalize to new objects, we conduct evaluations using the 3D-FUTURE dataset . As shown in Table 2, our proposed method outperforms the baseline and the two ablations. We also provide qualitative results in Figure 3.
Human Perceptual Study.
We conduct two human perceptual studies to further complement the evaluation of our approach. The first study assesses the consistency between the generated interactions and the text input. The second study evaluates the overall quality of these generated interactions. For each of these studies, we generate 100 sequences using each method, including our CHOIS, the baseline, our ablations, and the ground truth. This results in a set of 400 pairs. We employ Amazon Mechanical Turk (AMT) for evaluation. Each sequence pair is reviewed by 10 different AMT workers. The results are illustrated in Figure 4. Given that OMOMO generates human motions based solely on interpolated object states and does not incorporate language conditions, our method demonstrates superior performance in aligning with text input. Moreover, our approach shows improvements over the CHOIS w/o , with both CHOIS w/o and CHOIS exhibiting comparable proficiency in text alignment. Regarding interaction quality, our method also surpasses all baseline and model ablations.
4 Ablation Study
We conduct an ablation study to validate the effectiveness of our proposed guidance terms. As shown in Table 3, our hand-object contact guidance and feet-floor contact guidance are both critical. Without the hand-object contact guidance, the contact percentage degrades obviously. Without the feet-floor contact guidance, the height of the feet increases indicating there exists severe foot floating issues. We are not ablating object-floor penetration guidance as object-floor penetration issues are not common and this term is primarily designed for preventing penetration artifacts in qualitative results.
5 Application
This section presents a practical application of our method, enabling the synthesis of human-object interactions within 3D scenes, driven by language descriptions. We utilize 3D scenes from the Replica Dataset .
The process begins by composing language descriptions that specify the desired interactions, identifying both the objects involved and their intended positions. For example, the language description can be “pull the floor lamp to be close to a shelf”. We also define a set of primitive functions used to sample target 3D positions from 3D scenes. This set includes functions like sampling points on an object’s surface or near it. GPT-3 is used to extract key information including the interaction object and target objects, and to select the appropriate primitive functions from our predefined function set. Combining the information with the semantic labels of the scene point cloud, we can determine the target 3D positions.
We leverage Habitat to generate collision-free paths within the scene given the start and target object positions. However, as Habitat provides waypoints without corresponding time steps, we need to adapt these to our learned module. We apply heuristics to create waypoints at fixed intervals of 30 frames, which serve as the input conditions for our interaction synthesis model. An example of this application is shown in Figure 5, demonstrating how our learned interaction synthesis model effectively synthesizes human-object motion following a description in a 3D scene. Table 4 includes a quantitative evaluation of the generated motion. In addition, we showcase the results using the same text input but different waypoints in Figure 6, demonstrating the effectiveness of the control using object waypoints.
Conclusion
In conclusion, our work addresses the problem of human-object interaction synthesis conditioned on language descriptions and sparse object waypoints. By employing a conditional diffusion model, we successfully generate object and human motions that are not only synchronized but also resonate with given language descriptions. We incorporate object geometry loss during training which significantly improves the performance of object motion generation. We also propose effective guidance terms used during the sampling process which enhance the realism of the generated results. Moreover, we demonstrate that our learned interaction module can be integrated into a pipeline that synthesizes long-term interactions given language and 3D scenes.
This work is in part supported by the Stanford Institute for Human-Centered AI (HAI), NSF CCRI #2120095, ONR MURI N00014-22-1-2740, and Meta. Part of the research was done during Jiaman Li’s internship at FAIR, Meta.