GOAL: Generating 4D Whole-Body Motion for Hand-Object Grasping

Omid Taheri, Vasileios Choutas, Michael J. Black, Dimitrios Tzionas

Introduction

Virtual humans are important for movies, games, AR/VR and the metaverse. Not only do they need to look realistic, but also move and interact realistically. Most work on human motion generation has focused only on bodies, without the head and hands. Often, these bodies are considered in “isolation”, with no scene or object context. Other work focuses on bodies interacting with scenes, but ignores the hands. Similarly, work on generating hand grasps often ignores the body. We argue that these are all just parts of the problem. What we really need, instead, is to generate motion of full-body avatars grasping objects, by jointly considering the body, head, hands, and the object. We address this here for the first time.

The problem is challenging and multifaceted. Think of how we grasp objects in real life (see Fig. 2); we walk towards the object with our feet contacting the floor, we orient our head to look at the object, lean our torso and extend our arms to reach it, and dexterously pose our hands to establish fine contact and grasp it. Humans are able to gracefully execute these steps, yet, these are challenging and involve motion planning, motor control, and spatial awareness. Some of these steps have been studied separately, but we cannot simply combine the partial solutions since the entire action must be coordinated. This is challenging because: (1) full bodies have a much higher-dimensional state space than bodies or hands alone; (2) the body and hands have very different sizes, motion scales and level of dexterity; (3) the body, head, and hands must move in a coordinated fashion. Currently, there are no automatic tools to generate such coordinated full-body grasping motions.

We address this with GOAL, which stands for Generating Object-interActing whoLe-Body motions. GOAL generates whole-body avatar motion for grasping an unknown object, by jointly considering the body, head, hands, and the object. GOAL takes three inputs: (1) a 3D object, (2) its position and orientation, and (3) a “starting” 3D body pose and shape, positioned near the object and roughly oriented towards it. As output, GOAL generates a sequence of 3D body poses from the starting pose through to an object grasp. To do so, GOAL uses two novel networks (for an overview see Fig. 3): (1) First, GNet generates a “goal” whole-body grasp, with a realistic body pose, head pose, arm pose, and hand pose, as well as realistic finger-object and foot-ground contact. GNet is formulated as a conditional variational auto-encoder (cVAE), thus, it learns a distribution over grasping poses, and can generate a variety of “goal” grasps. (2) Then, MNet inpaints the motion between the “starting” and “goal” poses, by generating a sequence of whole-body poses in an auto-regressive fashion. This is challenging because the avatar needs to (see Fig. 1) walk by taking a number of steps proportional to the distance to the object, while having natural foot-floor contact without “skating”, and continuously orient the head to look at the object. Then, when it is near the object, it needs to slow down, stop walking, lean the torso, extend the arms to reach the object. It must also pose the hand to contact the object and grasp it. All body parts need to move gracefully and in full coordination, so that the motion looks natural.

Achieving this level of realism requires technical novelties. GOAL draws inspiration by recent work , but goes beyond this to uniquely infer both SMPL-X parameters and 3D offsets. GNet infers 3D hand-to-object vertex offsets to give spatial awareness and guide object grasping. MNet infers 3D SMPL-X vertex offsets to guide SMPL-X deformation from the previous to the current frame. These offsets lie in 3D Euclidean space, thus, they can be more accurately inferred than SMPL-X parameters, and are used in an offline optimization scheme to refine SMPL-X poses. We train GNet and MNet on the GRAB dataset, which contains whole-body SMPL-X humans grasping objects.

We evaluate GOAL, both quantitatively and qualitatively, on withheld parts of the GRAB dataset. Specifically, we withhold 5 objects for testing. Results show that GOAL generalizes well and produces natural motions for full-body walking and object grasping; see Fig. 1. Quantitative evaluation shows that GOAL outperforms baselines, and ablation studies show a positive contribution of all major components. A perceptual study, verifies the above, while showing that GOAL’s generated motions achieve a level of realism comparable to GRAB’s ground-truth motions.

To conclude, GOAL takes a step towards automatic whole-body grasp motion generation for realistic avatars. Models and code will be available for research purposes.

Related Work

Motion generation for bodies “in isolation”: Research on human motion generation has a long history . However, even recent methods , mostly study the body “in isolation”; i.e., with no scene context. Most methods generate the motion of 3D skeletons , while others generate the motion of a human model like SMPL . Typically, 11-22 seconds of motion synthesis is referred to as “long term”. Early deep-learning methods employ RNNs , however, they struggle with discontinuities between the observed and predicted poses, and with long-range spatial relations across time. Other methods account for these with phase-functioned feed-forward neural networks , i.e. by conditioning the network weights on phase. However, these focus on cyclic motions. More recent methods adopt an attention mechanism.

Motion generation for bodies in 3D scenes: Most early methods extend MoCap databases with point annotations for foot and hand contact . Then, they fit motion to contacts with optimization and space-time constraints for 3D body motion re-targeting , and animating bodies that move in 3D terrains .

To avoid big MoCap datasets, some methods use deep reinforcement learning (RL) for body-scene or hand-object interactions. These methods show promising results for navigating terrains with varying height and gaps , sitting on chairs , using a hammer and opening a door , and for in-hand object re-orientation . Generalization to new bodies, object geometry, and interaction types remains a challenge.

Others follow a 3D geometric approach. Pirk et al. place virtual sensors on objects to sense the flow of points sampled on an agent interacting with these, and build functional object descriptors. Al-Asqhar et al. re-target body motion by encoding human joints w.r.t. fixed points sampled on a scene. Ho et al. use body and object vertices to compute per-frame “interaction meshes”, and minimize their Laplacian deformation to re-target body motion. These pure geometric methods are not robust to real-world noise.

In contrast, we fall in the category of data-driven methods. Corona et al. generate the context-aware motion of a human skeleton interacting with objects, where “context” is encoded as a directed graph connecting person and object nodes. More relevant are methods for generating motion between a “start” and a “goal” pose in a 3D scene. Hassan et al. estimate a “goal” position and interaction direction on an object, plan a 3D path from a start body pose to this, and finally generate a sequence of body poses with an auto-regressive cVAE for walking and interacting, e.g., sitting on a chair. Wang et al. first estimate several “sub-goal” positions and bodies, divide these into short start/end pairs to synthesize short-term motions, and finally stitch these together in a long motion with an optimization process.

Motion generation for hands: ElKoura1 et al. estimate physically plausible hand poses for playing musical instruments, using a low dimensional pose space, with a data-driven approach. Pollard et al. use MoCap to learn a controller for physically-based grasping. Kry et al. capture hand MoCap and forces with sensors on objects, and use these to build “interaction trajectories”, and synthesize and re-target motions with physics simulation. More related to us, Lie et al. take as input MoCap data of body and object motion, and add the missing hand motion to the body, by first searching for feasible contact point trajectories, and then generating smooth hand motion with space-time optimization that satisfies the estimated contacts.

Pose generation for bodies in 3D scenes: Early methods use either contact annotations or detections on 3D objects, and fit 3D skeletons to these. Other methods use physics simulation to reason about contacts and sitting comfort . Focusing on rooms instead of single objects, Grabner et al. predict all areas on a 3D scene mesh where a 3D human mesh can sit, using proximity and intersection metrics. Recent methods use deep learning to generate static humans interacting with a scene. Zhang et al. learn a cVAE to generate SMPL-X poses, conditioned on an input depth image and semantic segmentation of the scene. Zhang et al. use an explicit scene-centric representation of interaction, while Hassan et al. use a human-centric representation.

Pose generation for hand-object grasps: Taheri et al. predict MANO hand grasps for unseen 3D object meshes, by first predicting a rough hand grasp, and then refining it with distance and contact metrics. Grady et al. refine grasps by first estimating contacts on both the hand and the object, and then refining the hand with optimization to satisfy the inferred contacts.

Motion for full-body interactions: People use their body and hands together for interacting with the world. Hsiao et al. build a database of whole-body grasps with a human operating an avatar, and perform imitation learning. Borras et al. capture whole-body MoCap data of people interacting with scene objects and handheld objects, using a humanoid model, and define a pose taxonomy. Taheri et al. capture whole-body SMPL-X interactions with handheld objects, but learn a cVAE that generates only static grasping hands, due to the task complexity. Merel et al. use deep RL and human MoCap demonstrations to learn a vision-guided neural controller for picking up and carrying boxes, or catching/throwing a ball.

Summary: The community has focused on parts of the problem (either the body or the hands) or used unrealistic bodies. GOAL learns to generate full-body SMPL-X motions, from walking to approach an object up to grasping it, given only a 3D object and a starting human pose.

Method

An overview of our method, GOAL, is shown in Fig. 3. GOAL takes three inputs, namely: (1) a 3D object, (2) its position and orientation, and (3) a “starting” 3D body pose and shape, positioned near the object (roughly 0.5−1.50.5-1.5 m) and oriented towards it (roughly ±10∘\pm 10^{\circ}). Then, as output, GOAL generates SMPL-X motion with two main networks: (1) GNet synthesizes a “goal” SMPL-X mesh that grasps the 3D object with a realistic body pose and hand-object contact; (2) MNet “inpaints” the motion from the starting to the “goal” frame, by generating a sequence of “moving” SMPL-X bodies in an auto-regressive way. Without loss of generality, we model right-handed grasps.

2 Interaction-Aware Attention

Two common representations for body-object interaction are vertex-to-vertex distances between meshes and contact maps on meshes. However, the former carries information that is irrelevant to the interaction (e.g., vertices far away from the object), while the latter is too compact and carries no information about 3D proximity before/after contact.

Here, we use vertex-to-vertex distances, but introduce a novel “interaction-aware” attention that focuses more on body vertices that are important for interaction (e.g., hands for grasping, feet for walking) and less to irrelevant vertices (e.g., knees are less relevant than the hand for grasping). Our “interaction-aware” attention is formulated as:

3 “Goal” Network (GNet)

GNet is a conditional variational auto-encoder (cVAE) that generates a whole-body grasp, conditioned on the given object and its location. To do this, we first encode whole-body grasps into an embedding space.

Finally let db→o=d(v,vo)\bm{d}^{b\rightarrow o}=\bm{d}(\bm{v},\bm{v}^{\mathtt{o}}), i.e. the offset vectors from the sampled body vertices, v\bm{v}, to the closest object vertices, vo\bm{v}^{\mathtt{o}}.

Both the encoder and decoder use fully-connected layers with skip connections, GNet is trained end-to-end and the loss is defined as L=λvLv+\mathcal{L}=\lambda_{\bm{v}}\mathcal{L}_{\bm{v}}+

where Lv=∥v−v^∥1\mathcal{L}_{\bm{v}}=\lVert\bm{v}-\hat{\bm{v}}\rVert_{1}, Lvh=∥vh−v^h∥1\mathcal{L}_{\bm{v}}^{h}=\lVert\bm{v}^{h}-\hat{\bm{v}}^{h}\rVert_{1}, Lp=∥Θ−Θ^∥2\mathcal{L}_{p}=\lVert\bm{\Theta}-\hat{\bm{\Theta}}\rVert_{2}, Lh=∥h−h^∥2\mathcal{L}_{\mathtt{h}}=\lVert\bm{h}-\hat{\bm{h}}\rVert_{2}, Ldrh=∥dh−d^h∥1\mathcal{L}_{\bm{d}}^{rh}=\lVert\bm{d}^{h}-\hat{\bm{d}}^{h}\rVert_{1}, and LKL\mathcal{L}_{KL} denotes the Kullback-Leibler divergence. The hat denotes regressed quantities; the non-hat variables are ground truth. For the exact architecture of GNet, see Sup. Mat.

We make two empirical observations: (1) Networks struggle to predict accurate SMPL-X parameters, possibly due to their non-Euclidean space. (2) Networks predict interaction features in a Euclidean space much more precisely. These observations are in line with recent work , but we go beyond them in regressing 3D offsets together with SMPL-X parameters, instead of regressing point positions and fitting SMPL-X to these. We leverage offsets in an optimization step to refine our SMPL-X predictions.

GNet Optimization: We leverage the predicted offsets to refine our SMPL-X predictions with optimization post processing. Specifically, we optimize over SMPL-X pose, θ\bm{\theta}, and translation, t\bm{t}, initialized with GNet’s predictions. Instead of hand-crafted contact constraints during optimization, we use data-driven constraints generated from GNet. Specifically, we use: (1) hand-to-object vertex offsets, (2) head-orientation and (3) pose coupling to the initial value, and (4) foot-ground penetration.

In technical terms, for refining the hands to realistically grasp the objects, we define a l1l_{1} term between the offsets d^h\hat{\bm{d}}^{h} generated from the GNet, and offsets computed online from SMPL-X’s hand vertices to the closest object vertices:

Pose and translation coupling discourages deviations from the initial ones:

Similarly, head-orientation coupling is formulated as:

Finally, we find the lowest vertex of the body along the y-axis (vertical axis) and enforce its y-coordinate to be zero to have contact and prevent penetration using:

Our final energy is a combination of the above five terms:

The efficacy of our optimization post processing using the predicted Euclidean-space interaction features, is evaluated in the next section with a perceptual study (Tab. 2).

4 Motion Network (MNet)

MNet generates the motion from the starting to the “goal” frame; the latter is generated by GNet above. The length of a sequence depends on several factors, like the object location w.r.t. the body and the speed of motion. Therefore, to generate motion of arbitrary length, we use an auto-regressive network architecture .

Input: MNet takes as input (auto-regressive fashion):

where Θt−5:t\bm{\Theta}_{t-5:t} are SMPL-X parameters of the last 55 frames, β\bm{\beta} is the subject’s shape, vt\bm{v}_{t} and v˙t\dot{\bm{v}}_{t} are the locations and velocities of the sampled body vertices in the current frame, dthd_{t}^{h} are the hand vertex offsets from the current to the “goal” pose, and bghb_{g}^{h} is the BPS representation of the hand in the “goal” grasping frame. For this, we use the same BPS points as for the object. The BPS representation encodes the spatial relationship between the hand and the object in the “goal” frame, and is empirically important for “guiding” the motion towards a grasp with a good hand pose and hand-object contact. For our auto-regressive scheme, in agreement with , we empirically find that using more than 11 past frame leads to a smoother motion prediction; more than 55 frames do not lead to noticeable improvement.

Similar to GNet, along with predicting SMPL-X model parameters Θ∈SE(3)\bm{\Theta}\in SE(3), we also predict interaction features that lie in the Euclidean space. Empirically, this improves network inference, i.e. the generated motion is smoother and better “reaches” the “goal” grasp. Unlike , where each motion frame depends only on 11 past frame, we find that MNet’s generated motion quality improves as the number of future frames it generates grows; see Tab. 3.

where t+10t+10 denotes the future 1010 motion frames, Δθt+10,Δtt+10\Delta\theta_{t+10},\Delta{t}_{t+10}, denote the change of SMPL-X pose and translation parameters, Δvt+10\Delta{v}_{t+10} is the change of SMPL-X vertex locations, and Δdt+10h\Delta{d}^{h}_{t+10} is the change of hand vertex offsets. All changes, Δ\Delta, are relative to the current frame. In an auto-regressive fashion, MNet estimates SMPL-X parameters for “motion” poses, and then these are fed back to MNet as inputs (along with other ones) for the next iteration. For the exact architecture of MNet, see Sup. Mat.

MNet is trained end-to-end, with a loss similar to GNet. Specifically, we use a loss term on hand-to-object offsets, body parameters, body and hand vertices, similar to Ldh\mathcal{L}_{\bm{d}}^{h}, Lp\mathcal{L}_{p}, Lv\mathcal{L}_{\bm{v}}, Lvh\mathcal{L}_{\bm{v}}^{h} in Eq. 4 respectively.

One common limitation of motion generation methods is “skating”, i.e. foot sliding on the ground. To account for this, we define an additional loss term on foot vertices, when these are close to the ground. This loss, along with the computed input velocities for Eq. 10, result in more realistic foot-ground contact; see video in Sup. Mat.

MNet Optimization: We refine MNet’s generated motion with post processing based on optimization; this refines the motion for better “reaching” the “goal” grasping pose generated by GNet. Since we need precision only when the hand is very close to the object, we apply the optimization step only when MNet’s estimated hand vertices get closer than 1010 cm to the “goal” hand vertex positions.

We follow GNet’s scheme, and use MNet’s predictions (Eq. 11) as constraints, instead of hand-crafted ones. We first compute the average value of MNet’s predicted hand-vertex velocities, vt˙h\dot{v_{t}}^{h}. Then, we linearly interpolate between the “goal”, vghv_{{g}}^{h}, and “current”, vthv_{t}^{h}, hand vertices:

where ∥vt˙h∥\lVert\dot{v_{t}}^{h}\rVert is the average-velocity magnitude, and l^\hat{l} is the (unit) vector pointing from “current” to the “goal” hand vertices. In practice, we “force” hands to move towards the “goal” grasp in a (locally) linear trajectory. Since our focus here is the hand grasp, for the rest of the body we keep the pose and velocity that MNet predicts.

The optimization objective function LL uses loss terms on hand vertices, Lvh\mathcal{L}_{\bm{v}}^{h}, and on SMPL-X pose parameters, Lp\mathcal{L}_{p}, similar to the ones described for Eq. 4, and has the form:

5 Implementation Details

Optimization details: For both GNet’s and MNet’s optimization-based post processing, we perform gradient descent with Adam to optimize SMPL-X parameters.

Training data: For training both GNet and MNet, we use the GRAB dataset , which contains whole-body 3D SMPL-X humans grasping 3D objects. Please refer to Sup. Mat. for the details of data preparation.

Experiments

We show examples of GNet’s generated grasp before and after optimization in Fig. 5. Results show that GNet generates plausible body pose and head orientation for static grasps, but the hand grasps have room for improvement. The optimization step refines hand grasps so that they are more realistic and physically plausible. We show several representative motions generated by MNet with different objects, locations as well as various body shapes in Fig. 6. For more results, please see our video and Sup. Mat.

2 Quantitative Experiments

Perceptual Study: To quantitatively evaluate the generated results from GNet and MNet, we perform a perceptual study through Amazon Mechanic Turk (AMT).

GNet: For each test-set object, we use GNet to generate 22 “goal” whole-body grasps. We render a “turntable animation” of the generated grasps, before and after optimization, as well as the corresponding ground-truth grasps. Participants are asked to rate the quality of 44 features: (1) grasping pose, (2) foot-ground contact, (3) hand-object grasp, and (4) head orientation. They rate the realism of each feature using a Likert scale of scores between 11 (unrealistic) to 55 (very realistic). Each grasp is evaluated by at least 1010 participants. To remove invalid ratings, e.g., participants that do not understand the task, we use catch trials similar to GRAB . The results of the evaluation are reported in Tab. 1. The study shows the effectiveness of the optimization step, especially on making the hand grasps more realistic.

The study shows that the optimized grasps have a better quality in grasping pose and head orientation compared to the ground truth. This is because in a subset of the GRAB dataset the subject looks away while grasping the object, but in GNet results the head is always oriented towards the object. The higher rating in the feet-ground penetration is due to the direct loss term in our optimization process which results in a better feet-ground contact. Overall, the quality of the generated grasps are close to the ground truth.

MNet: We use MNet to generate grasping motions on the test set. In a perceptual study we show participants generated sequences and ground-truth ones, and ask them to rate: (1) the overall body motion quality, (2) foot-ground contact and sliding, (3) hand-object grasp at the end of the motion, (4) and head orientation. Table 2 shows that GOAL generates realistic grasping motions, that approach the realism of ground truth. Note that MNet has a harder task than GNet, as it generates a full motion instead of a static pose. By comparing Tabs. 2 and 1, note that ground truth is rated higher for motions than for static poses and this is harder for MNet to match, though scores are not much lower.

Foot-Sliding Metric: We evaluate the physical plausibility of the generated motion using a “foot-sliding” metric. For each sequence, both generated and ground-truth ones, we find the closest vertex of the body to the ground and measure its velocity. We consider a frame to contain a “sliding” foot if the change in location of the selected foot vertex is higher than 1cm1cm per frame. The percentage of “foot-sliding” frames in the ground truth and in the GOAL-generated sequences is 6.7%6.7\%, and 13.7%13.7\% respectively. Although there is room for improvement, empirically GNet’s motions have less sliding than existing work (on other data).

3 Ablation Study

Number of Output Frames: Here we study the effect of the number of output motion frames in MNet. To study this, we trained 55 networks with different numbers of output frames ranging from 11 to 1010. We report the comparison in terms of the pose errors of the SMPL-X model, vertex-to-vertex distance for the body, feet, and hands between the generated results and the ground-truth. Results in Tab. 3 show that generating more frames in each iteration of the auto-regressive network helps generate more accurate results. Additionally, we have observed qualitatively that, for networks with a lower number of frames as output, sometimes the motion does not converge to a final grasp, and the hands gradually deviate from the object.

Conclusion and Future Work

We introduce GOAL, the first model to generate realistic human motions to grasp previously unseen 3D objects. We use two novel networks (GNet and MNet) to first generate a static “goal” grasp and then inpaint the motion between the frames. We exploit the ability of both networks to infer interaction features in Euclidean space and introduce an optimization step after each network to improve the quality of the grasps and motion based on the regressed features. The evaluation shows that our framework is able to synthesize natural and physically plausible grasping motions.

GOAL opens up many possibilities for future studies on grasping motion generation. Even though GOAL generates realistic grasping motions, it is constrained to be in a close distance to the object and can not generate motions when the body is far from the object. Future work should extend this to synthesize longer walking motions, prior to interaction with objects. In addition, in this work we focus on human-object interaction; in future work we would like to combine GOAL with human-scene interaction models to generate scene-aware grasping motions.

Social Impact: While realistic motion generation has mostly positive use cases in VR/AR, games, and movies, with the recent advances in neural rendering and deepfakes, we see a possibility that our results could be used for full-body deepfakes. Being aware of this, we will make our models available only for research purposes.

Acknowledgements: This research was partially supported by the International Max Planck Research School for Intelligent Systems (IMPRS-IS), the Max Planck ETH Center for Learning Systems (CLS), and the German Federal Ministry of Education and Research (BMBF): Tubingen AI Center, FKZ: 01IS18039B. We thank Tsvetelina Alexiadis for the Mechanical Turk experiments, and Taylor McConnell for the voice recordings. Disclosure: MJB has received research gift funds from Adobe, Intel, Nvidia, Facebook, and Amazon. While MJB is a part-time employee of Amazon, his research was performed solely at, and funded solely by, Max Planck. MJB has financial interests in Amazon, Datagen Technologies, and Meshcapade GmbH.

References

Data Preparation

GNet data preparation: GNet generates static grasps. Therefore, from the GRAB dataset, we collect all frames with right-hand grasps, for which participants grasp the object in a stable way. For this, we follow the selection criteria used for GrabNet’s training data. We then center the object at the origin along the horizontal plane, i.e., while preserving its height. In total, we collect 160K160K, 26K26K, and 12.5K12.5K frames for the training, testing, and validation set, respectively.

MNet data preparation: MNet generates motion. Therefore, from each sequence of GRAB, we gather all frames from the starting one up to the frame where the right hand first establishes a stable grasp. For this, we use the same selection criteria as above for GNet. We then create several sub-sequences by sliding a 2121-frame long window over each sequence with a stride of 11 frame. For each sub-sequence, we consider the first 1010 frames as “past” motion, the last 1010 frames as “future” motion, and the middle one as the “current” frame. Then, following , we make all “past” and “future” frames relative to the body coordinate system of the “current” frame, while keeping the gravity direction always upward. In total, we collect roughly 40K40K, 7K7K, and 3K3K motion sub-sequences for the training, testing, and validation sets, respectively.

GNet Architecture

For an architectural overview of GNet and its optimization-based post processing, see R.7.

MNet Architecture

For an architectural overview of MNet and its optimization-based post processing, see R.8.

Video

We provide a narrated video: (1) explains our motivation, (2) explains out method, and (3) shows many results, including qualitative motion results.