GRAB: A Dataset of Whole-Body Human Grasping of Objects

Omid Taheri, Nima Ghorbani, Michael J. Black, Dimitrios Tzionas

Introduction

A key goal of computer vision is to estimate human-object interactions from video to help understand human behavior. Doing so requires a strong model of such interactions and learning this model requires data. However, capturing such data is not simple. Grasping involves both gross and subtle motions, as humans involve their whole body and dexterous finger motion to manipulate objects. Therefore, objects contact multiple body parts and not just the hands. This is difficult to capture with images because the regions of contact are occluded. Pressure sensors or other physical instrumentation, however, are also not a full solution as they can impair natural human-object interaction and do not capture full-body motion. Consequently, there are no existing datasets of complex human-object interaction that contain full-body motion, 3D body shape, and detailed body-object contact. To fill this gap, we capture a novel dataset of full-body 3D humans dynamically interacting with 3D objects as illustrated in Fig. 1. By accurately tracking 3D body and object shape, we reason about contact resulting in a dataset with detail and richness beyond existing grasping datasets.

Most previous work focuses on prehensile “grasps” ; i.e. a single human hand stably lifting or using an object. The hands, however, are only part of the story. For example, as infants, our earliest grasps involve bringing objects to the mouth . Consider the example of drinking from a bowl in Fig. 1 (right). To do so, we must pose our body so that we can reach it, we orient our head to see it, we move our arm and hand to stably lift it, and then we bring it to our mouth, making contact with the lips, and finally we tilt the head to drink. As this and other examples in the figure illustrate, human grasping and using of everyday objects involves the whole body. Such interactions are fundamentally three-dimensional, and contact occurs between objects and multiple body parts.

Dataset. Such whole-body grasping has received much less attention than single hand-object grasping . To model such grasping we need a dataset of humans interacting with varied objects, capturing the full 3D surface of both the body and objects. To solve this problem we adapt recent motion capture techniques, to construct a new rich dataset called GRAB for “GRasping Actions with Bodies.” Specifically, we adapt MoSh++ in two ways. First, MoSh++ estimates the 3D shape and motion of the body and hands from MoCap markers; here we extend this to include facial motion. For increased accuracy we first capture a 3D scan of each subject and fit the SMPL-X body model to it. Then MoSh++ is used to recover the pose of the body, hands and face. Note that the face is important because it is involved in many interactions; see in Fig. 1 (second from left) how the mouth opens to eat a banana. Second, we also accurately capture the motion of 3D objects as they are manipulated by the subjects. To this end, we use small hemi-spherical markers on the objects and show that these do not impact grasping behavior. As a result, we obtain detailed 3D meshes for both the object and the human (with a full body, articulated fingers and face) moving over time while in interaction, as shown in Fig. 1. Using these meshes we then infer the body-object contact (red regions in Fig. 1). Unlike Brahmbhatt et al. this gives both the contact and the full body/hand pose over time. Interaction is dynamic, including in-hand manipulation and re-grasping. GRAB captures 1010 different people (55 male and 55 female) interacting with 5151 objects from . Interaction takes place in 44 different contexts: lifting, handing over, passing from one hand to the other, and using, depending on the affordances and functionality of the object.

Applications. GRAB supports multiple uses of interest to the community. First, we show how GRAB can be used to gain insights into hand-object contact in everyday scenarios. Second, there is a significant interest in training models to grasp 3D objects . Thus, we use GRAB to train a pair of neural networks (coarse prediction followed by refinement) to generate plausible grasps for unseen 3D objects. Given a randomly posed 3D object, we predict plausible hand parameters ( 66 DoF wrist pose and full finger articulation) appropriate for grasping the object. To encode arbitrary 3D object shapes, we employ the recent basis point set (BPS) representation , whose fixed size is appropriate for neural networks. Then, by conditioning on a new 3D object shape, we sample from the learned latent space, and generate hand grasps for this. We quantitatively and qualitatively evaluate the resulting grasps and show that they look natural.

In summary, this work makes the following contributions: (1) we introduce a unique dataset capturing real “whole-body grasps” of 3D objects, including full-body human motion, object motion, in-hand manipulation and re-grasps; (2) to capture this, we adapt MoSh++ to solve for the body, face and hands of SMPL-X to obtain detailed moving 3D meshes; (3) using these meshes and tracked 3D objects we compute plausible contact on the object and the human and provide an analysis of observed patterns; (4) we show the value of our dataset for machine learning, by training a novel conditional neural network to generate 3D hand grasps for unseen 3D objects. The dataset, models, and code are available for research purposes at https://grab.is.tue.mpg.de.

Related Work

Hand Grasps: Hands are crucial for grasping and manipulating objects. For this reason, many studies focus on understanding grasps and defining taxonomies , by exploring the object shape and purpose of grasps , the contact areas on the hand captured by sinking objects in ink , the pose and contact areas captured with an integrated data-glove and tactile-glove , or the number of fingers in contact with the object and thumb position . A key element for these studies is capturing accurate hand poses, relative hand-object configurations and contact areas.

Whole-Body Grasps: Often people use more than a single hand to interact with objects. However, there are not many works in the literature on this topic . Borras et al. use MoCap data of people interacting with a scene with multi-contact, and present a body pose taxonomy for such whole-body grasps. Hsiao et al. focus on imitation learning with a database of whole-body grasp demonstrations with a human teleoperating a simulated robot. Although these works go in the right direction, they use unrealistic humanoid models and simple objects or synthetic ones . Instead, we use the SMPL-X model to capture “whole-body”, face and dexterous in-hand interactions.

Capturing Interactions with MoCap: MoCap is often used to capture, synthesize or evaluate humans interacting with scenes. Lee et al. capture a 3D body skeleton interacting with a 3D scene and show how to synthesize new motions in new scenes. Wang et al. capture a 3D body skeleton interacting with large geometric objects. Han et al. present a method for automatic labeling of hand markers, to speed up hand tracking for VR. Le et al. capture a hand interacting with a phone to study the “comfortable areas”, while Feit et al. capture two hands interacting with a keyboard to study typing patterns. Other works focus on graphics applications. Kry et al. capture a hand interacting with a 3D shape primitive, instrumented with a force sensor. Pollard et al. capture the motion of a hand to learn a controller for physically based grasping. Mandery et al. sit between the above works, capturing humans interacting with both big and handheld objects, but without articulated faces and fingers. None of the previous work captures full 3D bodies, hands and faces together with 3D object manipulation and contact.

Capturing Contact: Capturing human-object contact is hard, because the human and object heavily occlude each other. One approach is instrumentation with touch and pressure sensors, but this might bias natural grasps. Pham et al. predefine contact points on objects to place force transducers. Recent advances in tactile sensors allow accurate recognition of tactile patterns and handheld objects . Some approaches use a data glove with an embedded tactile glove , but this combination is complicated and the two modalities can be hard to synchronize. A microscopic-domain tactile sensor is introduced in , but is not easy to use on human hands. Mascaro et al. attach a minimally invasive camera to detect changes in the coloration of fingernails. Brahmbhatt et al. use a thermal camera to directly observe the “thermal print” of a hand on the grasped object. However, for this they only capture static grasps that last long enough for heat transfer. Consequently, even recent datasets that capture realistic hand-object or body-scene interaction avoid directly measuring contact.

3D Interaction Models: Learning a model of human-object interactions is useful for graphics and robotics to help avatars or robots interact with their surroundings, and for vision to help reconstruct interactions from ambiguous data. However, this is a chicken-and-egg problem; to capture or synthesize data to learn a model, one needs such a model in the first place. For this reason, the community has long used hand-crafted approaches that exploit contact and physics, for body-scene , body-object , or hand-object scenarios. These approaches compute contact approximately. This approximation may be rough when humans are modeled as 3D skeletons or shape primitives . However, it gets relatively accurate when using 3D meshes that are generic , personalized , or based on 3D statistical models .

To collect training data, several works use synthetic Poser hand models, manually articulated to grasp 3D shape primitives. Contact points and forces are also annotated through proximity and inter-penetration of 3D meshes. In contrast, Hasson et al. use the robotics method GraspIt to automatically generate 3D MANO grasps for ShapeNet objects and render synthetic images of the hand-object interaction. However, GraspIt optimizes for hand-crafted grasp metrics that do not necessarily reflect the distribution of human grasps (see Sup. Mat. Sec. C.2 of , and ). Alternatively, Garcia-Hernando et al. use magnetic sensors to reconstruct a 3D hand skeleton and rigid object poses; they capture 66 subjects interacting with 44 objects. This data is used for 3D hand/object pose estimation or motion generation , but suffers from noisy poses and severe inter-penetrations (see Sec. 5.2 of ).

For bodies, Kim et al. use synthetic data to learn to detect contact points on a 3D object, and then fit an interacting 3D body skeleton to them. Savva et al. use RGB-D to capture 3D body skeletons of 55 subjects interacting in 3030 3D scenes, to learn to synthesize interactions , affordance detection , or to reconstruct interaction from videos . Mandery et al. use optical MoCap to capture 4343 subjects interacting with 4141 tracked objects, both large and small. This is similar to our effort but they do not capture fingers or 3D body shape, so cannot reason about contact. Corona et al. use this dataset to learn context-aware body motion prediction. Starke et al. use Xsens IMU sensors to capture the main body of a subject interacting with large objects, and learn to synthesize avatar motion in virtual worlds. Hassan et al. use RGB-D and 3D scene constraints to capture 2020 humans as SMPL-X meshes interacting with 1212 static 3D scenes, but do not capture object manipulation. Zhang et al. use this data to learn to generate 3D scene-aware humans.

We see that only parts of our problem have been studied. We draw inspiration from prior work, in particular . We go beyond these by introducing a new dataset of real “whole-body” grasps, as described in the next section.

Dataset

To manipulate an object, the human needs to approach its 3D surface, and bring their skin to come in physical contact to apply forces. Realistically capturing such human-object interactions, especially with “whole-body grasps”, is a challenging problem. First, the object may occlude the body and vice-versa, resulting in ambiguous observations. Second, for physical interactions it is crucial to reconstruct an accurate and detailed 3D surface for both the human and the object. Additionally, the capture has to work across multiple scales (body, fingers and face) and for objects of varying complexity. We address these challenges with a unique combination of state-of-the-art solutions that we adapt to the problem.

There is a fundamental trade-off with current technology; one has to choose between (a) accurate motion with instrumentation and without natural RGB images, or (b) less accurate motion but with RGB images. Here we take the former approach; for an extensive discussion we refer the reader to Sup. Mat.

We use a Vicon system with 5454 infrared “Vantage 1616” cameras that capture 1616 MP at 120120 fps. The large number of cameras minimizes occlusions and the high frame rate captures temporal details of contact. The high resolution allows to capture even small (1.51.5 mm radius) hemi-spherical markers. This minimizes their influence on finger and face motion and does not alter how people grasp objects. Details of the marker setup are shown in Fig. 2. Even with many cameras, motion capture of the body, face, and hands, together with objects, is uncommon because it is so challenging. MoCap markers become occluded, labels are swapped, and ghost makers appear. MoCap cleaning was done by four trained technicians using Vicon’s Shōgun-Post software.

Capturing Human MoCap: To capture human motion, we use the marker set of Fig. 2 (left). The body markers are attached on a tight body suit with velcro-based mounting at a distance of roughly db=9.5d_{b}=9.5 mm from the body surface. The hand and face markers are attached directly on the skin with special removable glue, therefore the distance to it is roughly dh=df≈0d_{h}=d_{f}\approx 0 mm. Importantly, no hand glove is used and hand markers are placed only on the dorsal side, leaving the palmar side completely uninstrumented, for natural interactions.

Capturing Objects: To reconstruct interactions accurately, it is important to know the precise 3D object surface geometry. We therefore use the CAD object models of , and 3D print them with a Stratasys Fortus 360360mc printer; see Fig. 2 (right). Each object oo is then represented by a known 3D mesh with vertices VoV_{o}. To capture object motion, we attach 1.51.5 mm radius hemi-spherical markers directly on the object surface with strong glue. We use at least 88 markers per object, empirically distributing them on the object so that at least 33 of them are always observed. The size and placement of the markers makes them unobtrusive. In Sup. Mat. we show empirical evidence that makers have minimal influence on grasping.

2 From MoCap Markers to 3D Surfaces

Model-Marker Correspondences: For the human body we define, a priori, the rough marker placement on the body as shown in Fig. 2 (left). Exact marker locations on individual subjects are then computed automatically using MoSh++ . In contrast to the body, the objects have different shapes and mesh topologies. Markers are placed according to the object shape, affordances and expected occlusions during interaction; see Fig. 2 (right). Thus, we annotate object-specific vertex-marker correspondences, and do this once per object.

Human and Object Tracking: To ensure accurate human shape, we capture a 3D scan of each subject and fit SMPL-X to it following . We fit these personalized SMPL-X models to our cleaned 3D marker observations using MoSh++ . Specifically we optimize over pose, θ\theta, expressions, ψ\psi, and translation, γ{\gamma}, while keeping the known shape, β\beta, fixed. The weights of MoSh++ for the finger and face data terms are tuned on a synthetic dataset, as described in Sup. Mat. An analysis of MoSh++ fitting accuracy is also provided in Sup. Mat.

3 Contact Annotation

Since contact cannot be directly observed, we estimate it using 3D proximity between the 3D human and object meshes. In theory, they come in contact when the distance between them is zero. In practice, however, we relax this and define contact when the distance, dd, is below a tolerance, threshold d≤ϵcontactd\leq\epsilon_{contact}. This helps address: (1) measurement and fitting errors, (2) limited mesh resolution, (3) the fact that human soft tissue deforms when grasping an object, while the SMPL-X model cannot model this.

Given these issues, accurately estimating contact is challenging. Consider the hand grasping a wine glass in Fig. 4 (right), where the color rings indicate intersections. Ideally, the glass should be in contact with the thumb, index and middle fingers. “Contact under-shooting” results in fingers hovering close to the object surface, but not on it, like the thumb. “Contact over-shooting”, results in fingers penetrating the object surface around the contact area, like the index (purple intersections) and middle finger (red intersections). The latter case is especially problematic for thin objects where a penetrating finger can pass through the object, intersecting it on two sides. In this example, we want to annotate contact only with the outer surface of the object and not the inner one.

We account for “contact over-shooting” cases with an efficient heuristic. We use a fast method to detect intersections, cluster them in connected “intersection rings”, Rb⊊Vb\mathcal{R}_{b}\subsetneq V_{b} and Ro⊊Vo\mathcal{R}_{o}\subsetneq V_{o}, and label them with the intersecting body part, seen as purple and red rings in Fig. 4 (right). The “intersection ring”, Rb\mathcal{R}_{b}, segments the body mesh, MbM_{b}, to give the “penetrating sub-mesh”, Mb⊊Mb\mathcal{M}_{b}\subsetneq M_{b}. We then identify two cases: (1) When a body part gives only one intersection, we annotate as contact points on the object, VoC⊂VoV_{o}^{\mathcal{C}}\subset V_{o}, all vertices enclosed by the ring Ro\mathcal{R}_{o}. We then annotate as contact points on the body, VbC⊂VbV_{b}^{\mathcal{C}}\subset V_{b}, all vertices that lie close to VoCV_{o}^{\mathcal{C}} with a distance do→b≤ϵcontactd_{o\to b}\leq\epsilon_{contact}. (2) In case of multiple intersections, ii, we take into account only the ring Rbi\mathcal{R}_{b}^{i} corresponding to the largest intersection subset, Mbi\mathcal{M}_{b}^{i}.

For body parts that are not found in contact above, there is the possibility of “contact under-shooting”. To address this, we compute the distance from each object vertex, VoV_{o}, to each non-intersecting body vertex, VbV_{b}. We then annotate as contact vertices, VoCV_{o}^{\mathcal{C}} and VbCV_{b}^{\mathcal{C}}, the ones with do→b≤ϵcontactd_{o\to b}\leq\epsilon_{contact}. We empirically find that \epsilon_{contact}={{\color[rgb]{0,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,0}\pgfsys@color@gray@stroke{0}\pgfsys@color@gray@fill{0}{4.5}}} mm works well for our purposes.

4 Dataset Protocol

Human-object interaction depends on various factors including the human body shape, object shape and affordances, object functionality, or interaction intent, to name a few. We therefore capture 1010 people (55 men and 55 women), of various sizes and nationalities, interacting with the objects of ; see example objects in Fig. 2 (right). All subjects gave informed written consent to share their data for research purposes.

For each object we capture interactions with 44 different intents, namely “use” and “pass” (to someone), borrowed from the protocol of , as well as “lift” and “off-hand pass” (from one hand to the other). Figure 3 shows some example 3D capture sequences for the “use” intent. For each sequence we: (i) we randomize initial object placement to increase motion variance, (ii) we instruct the subject to follow an intent, (iii) the subject starts from a T-pose and approaches the object, (iv) performs the instructed task, and (v) leaves the object and returns to a T-pose. The video on our website shows the richness of our protocol with a wide range of captured sequences.

5 Analysis

Dataset Analysis. The dataset contains 1334 sequences and over 1.6M frames of MoCap; Table 1 provides a detailed breakdown. Here we analyze those frames where we have detected contact between the human and the object. We assume that the object is static on a table and can move only due to grasping. Consequently, we consider contact frames to be those in which the object’s position deviates in the vertical direction by at least 55 mm from its initial position and in which at least 5050 body vertices are in contact with the object. This results in 952,514952,514 contact frames that we analyze below. The exact thresholds of these contact heuristics have little influence on our analysis, see Sup. Mat.

By uniquely capturing the whole body, and not just the hand, interesting interaction patterns arise. By focusing on “use” sequences that highlight the object functionality, we observe that 92%92\% of contact frames involve the right hand, 39%39\% the left hand, 31%31\% both hands, and 8%8\% involve the head. For the first category the per-finger contact likelihood, from thumb to pinky, is 100%100\%, 96%96\%, 92%92\%, 79%79\%, 39%39\% and for the palm 24%24\%. For more results see Sup. Mat.

To visualize the fine-grained contact information, we integrate over time the binary per-frame contact maps, and generate “heatmaps” encoding the contact likelihood across the whole body surface. Figure 5 (left) shows such “heatmaps” for “use” sequences. “Hot” areas (red) denote high likelihood of contact, while “cold” areas (blue) denote low likelihood. We see that both the hands and face are important for using everyday objects, highlighting the importance of capturing the whole interacting body. For the face, the “hot” areas are the lips, the nose, the temporal head area, and the ears. For hands, the fingers are more frequently in contact than the palm, with more contact on the right hand than the left one. The palm seems more important for right-hand grasps than for left-hand ones, possibly because all our subjects are right-handed. Contact patterns are also influenced by the size of the object and the size of the hand; see Sup. Mat. for a visualization.

Figure 6 shows the effect of the intent. Contact for “use” sequences complies with the functionality of the object; e.g. people do not touch the knife blade or the hot area of the pan, but they do contact the on/off button of the flashlight. For “pass” sequences subjects tend to contact one side of the object irrespective of affordances, leaving the other one free to be grasped by the receiving person.

For natural interactions it is important to have a minimally intrusive setup. While our MoCap markers are small and unobtrusive, as seen in Figure 2 (right), we ask whether subjects may be biased in their grasps by these markers. Figure 5 (right) shows contact “heatmaps” for some objects across all intents. These clearly show that markers are often located in “hot” areas, suggesting that subjects do not avoid grasping these locations. Further analysis based on K-means clustering of grasps can be found in Sup. Mat.

GrabNet: Learning to Grab an Object

We show the value of GRAB with a challenging application; we train on it a model that generates plausible 3D MANO grasps for an unseen 3D object. Our model, GrabNet, is comprised of two modules: coarse prediction and refinement. This is similar to , but with several key differences. We predict full human hand pose instead of a robotic gripper, train using captured human grasps, employ a different object representation which generalizes to new objects, and make different choices for the model and losses. Importantly, our refinement is done by a neural network, not an optimization process.

We first employ CoarseNet, a conditional variational autoencoder (cVAE) , that generates an initial grasp. For this it learns a grasping embedding space, ZZ, conditioned on the object shape, that is encoded using the Basis Point Set (BPS) representation as a set of distances from the basis points to the nearest object points. In contrast to , ZZ captures not only the 66 DoF pre-grasp pose (gripper for /wrist for MANO), but also the fully articulated human hand pose. CoarseNet’s grasps are reasonable, but realism can improve by refining contacts based on the distances, DD, between the hand and object meshes. We do this with a second network, called RefineNet, that employees explicit MANO contact points learned on our GRAB dataset to refine the initial pose. In contrast, learn an evaluator of the 66 DoF gripper pose that they differentiate to optimize the coarse pose. The architecture of GrabNet is shown in Fig. 7, for more details see Sup. Mat.

Pre-processing. For training, we gather all frames with right-hand grasps that involve some minimal contact, for details see Sup. Mat. We then center each training sample, i.e. hand-object grasp, at the centroid of the object and compute the {BPS}_{o}\in R^{{{\color[rgb]{0,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,0}\pgfsys@color@gray@stroke{0}\pgfsys@color@gray@fill{0}{4096}}}} representation for the object, used for conditioning.

CoarseNet. We pass the object shape BPSo{BPS}_{o} along with initial MANO wrist rotation θwrist\theta_{wrist} and translation γ\gamma to the encoder Q(Z∣θwrist,γ,BPSo)Q(Z|\theta_{wrist},\gamma,{BPS}_{o}) that produces a latent grasp code Z\in R^{{{\color[rgb]{0,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,0}\pgfsys@color@gray@stroke{0}\pgfsys@color@gray@fill{0}{{16}}}}}. The decoder P(θˉ,γˉ∣Z,BPSo)P(\bar{\theta},\bar{{\gamma}}|Z,{BPS}_{o}) maps ZZ and BPSo{BPS}_{o} to MANO parameters with full finger articulation θˉ\bar{\theta}, to generate a 3D grasping hand. For the training loss, we use standard cVAE loss terms (KL divergence, weight regularizer), a data term on MANO mesh edges (L1), as well as a penetration and a contact loss. For the latter, we learn candidate contact point weights from GRAB, in contrast to handcrafted ones or weights learned from artificial data . At inference time, given an unseen object shape, BPSo{BPS}_{o}, we sample the latent space, ZZ, and decode our sample to generate a MANO grasp.

RefineNet. The grasps estimated by CoarseNet are plausible, but can be refined for improved contacts. For this, RefineNet takes as input the initial grasp (θˉ\bar{\theta}, γˉ\bar{{\gamma}}) and the distances DD from MANO vertices to the object mesh. The distances are weighted according to the vertex contact likelihood learned from GRAB. Then, RefineNet estimates refined MANO parameters (θ^\hat{\theta}, γ^\hat{{\gamma}}) in 33 iterative steps as in , to give the final grasp. To train RefineNet, we generate a synthetic dataset; we sample CoarseNet grasps as ground truth and we perturb their hand pose parameters to simulate noisy input estimates. We use the same training losses as for CoarseNet.

GrabNet. Given an unseen 3D object, we first get an initial grasp estimate with CoarseNet, and pass this to RefineNet to get the final grasp estimate. For simplicity, the two networks are trained separately, but we expect end-to-end refinement to be beneficial, as in . Figure 8 (right) shows generated examples; our generations look realistic, as explained later in the evaluation section. For more qualitative results, see the video on our website and images in Sup. Mat.

Contact. As a free by-product of our 3D grasp predictions, we can compute contact between the 3D hand and object meshes, following Sec. 3.3. Contacts for GrabNet are shown with red in Figure 8 (right). Other methods for contact prediction, like , are pure bottom-up approaches that label a vertex as in contact or not, without explicit reasoning about the hand structure. In contrast, we follow a top-down approach; we first generate a 3D grasping hand, and then compute contact with explicit anthropomorphic reasoning.

Evaluation - CoarseNet/RefineNet. We first quantitatively evaluate the two main components, by computing the reconstruction vertex-to-vertex error. For CoarseNet the errors are 12.112.1 mm, 14.114.1 mm and 18.418.4 mm for the training, validation and test set respectively. For RefineNet the errors are 3.73.7 mm, 4.14.1 mm and 4.44.4 mm. The results show that the components, that are trained separately, work reasonably well before plugging them together.

Evaluation - GrabNet. To evaluate GrabNet generated grasps, we perform a user study through AMT . We take 6 test objects from the dataset and, for each object, we generate 2020 grasps, mix them with 2020 ground-truth grasps, and show them with a rotating 3D viewpoint to subjects. Then we ask participants how they agree with the statement “Humans can grasp this object as the video shows” on a 55-level Likert scale (55 is “strongly agree” and 11 is “strongly disagree”). To filter out the noisy subjects, namely the ones who do not understand the task or give random answers, we use catch trials that show implausible grasps. We remove subjects who rate these catch trials as realistic; see Sup. Mat. for details. Table 2 (left) shows the user scores for both ground-truth and generated grasps.

Evaluation - Contact. Figure 9 shows examples of contact areas (red) generated by (left) and our approach (right). The method of gives only 1010 predictions per object, some with zero contact. Also, a hand is supposed to touch the whole red area; this is often not anthropomorphically plausible. Our contact is a by-product of MANO-based inference and is, by construction, anthropomorphically valid. Also, one can draw infinite samples from our learned grasping latent space. For further evaluation, we follow a protocol similar to for our data. For every unseen test object we generate 2020 grasps, and for each one we find both the closest ground-truth contact map and the closest ground-truth hand vertices, for comparison. Table 2 (right) reports the average error over all 2020 predictions, in %\% for the former and cm for the latter case.

Discussion

We provide a new dataset to the community that goes beyond previous motion capture or grasping datasets. We believe that GRAB will be useful for a wide range of problems. Here we show that it provides enough data and variability to train a novel network to predict human grasping of objects, as we demonstrate with GrabNet. But there is much more that can be done. Importantly, GRAB includes the whole-body motion, enabling a much richer modeling than GrabNet.

Limitations: By focusing on accurate MoCap, we do not have synced image data. However, GRAB can support image-based inference by enabling rendering of synthetic human-object interaction or learning priors to regularize ill-posed inference of human-object interaction from 2D images .

Future Work: GRAB can support learning human-object interaction models , robotic grasping from imitation , mapping MoCap markers to meshes , rendering synthetic images , inferring object shape/pose from interaction , or analysis of temporal patterns .

Acknowledgements: We thank S. Polikovsky, M. Höschle (MH) and M. Landry (ML) for the MoCap facility. We thank F. Mattioni, D. Hieber, and A. Valis for MoCap cleaning. We thank ML and T. Alexiadis for trial coordination, MH and F. Grimminger for 3D printing, V. Callaghan for voice recordings and J. Tesch for renderings. Disclosure: In the last five years, MJB has received research gift funds from Intel, Nvidia, Facebook, and Amazon. He is a co-founder and investor in Meshcapade GmbH, which commercializes 3D body shape technology. While MJB is a part-time employee of Amazon, his research was performed solely at, and funded solely by, MPI.

References