Grasping Field: Learning Implicit Representations for Human Grasps
Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael Black, Krikamol Muandet, Siyu Tang
Introduction
Capturing and synthesizing hand-object interaction is essential for understanding human behaviours, and is key to a number of applications including augmented and virtual reality, robotics and human-computer interaction. Despite substantial progress, fully automatic synthesis of highly realistic human grasps remains an unsolved problem. The anatomical complexity of the human hand and the variety of manufactured and natural objects make it extremely challenging to pose the hand such that it interacts with the object in a natural and physically plausible way. Recent data-driven approaches explore deep learning technology to learn and leverage powerful object representations, yet they are mainly limited to simple robotic end effectors, such as parallel jaw grippers . In this work, we seek to understand: 1) what is an efficient and expressive representation for modeling hand-object interaction, that can facilitate realistic human grasp synthesis given an unseen 3D object; and 2) how can we learn such a representation from data.
Our key observation is that human grasping is rooted in physical hand-object contact. Through this contact, humans are able to grasp and manipulate objects naturally. To better model hand-object interaction, we must find a way to effectively represent the contact between hands and objects. To this end, we propose a novel interaction representation that is based on regressing a continuous function that we call the Grasping Field. The grasping field maps any 3D point to a 2D space, where each dimension of the 2D space indicates the signed distance to the surface of the hand and the object respectively (see Sec. 3.1 for a formal definition). Inspired by , we further utilize a deep neural network to parameterize the grasping field and learn it from data. As a result, the learned grasping field serves as a powerful representation to facilitate hand-object interaction modelling.
Based on the grasping field representation, we propose a generative model, in which we generate plausible hand grasps given an object point cloud. We show that our model can produce physically and semantically plausible synthetic grasps, which are similar to the ground truth. Generated grasps on unseen objects are shown in Fig. 1 and Fig. 4.
We further demonstrate the effectiveness of the grasping field representation by considering the task of 3D hand and object reconstruction from a single RGB image. In recent work, Hasson et al. introduce an end-to-end learnable model to reconstruct 3D meshes of the hand and object simultaneously, producing the state-of-the-art results on several datasets. Physical constraints, such as no interpenetration and proper contact, are enforced during the training. However, there are several drawbacks of their mesh-based representation for hand-object interaction modeling. First, they heuristically pre-define regions of the hand that can be in contact with objects. Second, their object representation is limited to objects of genus zero. Third, the resolution of their contact inference is limited by the resolutions of the hand and object meshes. In contrast, with the grasping field representation, it is not necessary to first compute the hand and object meshes, and then compute the contact region. Instead, one can easily infer the contact region by querying the signed distances of input 3D points. Furthermore, the physical constraints, such as no inter-penetration and proper contact, can be efficiently computed and enforced. As demonstrated in our experiments, our model considerably reduces the interpenetration between the reconstructed hand and the object, and improves the quality of 3D hand reconstruction, compared with .
In summary, our contributions are: (1) We propose the grasping field, a simple and effective representation for hand-object interaction; (2) Based on the grasping field, we present a generative model to yield semantically and physically plausible human grasps given a 3D object point cloud; (3) We further propose deep neutral networks to reconstruct the 3D hand and object given an RGB input in a single pass; (4) We perform extensive experiments to show that our method outperforms the baseline on 3D hand reconstruction and on synthesizing grasps that appear natural.
Related work
Human grasp and contact. There is a large body of work on capturing and recognizing human grasps . Recently, introduced a stretch-sensing soft glove to capture accurate hand pose without extra optical sensors. Puhlmann et al. utilized a touch screen to facilitate the capturing process of human grasping. As physical contact is fundamental to hand-object interaction, researchers have proposed methods to capture and modeling contact from diverse modalities , but these often interfere with natural movement. Concurrent to this work, proposed a new dataset of hand-object contact paired with RGB-D images. Our work differs in that our focus is on learning an interaction representation, which is efficient and easy to interface with deep neural networks.
Grasp synthesis. Grasp synthesis is a longstanding problem in robotics and graphics, resulting in an extensive literature . As early as 1991, Rijpkema and Girard proposed a knowledge-based approach to incorporate the role of the human hand, object, environment and animator for the task of computer-animated grasping. More recent works can be categorized into three types of approaches: analytic, data-driven and hybrid approaches. For the analytic approaches , the grasps are often synthesized by formulating the problem as a constrained optimization problem that satisfies a set of criteria measuring the stability or other properties of the grasps. The data-driven approaches often employ machine learning methods to learn representations for synthesizing grasps. An excellent survey of data-driven grasp generation is presented in . Recent hybrid approaches combine analytic models and deep learning tools to synthesize grasps for various end effectors. Finally, the most related approaches to our work was presented in and , where neural networks are used to predict hand parameters of the MANO hand model given object information. In , the model learns to predict the best grasp type from the grasp taxonomy according to the RGB images of the objects. Then, the predicted hands are optimized together with the object meshes to refine the contact points. While in , the parameters are generated directly from the given Basis Point Set of the objects. Our work differs from the previous works in that, by also considering the object distance field, we propose a learnable representation for modelling hand-object interaction that can be used without contact post-processing. Empowered by deep neural networks, the learned representation enables us to synthesize realistic human hand grasping a given object naturally.
Hand pose estimation. Hand pose estimation is a long-standing problem, and various input modalities have been considered, e.g., RGB images or RGB-D and depth sensors . Due to the lack of large scale 3D ground truth data, synthetic data has often been used for training . Recently, instead of estimating the hand skeleton, recovering the pose and the surface of the hand has become popular using statistical hand models, e.g., the MANO model , that can represent a variety of hand shapes and poses . Using the template derived from MANO, show that it is also possible to regress hand meshes directly using mesh convolution. In this work, we represent the 3D hand by a signed distance field, instead of a parametric hand model, due to the difficulty of incorporating object interaction into the model parameter space. For fair comparison with the parametric hand model representation, we fit the MANO model into our resultant signed distance fields. The experimental results indicate the advantage of our new interaction representation.
Object model representation. Learning 3D object models using various types of representations has also been explored . Recently, the community has focused on using the implicit functions such as the Signed Distance Function (SDF) , Occupancy Networks , Implicit Field , and their derivations , as these can model arbitrary object topology with adjustable resolution. Due to these advantages, we also adopt implicit functions to capture hand-object interaction.
Hand-object interaction. Reconstructing hand and object jointly has been studied with both RGB input and RGB-D input . Recently, Hasson et al. achieved promising results on explicitly modeling the contact by combining a parametric hand model MANO , with the mesh based representation for the object. As data for hand-object interaction is limited, we opt to use their synthetic dataset, the ObMan dataset , which is sufficiently large for training a neural network. Our work differs from previous hand-object reconstruction work mainly by focusing on the novel representation of contact and learning both hand and object in the signed-distance space, which allows arbitrary shape modelling and easier distance field manipulation. Furthermore, we go beyond the reconstruction task by proposing generative models to synthesize realistic human grasps given a 3D object.
Method
Inspired by and , we propose to model using a deep neural network, and learn it from data. Therefore, one can infer hand-object interaction in 3D space without the explicit hand and object surfaces. The learned GF can be considered as an interaction prior, which enables us to infer various grasping poses of the hand, only based on the 3D object. Furthermore, in contrast to previous works, e.g. , which can only evaluate body-object interactions after obtaining the body and the object meshes, when using GF as the representation in hand-object reconstruction from images, we model the hand, the object, and the contact area by the implicit surfaces in a common space, largely improving the physical plausibility of the reconstruction.
According to the aforementioned merits of the GF, we use it to address two tasks in this paper; i.e. hand grasp generation given 3D objects and hand-object reconstruction from RGB images. Different GF networks are designed specifically for different tasks.
2 Grasping field for human grasp synthesis
In this section, we show how to use GF to synthesize human grasps. Given an object point cloud, the goal is to generate diverse hand grasps that interact with the object in a natural manner.
The network architecture is shown in Fig. 2(a); we adopt the encoder-decoder framework. To extract features from point clouds, we use the PointNet encoder with residual connection. The encoder is trained jointly with other network layers from scratch. The encoder-decoder network takes a query 3D point, and two point clouds of the hand and the object as input, and produces the signed distances of the query point to the hand and the object surfaces. In addition, the encoded object point cloud feature is fed into the hand point cloud encoder, leading to a hand distribution conditioned on the object. Note that this variational encoder-decoder network only requires both hand and object point clouds during the training. During inference, only the conditioning object point cloud and the query point are required. The hand features are sampled from the learned latent space, as in a standard VAE . The training loss consists of the following terms:
The reconstruction loss : For each query point , the input object point cloud and the input hand point cloud , the reconstruction loss is designed for the hand and object individually:
where is the grasping field network (Fig. 2(a)). and are the ground truth SDF for the hand and object, respectively. In addition, is a function to constrain the distance within . is set to 1cm in all experiments.
KL-Divergence : In order to generate new hand grasps, we use a KL-divergence loss to regularize the distribution of hand latent vector , obtained from the hand point cloud encoder , to be a normal distribution. The loss is given by
where denotes a standard high-dimensional normal distribution, and denotes mean and standard deviation. For generation, the hand latent vector is sampled from a standard normal distribution.
Classification loss : Besides predicting the signed distances of a query point, we also train the network to produce the hand part label of a query point to parse the hand semantically. To achieve this, we introduce a classification loss, which is given by a standard cross-entropy loss. The hand part annotation is based on the MANO model as illustrated in Fig. 2(b).
3 Grasping field for 3D hand-object reconstruction from a single RGB image
Network architecture. The network architectures are illustrated in Fig. 3, which are designed to recover both hand and object in a single pass. To enable a direct comparison with , the two-branch network is employed (Fig. 3(a)), which addresses hand and object individually. Similar to , we introduce contact and inter-penetration losses during the training to facilitate a better 3D reconstruction on the contact regions of the hand and the object. To introduce hand-object interactions in early stages, we propose a one-branch network (Fig. 3(b)), which uses the same image encoder and has the same number of layers with the two-branch model. See Appendix A for architecture details.
The training loss consists of the following terms:
The reconstruction loss : For each query point and the input image , the reconstruction loss is designed for the hand and object individually, and is given by
in which is our conditinal grasping field network, and is the ground truth SDF for the component (hand or object). is the thresholding function to constrain the distance within as with the generative model proposed in Sec. 3.2.
The inter-penetration loss : To avoid surface inter-penetration between the reconstructed hand and object, we define the inter-penetration loss as
where is a 2D one-vector, and denotes a dot product. This loss function actually penalizes the negative sum of predicted signed distances to the object and to the hand. If the hand and the object are separate and have no contact, the signed distance sum of every point in 3D space is always positive, and hence is ignored by our inter-penetration loss. On the other hand, if the hand and the object have inter-penetration, then this inter-penetration loss does not only penalize the points in the intersection volume, but also all 3D points in the space, indicating that the predicted hand and object are incorrect. Compared to the inter-penetration methods in , which only penalize the intersection volume, our loss applies stronger constraints.
Contact loss : Our proposed contact loss encourages hand-object contact, and is given by
where is a hyper-parameter. We can see that corresponds to the hand-object contact surface. Therefore, it ignores points with predicted grasping field , and only encourages points with to be the contact points. In our study, we empirically set based on the hand-object interactions in the training data. Finally, we employ the same Classification loss as the one proposed in Sec. 3.2.
4 From grasping field to mesh
With the trained grasping field conditioned on images or point clouds, one can compute the signed distances to the hand and object of a query 3D point. To recover the hand, object and their interactions, we first randomly sample a large number of points, and evaluate their signed distances. The point clouds belonging to the hand and the object can be selected, according to point-object signed distances close to zero. Then, the hand mesh and the object mesh are obtain by marching cubes .
In addition, the hand mesh can be recovered by fitting the MANO model to the hand point cloud. In this case, we can obtain hand segmentation, hand joint positions, and a compact representation of the hand simultaneously, according to the pre-defined topology in MANO.
The implementation details are thoroughly presented in Appendix A.
Experiments
We demonstrate the effectiveness of the grasping field representation on two challenging tasks: human grasp generation given a 3D object and 3D hand-object reconstruction from a single image.
To train the generative model for human grasp synthesis, we need ground truth 3D meshes of interacting hands and objects. Unfortunately, existing datasets often lack the desired properties.The limitations include small dataset size and lack of 3D ground truth hand pose or shape. Consequently we use the synthetic ObMan dataset to train our model. The data is generated from a statistical hand model, MANO , and 2772 object meshes covering 8 classes of everyday objects from the ShapeNet dataset . Hand-object interaction is generated using a physics simulator, GraspIt , resulting in high-quality hand-object interaction. Due to the limited number of grasp types in the FHB dataset and the HO-3D dataset , they are not suitable for training the generative model (see Appendix B). Instead, we use them to test the generalization ability of the generative model trained on the ObMan grasps.
For the 3D reconstruction task, we also mainly use the ObMan dataset for training and testing. To test the effectiveness of our network on real-world images, however, we follow the same approach as to train and test on the FHB dataset.
Evaluation metrics.
Our goal is to generate physically plausible and semantically meaningful 3D human hand given an object. Therefore, we quantitatively evaluate the generated samples according to physics-based metrics and use large-scale perceptual studies to measure the visual realism of the grasps. For the 3D reconstruction task, we use Chamfer distance and hand joint error. Details of the evaluation metrics are in Appendix C.
(1) Physical metrics: A valid human grasp implies stable hand-object contact without interpenetration. Consequently, we use the following evaluation metrics: a) Intersection volume and depth. The hand and object mesh are voxelized and the interpenetration depth is the maximum distance from all the intersected voxels to the surface of another mesh. b) Ratio of samples with contact. We define a contact between the object and the hand when any point on the surface of the hand is on or inside the surface of the object. We calculate the ratio of samples over the entire dataset that have interpenetration depth more than zero. c) Grasp stability. Using physics simulation , we hold the hand constant, apply gravity, and measure the average displacement of the object’s center of mass during a fixed time period.
(2) Semantic metric: We perform perceptual studies using Amazon Mechanical Turk to evaluate the naturalness of our generated grasps. Details of the study are presented in Appendix C and D.
(3) 3D reconstruction quality: We use the Chamfer distance between reconstructed and ground truth hand surfaces to evaluate the hand reconstruction quality as in . Hand joint distance is computed following .
1 Evaluation: Human grasps generation
Baseline. To our knowledge, there is no previous model that learns to synthesize natural human grasps given a 3D object. Rather than randomly placing the hand around the object, we trained a strong baseline model for grasp generation. Specifically, we replace the decoder (i.e. the grasping field) of our conditional VAE model (Fig. 2(a)) with fully connected layers to regress MANO hand parameters. Then given a 3D object point cloud and a random sample, our baseline model generates MANO parameters directly. Generated grasps from the baseline are shown in Appendix E.
Results. We show the systematic quantitative evaluation of the generative method in Tab. 1 and qualitative results in Fig. 1, 4 and Appendix E (Fig. 3, 4). The baseline and the GF model are only trained on the ObMan training set, and tested extensively on the objects from the ObMan test set, FHB and HO3D. Our proposed GF performs substantially better than the baseline on ObMan and achieves comparable quality as the ground truth grasps. When the model is tested on the FHB objects, which are never seen during training, it achieves a comparable perceptual score compared to the ground truth grasps. Surprisingly, on HO3D, our synthesized grasps are judged more realistic than the ground truth grasps of real humans (3.29 vs 3.18). These perceptual studies suggest that our method makes an important step towards the fully automatic synthesis of realistic human grasps.
Regarding the physical plausibility, we observe that our model achieves a better contact ratio and grasp stability (physics simulation) than the ground truth grasps on FHB and HO3D. This is likely due to the GF results having a larger intersection volume. One reason is that there are a very limited number of objects in these two datasets. Some of the test objects are very different from the training objects, resulting in more inter-penetration for the generated grasps. Overall, the combination of visual realism and grasp stability suggests that our results are approaching the level of natural human grasps.
2 Evaluation: 3D hand-object reconstruction
Apart from serving as a powerful representation for the synthesis task, the proposed GF also facilitates the 3D reconstruction task. In the following, we analyze the different network architectures and training losses proposed in Sec. 3.3. We compare with the baseline method on the ObMan dataset and on the FHB dataset. The results are summarized in Tab. 4. Due to many limiting factors of the real-world datasets such as FHB and HO3D (see Appendix B for detailed data analysis), learning a reasonable model for joint object and hand reconstruction is extremely challenging. Instead, to evaluate the effectiveness of the GF representation for 3D reconstruction on real-world images, we follow the setting of the latest work , where the object 3D model is given as input. Note that this is a commonly used setting in previous works (e.g. ).
Network design. We first analyze the two different network architectures presented in Fig. 3 denoted with and without ‘2De’ respectively. Both architectures achieve comparable performance for hand reconstruction, however differ significantly for intersection error, where the one decoder model achieves considerably better performance, due to the efficient joint modeling of hand-object interaction. Compared with the baseline , the intersection volume and depth are reduced from 6.25 and 1.20 to 0.65 and 0.32, respectively. The contact ratio are comparable among two architectures and baseline model. All our models considerably improve the quality of hand reconstruction, compared with . The object reconstruction quality is behind hand quality for all model variations including the baseline model. Note that the ObMan dataset contains more than 1600 objects from 8 different classes. The object reconstruction performance is decreased as it remains unclear how to learn the implicit representations to reconstruct a large variety of object classes with a single model and such a task is beyond the focus of this work. Please see Appendix E (Fig. 1) for visualization.
Training losses. The effect of the contact and interpenetration loss (L) is shown in Tab. 4 (a), when the loss is imposed during the training of the two-decoder network, the intersection volume and depth are reduced and the overall quality of the interaction is considerably improved. In contrast, for the one-decoder model, our observation is that, for a large portion of 3D points, the signed distances to the object and to the hand are highly correlated, the model that jointly predicts both signed distance values does not need to enforce this auxiliary training loss.
MANO fitting. As shown in Tab. 4 (a), MANO fitting (indicated by GF-MANO) does not have a substantial influence on the reconstruction quality. This implies that on the one hand, the reconstructed hand of our GF model is realistic enough without a statistical model to regularize it, and that on the other hand, the output hand part labels are accurate enough for us to fit the MANO model and retrieve hand joints or shape parameters for applications that need these without undermining the shape and contact estimation. A qualitative illustration is presented in Fig. 2.
Hand reconstruction on real-world images. To analyse the effectiveness of the proposed grasping field representation for the 3D reconstruction task on real-world images, we compare our method with the latest work on the FHB dataset. Compared to , the key difference in is that the object is given as part of the input. We explore the same network architecture as and only replace the decoder part with the grasping field. The implementation details are presented in Appendix A (Fig. 3).
As stated in , definition of hand joint locations vary between datasets. Without hand surface annotation in the FHB dataset, it is difficult to train an accurate regressor that maps between the FHB markers and the MANO joints. Assuming that the joints are identically defined, we fit the MANO model to the FHB markers by minimizing the distance between the MANO joints and the FHB markers. Then the MANO joints obtained in such way are considered as our pseudo ground truth joints, and the obtained MANO surface is used to supervise the training.
We compare the predicted MANO joints with the pseudo ground truth joints as well as the original FHB markers assuming identical joints. As our model is not trained to optimise for the FHB marker locations, the reconstruction error is larger than as shown in Tab. 4 (b). When we evaluate our prediction on the pseudo ground truth MANO joints, the reconstruction error decreases from to . This suggests that the proposed grasping field representation is effective for the task of 3D hand reconstruction from a single image, achieving comparable performance with respect to the start-of-the-art.
Conclusion and Discussion
In this work, we propose a novel representation for hand-object interaction, namely the grasping field. Learning from data, the GF captures the critical interactions between hand and object by modeling the joint distribution of hand and object shape in a common framework. To verify the effectiveness, we address two challenging tasks: human grasp generation given a 3D object and shape reconstruction given a single RGB image. The experiments show that the generated hand grasps appear natural and are physically plausible while the hand reconstruction achieves comparable performance as the state-of-the-art.
A limitation of our work is that there is no explicit modeling of the object functionality and human action in the current grasping field representation. In reality, a person holds an object differently based on different intentions. For instance, using a knife or passing it to someone else result in completely different human grasps. One promising future research direction is to incorporate human intention and object affordances into the grasping field for action specific grasps generation. Furthermore, we believe the proposed grasping field representation opens up avenues for several other future research directions. For instance, 3D human hand generation given only an object image and synthesizing the motion of hand-object interactions.
Acknowledgement. We sincerely acknowledge Lars Mescheder and Michael Niemeyer for the detailed discussions on implicit function, Dimitrios Tzionas, Omid Taheri, and Yana Hasson for insightful discussions on hand interaction, Partha Ghosh and Qianli Ma for the help with VAE. Disclosure. MJB has received research gift funds from Intel, Nvidia, Adobe, Facebook, and Amazon. While MJB is a part-time employee of Amazon, his research was performed solely at MPI. He is an investor in Meshcapde GmbH.
References
Appendix A Implementation Details
In Sec. 3.2 and Sec. 3.3, we present the neural networks that are used for human grasps generation and reconstruction, respectively. Here we discuss the implementation details.
In this section, we explain the network architectures used in our experiments. The same decoder architecture is used in our image reconstruction and the hand generation tasks. We change the encoder architectures according to the input type. In our experiments, both the encoder and decoder are jointly trained end-to-end. Figure 1 illustrates the one-branch decoder with 8 fully-connected layers used in all tasks.
For image reconstruction, we use the ResNet18 model pretrained on the ImageNet dataset as an encoder. We change the last layer of the encoder to produce a latent vector of size 256 for the decoder.
For the point cloud input, we use two separated PointNet encoders with additional pooling and expansion layers presented in . In each encoder, 3D points are first mapped to 512-dimension feature vectors followed by 5 ResNet-blocks, producing a latent vector of size 256. The latent codes for hand and object are then concatenated to make a 512-dimension latent code.
For the hand generation task, we change the first layer in the hand encoder to produce a 256-dimension vector for each point then concatenate it with the 256-dimension object latent vector. Figure 2 shows the details of the point cloud model.
For image reconstruction with known objects, we assume that the object mesh in the normalized pose is given. We sample surface points from the given object and use a PointNet encoder to compute object latent vector of size 128. The object latent vector is then concatenated with a hand latent vector of size 128 from ResNet18 encoder, producing a latent code of size 256 for the decoder. The overview of the network is shown in Figure 3.
A.2 Data preparation
To prepare the sampled 3D points and their distances to the hand and object surfaces for training, we follow the point sampling method provided by : For each pair of the hand and object meshes, we translate both meshes such that the hand root joint is at the origin then scale them to fit in a unit cube. The scaling factor is the same for the entire dataset to ensure the hands are normalized across dataset. After that, 40,000 points are sampled in a unit cube. Following , 95% of the total points were sampled near the surface to capture the details of both meshes. For the Chamfer distance calculation, we sample 30,000 points from the surface of the ground truth mesh and reconstructed mesh following . In case the reconstructed mesh contains more that one connected component, only the largest watertight connected component is retained.
A.3 Training
The contact loss is disabled in the beginning. When computing the reconstruction loss , hand points to the object surface, and object points to the hand surface, are not considered until the contact loss is enabled. In our trials, we observe dramatic degradation when such a mask is not used or when the contact loss is enabled in the beginning.
For the generative GF network conditioned on an object point cloud, the KL loss, , is employed in an annealing scheme; the loss weight is kept at 0 in the first 200 epochs and then linearly increased to 0.1 over the next 200 epochs. We find that such a annealing scheme is essential in our trials. Applying the KL loss in the beginning causes our generative network posterior to collapse.
In all experiments, we use Adam optimizer with learning rate of and decay it to at after 600 epochs. We train the models for 1,200 epochs without hand-part classification loss and another 100 epochs with the classification loss. Weight decay is used in all layers in the decoder.
A.4 Inference
During inference, we use Marching Cube with resolution 128 to obtain hand and object meshes. As the object can vary in size, we use a two-stage approach to dynamically scale the cube size in the Marching Cube algorithm. First, to find the boundary of the reconstructed meshes, we query equally space points in a unit cube centered at the origin point to locate the negative signed-distance values which indicate the inside of the mesh. Then, we query again with a cube that covers every negative-value point. Using this approach, no mesh is produced if no negative point is found in the first stage.
Appendix B Dataset Analysis
In this section, we provide detailed analyses on the FHB and the HO-3D datasets. Although these datasets considerably contribute to the studies of hand-object interactions with detailed 3D annotation, our analyse shows that they might not be suitable for learning human grasps and modelling the accurate contact relation between hand and object. First, the number of objects and the types of grasps are limited. As shown in Tab. 5, the number of object is 3 in the FHB dataset and 10 in the HO3D dataset. Second, the (pseudo) ground truth meshes of the interacting hand and object exhibit frequent interpenetration. Sampled ground truth meshes from the HO-3D dataset are illustrated in Fig. 1.
We evaluate the interpenetration between hand and object meshes quantitatively. We use the same evaluation metric as the one presented in the main paper, namely, the intersection volume (cm3) and depth (cm) (Sec. 4). The results are shown in Table 5.
For the HO-3D dataset, 91.94% of the training examples exhibit hand-object contact. However, among these training examples, the average intersection volume and depth are 10.91 cm3 and 1,56 cm respectively.
For the FHB dataset, we use the similar subset as the previous work , namely, we exclude the milk bottle related examples and the examples where the distance from hand joints to the object mesh is more than 1cm. We refer to this dataset as FHBc. We further fit the MANO hand model with the provided joint location. For FHBc, 97.1% of the training examples have hand-object contact. However, similar level of intersection between hand and object meshes can be observed in Table. 5.
Overall, the evaluation shows considerable intersection volume and depth of the training data. Therefore we use the ObMan dataset as our main training dataset, where the ground truth quality of the contact regions is more suitable for learning physically plausible human grasps.
Furthermore, as shown in the experiment section (Tab. 1), our generated grasps that are learned from the ObMan dataset obtain a higher perceptual score than the ground truth grasps from the HO3D data in the perceptual study, suggesting that the physical plausibility, i.e. no interpenetration and proper contact, plays an important role on the naturalness of human grasps.
Appendix C Details of the evaluation metrics
Evaluation Metrics. For human grasps synthesis, our goal is to generate physically plausible and semantically meaningful 3D human hand given an object. Therefore, we propose to quantitatively evaluate the generated samples using physics metrics and a large-scale perceptual study to measure the perceptual fidelity. In addition, for the quantitative evaluation of our reconstruction networks, we use Chamfer distance and hand joint error.
(1) Physical metric: A valid human grasp implies hand-object contact without interpenetration. Naturally we propose with the following evluation metrics:
Intersection volume and depth. We follow to report intersection volume and depth. The hand and object mesh are voxelized using a voxel with edge length of 0.5cm. The interpenetration depth is the maximum distance from all points on the interpenetrated surface to another surface. If the meshes do not overlap, the interpenetration depth is defined as 0.
Ratio of samples with contact. We define a contact between object and hand when any point on the surface of hand is on or inside the surface of the object. To measure the performance of models on hand-object contact quality, we calculate the ratio of samples over the entire dataset that have interpenetration depth more than zero. As all of the samples in the dataset should have contact between hand and object, the best ratio of frames with contact is 100%. A good hand and object reconstruction model should have high ratio of contact and small interpenetration volume and depth.
Simulation displacement. Following , we use physics simulation to evaluate the stability of the grasps. In the simulated environment , we fix the hand and measure the average displacement of the mass center of the object in a give time period. Small displacement suggests a stable grasp.
(2) Semantic metric: We perform perceptual studies on Amazon Mechanical Turk to evaluate the authenticity of our generated grasps. For each randomly generated sample, we render images from 6 different views, and request participants to score from 1 (low fidelity) to 5 (high fidelity).
(3) 3D reconstruction quality: We use the Chamfer distance between reconstructed and ground truth hand surfaces to evaluate the hand reconstruction quality. Surface distance is approximated by mean square point cloud Chamfer distance (cm2) as implemented in . The MANO wrist is sealed to form a watertight mesh for fair comparison. Joins distance is computed following . After MANO parameters are recovered from the predicted hand mesh as described in Sec. 3.4, we compute mean Euclidean distance over 21 joints following . Note, since scale and global translation can not be determined by a single image, for each predicted hand, we optimize the scale and global translation to match the ground truth by minimizing the Chamfer distance between them. Similarly to hand, we also use Chamfer distance as measurement of object surface quality. The predicted object mesh is transformed according to the corresponding predicted hand transformation estimated from the above to align with the ground truth object mesh.
Appendix D Details of the perceptual study
Figure 1 shows the user interface for evaluating the generated grasps on the Amazon Mechanical Turk (AMT). Users are asked to rate the plausibility of the hand-object interactions individually. Each entry consists of images from six different views and is rated by three different users.
Appendix E Qualitative results
Figure 5 shows the generated grasps from our baseline VAE model conditioned on the object surface point cloud. The MANO parameters are directly predicted by the decoder.
Figure 1 shows the reconstruction results on the test images of the ObMan dataset . We observe that our model can recover hand meshes with proper interaction with the object.
Figure 2 shows the comparison between reconstructed mesh before and after MANO fitting. The hand meshes also come from the single image reconstruction task. We observe that the MANO fitted meshes match the inferred meshes, even in the case where the rasterized hand mesh has merged fingers.
Figure 3 shows randomly sampled grasps from our VAE model conditioned on the object surface point cloud. We observe that our model can generate a variety of grasps given an object. Figure 4 shows the generated results of the same model conditioned on the objects from HO3D datset. It should be note that this model is only trained on the ObMan dataset and have never seen these objects before.