Hand-Object Contact Consistency Reasoning for Human Grasps Generation

Hanwen Jiang, Shaowei Liu, Jiashun Wang, Xiaolong Wang

Introduction

Capturing hand-object interactions has been an active field of study and it has wide applications in virtual reality , human-computer interaction and imitation learning in robotics . In this paper, we study the interactions via generation: As shown in Fig. 1, given only a 3D object in the world coordinate, we generate the 3D human hand for grasping it. Unlike predicting robot grasps with parallel jaw grippers , predicting human grasps is substantially more difficult because: (i) Human hands have a lot more degrees of freedom, which leads to much more complex contact; (ii) The generated grasp needs to be not only physically plausible but also presented in a natural way, consistent with how objects are usually grasped.

To synthesize physically plausible and natural grasp poses, recent works propose to use generative models supervised by large-scale datasets with grasp annotations and contact analysis on hands. Specifically, the large-scale dataset allows the model to generate realistic human grasps and the contact analysis encourages the hand contact points to be close with the object but without inter-penetration. While these methods put a lot of efforts into modeling the hand and its contact points, they ignore that the object itself also has more possible contact regions that need to be reached (see contact map in Fig. 1). In fact, recent work has studied the common contact regions on objects and trained neural networks to directly predict the contact map from the 3D object model .

In this paper, we argue that it is critical for the hand contact points and object contact regions to reach mutual agreement and consistency for grasp generation. To achieve this, we propose to unify two separate models for both the hand grasp synthesis and object contact map estimation. We show that the consistency constraint between hand contact points and object contact map is not only useful for optimizing better grasps during training time by designing new losses, but also provides a self-supervised task to adjust the grasp when testing on a novel object. We introduce the two components as follows.

First, we train a Conditional Variational Auto-Encoder (CVAE) based network which takes the 3D object point clouds as inputs and predicts the hand grasp parameterized by a MANO model , namely GraspCVAE. During training the GraspCVAE, we design two novel losses with one encouraging the hand to touch the object surface and another forcing the object contact regions touched by the ground truth hand close to the predicted hand. With these two consistent losses, we observe more realistic and physically plausible grasps.

Second, given the hand grasp pose and object point clouds as inputs, we train another network that predicts the contact map on the object. We name this model the ContactNet. The key role of the ContactNet is to provide supervision to finetune GraspCVAE during test time when no ground truth is available. We design a self-supervised consistency task, which requires the hand contact points produced by the GraspCVAE to be consistent and overlapped with the object contact map predicted by the ContactNet. We use this self-supervised task to perform test-time adaptation which finetunes the GraspCVAE to generate a better human grasp. This adaptation approach can be applied on each single test instance. We emphasize that this procedure does not require any extra outside supervision and it can flexibly adapt to different inputs by resuming to the model before adaptation.

We evaluate our approach on multiple datasets include Obman , HO-3D and FPHA datasets. We show that by utilizing the novel objectives based on the contact consistency constraints in training time, we achieve significant improvements on human grasps generation against state-of-the-art approaches. More interestingly, by optimizing with the proposed self-supervised task during test time, it generalizes and adapts our model to unseen and out-of-domain objects, leading to the large performance gain.

Our contributions of this paper include: (i) Novel hand-object contact consistency constraints for learning human grasp generation; (ii) A new self-supervised task based on the consistency constraints which allows the generation model to be adjusted even during test time; (iii) Significant improvement on grasp generation for both in-domain and out-of-domain objects.

Related Works

Hand-object interaction. Modeling and analysing hand-object interaction is an active field of study with two main paradigms: joint estimating hand-object poses simultaneously during interaction and studying from multi-modal hand-object interaction representations . To perform hand-object pose estimation, Tekin et al. proposed a 3D detection framework, where the hand-object poses are predicted by two output grids without explicit interaction between them. On the contrary, Hasson et al. leveraged the hand-centric physical constraints for modeling the interaction between hand-object to avoid penetration. Inspired by this work, we also applied explicit constraints for grasp generation.

Another paradigm of study is to analyze the forces on hand and contact regions on objects from the multi-modal data. For example, Sundaram et al. introduced a scalable tactile glove, and utilized the touching information for object classification, while Glauser et al. leveraged it for a more difficult hand pose estimation task. Instead of using the tactile sensors on hand, Brahmbhatt et al. proposed to use thermal cameras to capture object contact maps, which reflects the object common contact regions after grasping. Inspired by this work, Taheri et al. further built a GRAB dataset, which not only captures the contact map from hand, but also takes the whole human body into consideration. This line of research motivates us to go beyond modeling hand-centric grasp generation, and explore the object-centric contact map by designing an object-centric loss to encourage the common contact regions on object to be touched by the hand.

Grasp generation. Generating human grasp is very challenging due to the higher degree-of-freedom of the human hand . To generate a realistic grasp, Karunratanakul et al. proposed an implicit representation for modeling the joint distribution of hand-object shape. Instead of implicit representation, our work is more related to work by Brahmbhatt et al. which made use of the object contact maps to filter multiple generated grasps from GraspIt! . However, the contact maps are taken as a constraint rather than a learning target in this grasp generation framework. In our work, we leverage the consistency between hand-object contact regions as training targets with the use of object contact map. Moreover, a self-supervised task is also designed for adjusting generated grasps using the contact maps at test-time.

Affordance prediction. Predicting scene and object affordance plays an important role in visual understanding . For example, Corona et al. proposed a novel dataset and a generative network for learning grasps of multiple on-table objects. Williams et al. leveraged object affordance prediction for real-world robotic manipulation. Inspired by these works, our goal is to generate grasps by learning the object affordance with ensuring the perceptual naturalness and physical plausibility at the same time. Different from the previous works, our method also allows better generalization of grasping out-of-domain objects with the help of the proposed self-supervised task.

Learning on test instances. Improving the generalization ability of neural networks is one of the most important problem in machine learning . Recent research has started tackling the problem by leveraging self-supervision at test-time . For example, Shocher et al. proposed a self-supervised super-resolution framework where the network is only trained at test-time by up and downscaling a single test example. Sun et al. extended the test-time adaptation idea to more general applications with a joint training framework of an image recognition task and a self-supervised task. At test-time, the network can be adjusted to a single test image by tuning the self-supervised objective. While this approach is intriguing, it is still unclear how the self-supervised objective can affect the main task objective. Inspired by this work, our approach also leverage self-supervision on a single instance for test-time adaptation. Different from , our self-supervised task is directly optimizing the main goal of generating better human grasps, which ensures the performance gain.

Approach

Our goal is generating hand meshes as human grasps given object point clouds as inputs. The generated hand mesh not only needs to be presented in a natural and realistic way, but it should also hold the object tightly in a physically plausible manner. We emphasize that ensuring reasonable contact between the object and synthesized hand is the key to get high-quality and stable human grasps.

To deal with this problem, we utilize both hands and object contact information and make sure they are consistent with each other, as summarized in Fig. 2. We propose two networks, a generative GraspCVAE to synthesize grasping hand mesh, and a deterministic ContactNet for modeling the contact regions on the object.

Training Stage. As shown on the left side of Fig. 2, we optimize these two networks using ground-truth supervision separately to learn grasp generation and predicting object contact maps. In this stage, the inputs of GraspCVAE are both of hand and object, and GraspCVAE learns to synthesizing grasps in the hand reconstruction paradigm, where both of the its encoder and decoder will be used. Note that this follows the standard procedure in Conditional Variational Auto-Encoder (CVAE) . To train the GraspCVAE, we propose two novel losses to ensure the hand-object contact consistency: one loss forcing the prior hand contact vertices to be close to the object surface, and another loss encouraging the object common contact regions to be touched by hand at the same time. The object and generated hand will find mutual agreement on the form of contact with the two losses during training.

Testing Stage. As shown on the right side of Fig. 2, we unify the two networks and design a self-supervised task by leveraging the consistency between their outputs. Given a test object, we first generate an initial grasp from the GraspCVAE decoder (without the encoder). Different from the training stage, the reconstruction target – grasp is not provided in testing . Then, the generated grasp is forwarded together with the object to the ContactNet to predict a target contact map. Since ContactNet is trained with ground truth data, where penetration between hand-object does not exist and hand fingers are touching the object surface closely, it will model the pattern of the ideal hand-object contact. During testing, the predicted contact map from ContactNet will tend to contain the ideal contact pattern. We use the predicted contact map from the ContactNet as a target for finetuning and optimizing the grasps generated by GraspCVAE. If the grasp is predicted correctly from GraspCVAE, the object contact region from the predicted grasp should be consistent with the target object contact map. We use this consistency as a self-supervision signal to adapt grasps generated by GraspCVAE during test-time.

In the following, we will first introduce the individual framework for GraspCVAE and ContactNet, and then the test-time contact reasoning with both networks for better adaptation to the new objects.

The GraspCVAE is a Conditional Variational Auto-Encoder (CVAE) based generative network, which uses conditional information to control generation. For GraspCVAE, the conditional information is the object. We follow to use the GraspCVAE: In training, both of encoder and decoder of GraspCVAE is used to learn the grasp generation task in a hand reconstruction manner by taking in both of hand-object as input; At test-time, only its decoder is used to generate human grasp of a object with only the 3D object as the input (without using the grasp for input). The network architecture is shown in Fig. 4.

During testing, as shown in the bottom row of Fig. 4, we only utilize the decoder from the GraspCVAE for inference. Given only the extracted object point cloud feature Fo\mathcal{F}^{o} and a latent code zz randomly sampled from a Gaussian distribution as inputs, the decoder will generate the parameters for the MANO model which leads to the hand mesh output.

Given this architecture, we then introduce the training objectives as follows. We will first introduce the baseline objectives and then two novel losses which encourage the hand-object contact consistency.

The first objective for the baseline model is mesh reconstruction error, which is defined on both the vertices of the mesh as well as the parameters of the MANO model. We adopt the L2L_{2} distance to compute the error. We denote the reconstruction loss between the predicted vertices and the ground-truths as LV=∣∣V^−Ph∣∣22L_{\mathcal{V}}=||\hat{\mathcal{V}}-\mathcal{P}^{h}||^{2}_{2}. The losses on MANO parameters are defined in a similar way with LθL_{\theta} and LβL_{\beta}. The reconstruction error can be represented asL_{\mathcal{R}}=\lambda_{\mathcal{V}}\cdot L_{\mathcal{V}}+\lambda_{\theta}\cdot L_{\theta}+\lambda_{\beta}\cdot L_{\beta}\, where λV,λθ\lambda_{\mathcal{V}},\lambda_{\theta} and λβ\lambda_{\beta} are constants balancing the losses.

Following the training of VAE , we define a loss enforcing the latent code distribution Q(z∣μ,σ2)Q(z|\mu,\sigma^{2}) to be close to a standard Gaussian distribution, which is achieved by maximizing the KL-Divergence as LKLD=−KL(Q(z∣μ,σ2)∣∣N(0,I)) .L_{\mathcal{KLD}}=-KL(Q(z|\mu,\sigma^{2})||\mathcal{N}(0,I))\ .

We also encourage the grasp to be physically plausible, which means the object and hand should not penetrate into each other. We denote the object point subset that is inside the hand as Pino\mathcal{P}^{o}_{in}, then the penetration loss is defined as minimizing their distances to their closest hand vertices Lpenetr=1∣Pino∣∑p∈Pinomin⁡i∣∣p−V^i∣∣22 .L_{penetr}=\frac{1}{|\mathcal{P}^{o}_{in}|}\sum_{p\in\mathcal{P}^{o}_{in}}\min_{i}||p-\hat{V}_{i}||_{2}^{2}\ . In a short summary, the loss for training the baseline is:

where λKLD\lambda_{\mathcal{KLD}} and λp\lambda_{p} are constants balancing the losses.

There are two potential challenges in the baseline framework: First, the losses in the baseline model ignore physical contact between the hand-object, which cannot ensure the stability of the grasp; Second, grasp generation is multi-modal and the ground-truth hand pose is not the only answer. To tackle these challenges, we design two novel losses from both the hand and the object aspects to reason plausible hand-object contact and find the mutual agreement between them.

Hand-centric Loss. We define the prior hand contact vertices Vp\mathcal{V}^{p} as shown in Fig. 6, motivated by . Given the predicted locations of the hand contact vertices, we then take the object points nearby as possible points to contact. Specifically, for each object point Pio\mathcal{P}^{o}_{i}, we compute the distance D(Pio)=min⁡j∣∣Vjp−Pio∣∣22\textbf{D}(\mathcal{P}^{o}_{i})=\min_{j}||\mathcal{V}^{p}_{j}-\mathcal{P}^{o}_{i}||^{2}_{2}, and if it is smaller than a threshold, we take it as the possible contact point on the object. Our hand-centric objective is to push the hand contact vertices close to the object as,

for all the possible contact points on the object, where T=1 cm\mathcal{T}=1\ cm is the threshold. The final loss combining the two novel losses above is,

where λH\lambda_{\mathcal{H}} and λO\lambda_{\mathcal{O}} are constants balancing the losses. Intuitively, the LOL_{\mathcal{O}} generally answers the question Where to grasp? and does not specify which hand part should be close to the object contact regions. And LHL_{\mathcal{H}} is used to find the answer of Which finger should contact? dynamically. During training, with the two proposed losses, the hand contact points and object contact region will reach mutual agreement and be consistent to each other for generating stable grasps.

2 Learning ContactNet

During training time, the inputs for the ContactNet are directly obtained from the ground-truths.

3 Contact Reasoning for Test-Time Adaptation

During testing, we unify the GraspCVAE and ContactNet in a cascade manner as shown on the right side of Fig. 2. Given the object point clouds as inputs, the GraspCVAE will first generate a hand mesh M^\hat{\mathcal{M}} as the initial grasp. We compute its object contact map ΩM^\Omega_{\hat{\mathcal{M}}} correspondingly. Taking both the predicted hand mesh and object as inputs, the ContactNet will predict another contact map Ωc\Omega^{c}. If the grasp is predicted correctly, the two contact map ΩM^\Omega_{\hat{\mathcal{M}}} and Ωc\Omega^{c} should be consistent. Based on this observation, we define a self-supervised consistency loss as Lrefine=∣∣ΩM^−Ωc∣∣22L_{refine}=||\Omega_{\hat{\mathcal{M}}}-\Omega^{c}||^{2}_{2} for fine-tuning the GraspCVAE. Besides this consistency loss, we also incorporate the hand-centric loss LHL_{\mathcal{H}} and penetration loss LpenetrL_{penetr} to ensure the grasp is physically plausible. We apply the joint optimization with all three losses on a single test example as,

We use this loss to update the GraspCVAE decoder, and freeze other parts of the two networks.

Experiment

We show qualitative results of generated grasps from our methods, and compare the qualitative performance with other methods in Sec. 4.4. Then, we give ablation studies on the effectiveness of proposed novel losses during training and different Test-Time Adaptation (TTA) paradigms in Sec. 4.5.

We sample N=3000N=3000 points on the object mesh as the input object point clouds. In training, we use Adam optimizer and LR=1e−4LR=1e-4 with 100 epochs, where the LRLR is reduced half when model trained 30,60,80,9030,60,80,90 epochs. Batch size is 128128. The loss weights are λβ=0.1\lambda_{\beta}=0.1, λθ=0.1\lambda_{\theta}=0.1, λp=5\lambda_{p}=5, λH=1500\lambda_{\mathcal{H}}=1500 and λO=100\lambda_{\mathcal{O}}=100. For Test-Time Adaptation, we use optimizer SGD with Momentum 0.80.8, LR=6.25×10−6LR=6.25\times 10^{-6} is same as last epochs in training. For each sample, we use batch augmentations with batch size 3232. The loss weights are λp=5\lambda_{p}=5, λH=1\lambda_{\mathcal{H}}=1 and λO=5\lambda_{\mathcal{O}}=5.

2 Datasets

Obman Dataset is a synthetic dataset, which includes hand-object mesh pairs. The hands are generated by a non-learning based method GraspIt! and are parameterized by the MANO model. 2772 object meshes covering 8 classes of everyday objects from ShapeNet dataset are included. The model trained on this dataset will benefit from the diversified object models and grasp types. We train the two networks on this dataset as the initial model.

HO-3D and FPHA Dataset are two real datasets for studying hand-object interaction, and we use them for evaluating the generalization ability of our proposed framework. These two datasets collect video sequences annotated with object-hand poses. Because only a dozen of objects are included in these two datasets, they are not suitable for training the model. Besides, the objects in these two datasets have larger scales. We use the same split and data filtering of the two datasets with .

3 Evaluation Metrics

Penetration is measured by penetration depth and volume between objects and generated grasps following . We voxelize the hand-object meshes with voxel size 0.5 cm0.5\ cm, and calculate the intersection shared by the two 3D voxels.

Grasp displacement is used to measure the stability of the grasp. To test stability, we put the object and generated grasp in a simulator following . In general, the simulator calculates the motion of the object under the grasp. Specifically, the simulator calculates forces on fingertips, which are has a positive correlation with the penetration volume on fingertips. Then, it applies the calculated forces to hold the object against its gravity. The grasp stability is measured by the displacement of the object’s center of mass during a period in the simulation. In this period, the pose and location of hand is fixed. We measure the mean and variance of the simulation displacement for all test samples. Examples with smaller simulation displacement have better grasp stability. The correlation between the grasp stability and penetration of the grasp is discussed in Sec. 4.6.

Perceptual score is utilized for evaluating the naturalness of generated grasps. We perform the perceptual study following with Amazon Mechanical Turk.

Hand-object contact metrics are used for analyzing contact between hand-object. We calculate the sample-level hand-object contact ratio, individual object and hand contact points ratio, and the number of hand fingers contacting the object. We classify the contact status of a point by judging whether its distance to its nearest neighbor in the other point cloud is smaller than 0.5 cm0.5\ cm. We also calculate the object contact map score, as s=100⋅∑ΩN∈s=100\cdot\frac{\sum\Omega}{N}\in, which reflects the coverage area of grasps. Generally, larger contact areas can imply a better grasp, but this is not strictly correct.

4 Grasp Generation Performance

Qualitative results. We first visualize generated grasps for different objects. Fig. 7 shows that our framework is able to generate stable grasps with natural hand poses on both in-domain and out-of-domain objects. By sampling different object poses for the same object as inputs, our model can generate diverse grasps. Fig. 8 shows 5 different grasps generated by our model for each object in each row.

Quantitative results. The evaluation results on the three datasets are shown in Table 1. We train the models on the Obman training set, and test on Obman testset. We also test the model trained from the Obman training set extensively on HO-3D and FPHA to demonstrate the generalization ability of our method. All results are evaluated after Test-Time Adaptation (TTA). The objects in the Obman test set may overlap with its training set, while objects (with different poses) from HO-3D and FPHA are never seen in training.

On all of the three datasets, our framework shows significant improvement over the state-of-the-art approach in both of physical plausibility, grasp stability and perceptual score. And results on HO-3D and FPHA dataset imply that our model has a much stronger cross-domain generalization ability. For instance, achieves reasonably good stability on HO-3D and FPHA but suffers from huge penetration (they are correlated). However, our model performs much better on both of the two metrics with a great balance. Moreover, the perceptual scores of our framework on the three datasets are similar: 3.543.54 for in domain objects in Obman, 3.503.50 and 3.573.57 for out-of-domain data in HO-3D and FPHA. This shows the quality of generated grasps on out-of-domain objects are close to the in-domain objects. Besides, our results are close to or even outperform the ground truth, especially for the stability and perceptual score. These results imply our method’s capability to generating natural, physically plausible and steady grasps.

5 Ablation Study

We first perform ablation studies on Obman dataset for evaluating the two proposed losses LH\mathcal{L_{H}} and LO\mathcal{L_{O}}. We then analyse different designs of ContactNet. Finally, we compare different Test-Time Adaptation (TTA) paradigms on out-of-distribution HO-3D and FPHA dataset.

The results are shown in Table 2. With the hand-centric loss LH\mathcal{L_{H}}, the simulation displacement decreases and contact metrics grows significantly while the penetration grows slightly. After adding the object-centric loss LO\mathcal{L_{O}}, only object-related metrics, e.g. contact object vertices ratio and contact map score, and stability grows (displacement decreases). This implies that with LH\mathcal{L_{H}}, the LO\mathcal{L_{O}} acts as a regularizer on the object contact regions to improves the grasp stability, which matches the design of this loss function.

We also verify the effectiveness of two losses by comparing each with a modified version. First, we can force the fingers to touch the ground truth contact regions rather than finding them dynamically with LH\mathcal{L_{H}}. This loss is denoted as LH(gt)\mathcal{L_{H}}(gt). Experiments demonstrate that LH\mathcal{L_{H}} is better in all metrics than LH(gt)\mathcal{L_{H}}(gt). This implies that fitting the ground truth in the multi-solution grasp generation task may not be optimal. Second, in the loss LO\mathcal{L_{O}}, we verify the effectiveness of using the contact map as the representation of hand-object distance. We experiment with directly minimize the residual between predicted and ground truth object-hand distances D^\hat{\textbf{D}} and D without normalizing them into contact maps. We call this loss as LO(dist)\mathcal{L_{O}}(dist). Experiments show that with LO(dist)\mathcal{L_{O}}(dist), the performance even degenerates. The reason is that the LO(dist)\mathcal{L_{O}}(dist) is contributed almost by hand-object point pairs with large distance, while LO\mathcal{L_{O}} focus more on hand vertices close to object surface with the help of normalization.

5.2 ContactNet Designs

We compare three different kinds of ContactNet designs, as shown in Table 3. The first model (object-only) takes solely the object as input, while the second (h-o global) and third (h-o global-local) model take in both of hand and object. The difference between the latter two is to experiment whether using object local features helps predict the contact map by maintaining point permutation information.

Without the hand as an input, predicting object contact map is a very difficult one-to-many mapping. Considering only a small part of object points are in contact, the 0.1610.161 absolute error is actually huge. Experiments also show that without object local features, the gain from adding the hand as one of the input is trivial. With object local features, the error reduces 0.070.07 as a significant improvement of 50%50\%.

5.3 Test-Time Adaptation (TTA) for Generalization

TTA (offline): Learning-based TTA same as illustrated in Sec. 3.3, and network parameters are re-initialized before adapting each sample;

TTA-optm (offline): Optimization-based TTA, where the MANO parameters are directly optimized;

TTA-noise (offline): Learning-based TTA. When training ContactNet, the hand parameters are injected with random noise. The method is used to compare different methods to obtain the target contact map;

TTA-online: Learning-based TTA, and the network parameters are re-initialized only once for each sequence.

As shown in Table 4, on the two datasets, all TTA methods can improve the results. There are three comparisons between the different methods. First, on the HO-3D dataset, the TTA and TTA-optm achieve comparable results because they are both offline methods using the same objective function. The results of the learning-based TTA are slightly better, which can be explained by that the network parameters serve as a regularization and make the adaptation more steady. Second, training with injected noise, we expect the ContactNet can learn to predict ideal contact maps as target in TTA by "correting" the noise. However, the results deteriorate compared with the one trained on perfect ground truth data. This can be explained by: (i) Injecting noise hurts learning contact maps; (ii) It is hard to match the random noise with the noise pattern in initially predicted grasps. Third, the online version of TTA is much stronger than offline versions. Inspired by , for the online TTA, the target of the TTA can be optimized continually with the help of network parameters, and the model can fit the test distribution better. With online updating, the stability grows and penetration depth decreases simultaneously, indicating that the network leans better hand-object contact. A huge improvement in contact ratio also verifies the point. The improvement of online TTA on the FPHA dataset is not as big as on HO-3D dataset because the average video sequence length is 120\frac{1}{20} of HO-3D dataset, so the learning target cannot be optimized continually.

To show the effectiveness of TTA for improving both the naturalness and stability of generated grasps, we further visualize the grasps and object contact maps before and after TTA. As shown Fig. 10, after TTA, the hand penetration decreases with fingers closely contacting object surface. In Fig. 10, the object contact regions become larger, which indicates the grasps are more stable.

6 Penetration Volume vs. Grasp Displace.

Larger penetration volume can cause better grasp stability (reflected by smaller simulation displacement) during the simulation. However, ideal grasps should be with small penetration and simulation displacement simultaneously, rather than achieving reasonable stability by suffering from huge penetration volume. Thus, we draw Fig. 11 for demonstrating the balance between them on the three datasets. Overall, our results are very close the origin point, which demonstrated our generated grasps has both small penetration and superior stability at the same time. With TTA, the results move vertically in the figure, indicating the TTA is able to increase grasp stability without magnifying the penetration at the same time. Besides, the results are comparable to or even outperform the ground truth.

Conclusion

In this work, we propose a framework for generating human grasps given an object. To get natural and stable grasps, We reason the consistency of contact information between object and generated hand from two aspects: First, we design two novel training targets from the view of hand and object respectively, which helps them to find a mutual agreement on the form of contact. Second, we design two networks for grasp generation and predicting contact map respectively. We leverage the consistency between outputs of the two networks for designing a self-supervised task, which can be used at test-time for adapting generated grasps on novel objects. With the proposed method, we not only observe more natural and stable generated grasps, but also a strong generalization capability on cross-domain test inputs.

Acknowledgements. This work was supported, in part, by grants from DARPA LwLL, NSF 1730158 CI-New: Cognitive Hardware and Software Ecosystem Community Infrastructure (CHASE-CI), NSF ACI-1541349 CC*DNI Pacific Research Platform, and gifts from Qualcomm and TuSimple.

References

Appendix A: Network Architectures

We show structures of GraspCVAE and ContactNet in the following sections.

Table 5 and Table 6 show the architecture of GraspCVAE during training and testing respectively. The input of the two phase are different. During training, the input is both of hand and object point cloud and we train the network in a hand reconstruction manner. During testing, the only input is the object point cloud, and the network generates human hand mesh for grasping the object.

For training, we use two PointNet encoders to get features of hand and object point cloud as Fh\mathcal{F}^{h} and Fo\mathcal{F}^{o}. Then, they are concatenated and sent to the CVAE encoder for predicting the posterior distribution Q(z∣μ,σ2)Q(z|\mu,\sigma^{2}). Then, a latent code zz is sampled from this distribution, and concatenated with the object feature Fo\mathcal{F}^{o} as the input of CVAE decoder for regressing the MANO parameters. In the end, the parameters pass the MANO layer, where the output is the generated hand mesh M^\hat{\mathcal{M}}.

For testing, the latent code zz is randomly sampled from the standard Gaussian distribution. Thus, we do not need the CVAE encoder and the hand point cloud.

A.2. ContactNet

Table 7 shows the architecture of ContactNet, which takes in both hand-object point cloud to regress the object contact map. In the network, we use both of the object global and local features, Fgo\mathcal{F}^{o}_{g} and Flo\mathcal{F}^{o}_{l}, where the local features are used to maintain the point correspondence.

Appendix B: Details of Experiments and Evaluation

We follow to use HO-3D and FPHA datasets for evaluating the generalization ability of the proposed method. For FPHA dataset, the ground-truth hand mesh are fitted on the provided hand joints. We follow to exclude the huge objects (especially milk bottle) in the FPHA dataset.

B.2. Evaluation Metrics

Perceptual score. The perceptual score is evaluated with Amazon Mechanical Turk following , the layout is shown in Fig. 12. We show 3 views of each sample. The rating score ranges from 1 to 5. Every sample is rated by 3 workers.

Penetration. The penetration is to measure the collision between the hand and the object. We report the maximum penetration depth and penetration volume following . The former is calculated as the largest distance from the penetrating vertices of hand mesh to the closed object surface. And the latter is the volume of the intersecting voxels between the hand and object meshes. To compute this metric, we first voxelize both hand and object mesh using the voxel size of 0.5 cm0.5\ cm, and then compute the number of intersecting voxels. The result is computed by the voxel volume times the number of intersecting voxels.

Reconstruction Error. We do not use hand reconstruction error (mesh reconstruction error on hand mesh, or kinematics error on hand joints) as a metric for evaluating the quality of generated grasps. Because grasp generation has multiple solutions, a good and reasonable grasp can be far away from the GT (Note that only one GT grasp is provided for each sample in datasets we used). Thus it does not make sense in our case to measure the reconstruction error with only one GT.

B.3. Experiments

We introduce more experiments details and results in this section, including details of GraspCVAE training targets and Test-time Adaptation.

Table 8 shows the performance of the GraspCVAE trained by losses we proposed and losses from , which are tested on the Obman test set. Our training targets performs significantly better.

B.3.2. Test-time Adaptation

Details of TTA During TTA, each test sample is adapted in a self-supervised manner for 1010 iterations. For each iteration, the single test object is augmented into a batch which includes 3232 samples, where the augmentation is random translation in  cm\ cm. Due to the reason that the augmentation is supposed to maintain the geometry feature of the object, other augmentation methods, e.g. scaling and rotation, are harmful.

Details of Different TTA Paradigms In the Sec. 4.5.3 of the paper, we compare different TTA paradigms. And we give more details here.

TTA-optm (offline): In this method, we only optimize the 45-D hand joints axis-angle rotation tensor, rather than the 61-D full hand pose parameters as in other learning-based TTA. We observe that optimizing the 61-D full hand pose is not stable, and the results can even become worse.

TTA-noise (offline): In this method, when we train the ContactNet, we injecting random noise on the input 45-D hand joint rotation tensor. The model is denoted as ContactNet-noise. The reconstruction error of ContactNet-noise is 0.109, higher than the original 0.090 without injecting noise (Table 4 in paper). The increased error demonstrate that injecting noise is harmful for learning contact maps, and implies that the network cannot learn to "corret" the noise. It is also the reason for the worse results of TTA-noise compared with original TTA.

TTA-online: The HO-3D and FPHA datasets are video datasets, and the TTA-online is performed on the video clips. Because the object pose changes smoothly in the video frames, it provides the chance for the network to fit the test distribution continuously better. Besides, the TTA-online also demonstrate that the model after TTA does not overfit to the single test sample, because it can continually generates grasps of the following incoming test samples without re-initializing the network parameters.

Appendix C: Additional Results

More visualization are shown in Fig. 13 for in-domain Obman test set objects, and Fig. 14 for out-of-domain HO-3D objects. Each result is shown in a row. All results are chosen randomly.

More results are shown in Fig. 15 for in-domain Obman test set objects, and Fig. 16 for out-of-domain HO-3D and FPHA objects. We show 4 examples in each row, and each result is shown with 3 views. All results are chosen randomly.