GenDexGrasp: Generalizable Dexterous Grasping

Puhao Li, Tengyu Liu, Yuyang Li, Yiran Geng, Yixin Zhu, Yaodong Yang, Siyuan Huang

I Introduction

Humans’ ability to grasp is astonishingly versatile. In addition to the full grasp with five fingers, humans can efficiently generalize grasps when some fingers are occupied and imagine diverse grasping poses for various downstream tasks when given an unseen new type of hand, all happened rapidly with a high success rate. These criteria starkly contrast with most prior robot grasping methods, which primarily focus on specific end-effectors, requiring redundant efforts to learn the grasp model for every new robotic hand. On top of this challenge, prior methods often have difficulties quickly generating diverse hand poses for unseen scenarios, further widening the gap between robot and human capabilities. Hence, these deficiencies necessitate a generalizable grasping algorithm, efficiently handling arbitrary situations and allowing fast prototyping for new robots.

Fundamentally, the most significant challenge in generalizable dexterous grasping is to find an efficient and transferable representation for diverse grasp. The de facto representation, joint angles, is unsuitable for its dependency on the structure definition: two similar robotic hands could have contrasting joint angles if their joints are defined differently. Existing works use contact points , contact maps , and approach vectors as the representations, and execute the desired grasps with complex solvers. A simple yet effective representation is still in need.

In this paper, we denote generalizable dexterous grasping as the problem of generating grasping poses for unseen hands. We evaluate generalizable grasping in three aspects:

Speed: Hand-agnostic methods adopt inefficient sampling strategies , which leads to extremely slow grasp generation, ranging from 5 minutes to 40 minutes.

Diversity: Hand-aware methods rely on deterministic solvers, either as a policy for direct execution or predicted contact points for inverse kinematics, resulting in identical grasping poses for the same object-hand pair.

Generalizability: Hand-aware methods also rely on hand descriptors trained on two- and three-finger robotic hands, which hinders their generalizability to new hands that are drastically different from the trained ones.

To achieve a three-way trade-off among the above aspects and alleviate the aforementioned issues, we devise GenDexGrasp for generalizable dexterous grasping. Inspired by Brahmbhatt et al. , we first generate a hand-agnostic contact map for the given object using a conditional variational autoencoder . Next, we optimize the hand pose to match the generated contact map. Finally, the grasping pose is further refined in a physics simulation to ensure a physically plausible contact. GenDexGrasp provides generalizability by reducing assumptions about hand structures and achieves fast inference with an improved contact map and an efficient optimization scheme, resulting in diverse grasp generation by a variational generative model with random initialization.

To address contact ambiguities (especially for thin-shell objects) during grasp optimization, we devise an aligned distance to compute the distance between surface point and hand, which helps to represent accurate contact maps for grasp generation. Specifically, the traditional Euclidean distance would mistakenly label both sides of a thin shell as contact points when the contact is on one side, whereas the aligned distance considers directional alignment to the surface normal of the contact point and rectifies the errors.

To learn the hand-agnostic contact maps, we collect a large-scale multi-hand dataset, MultiDex, using force closure optimization . MultiDex contains 436,000 diverse grasping poses for 5 hands and 58 household objects.

We summarize our contributions as follows:

We propose GenDexGrasp, a versatile generalizable grasping algorithm. GenDexGrasp achieves a three-way trade-off among speed, diversity, and generalizability to unseen hands. We demonstrate that GenDexGrasp is significantly faster than existing hand-agnostic methods and generates more diversified grasping poses than hand-aware methods. Our method also achieves strong generalizability, comparable to existing hand-agnostic methods.

We devise an aligned distance for properly measuring the distance between the object’s surface point and hand. We represent a contact map with the aligned distance, which significantly increases the grasp success rate, especially for thin-shell objects. The ablation analysis in Tab. II shows the efficacy of such a design.

We collect and open-source a large-scale synthetic dataset, MultiDex, for generalizable grasping with 5 robotic hands, 58 household objects, and 436,000 diverse grasping poses. MultiDex is by far the largest multi-hand grasp dataset with diverse hand structures.

II Related Work

Existing solutions to generalizable grasping fall into two categories: hand-aware and hand-agnostic. The hand-aware methods are limited by the diversity of generated poses, whereas the hand-agnostic methods are oftentimes too slow for various tasks. Below, we review both methods in detail.

Hand-aware approaches learn a data-driven representation of the hand structure and use a neural network to predict an intermediate goal, which is further used to generate the final grasp. For instance, UniGrasp and EfficientGrasp extract the gripper’s PointNet features in various poses and use a PSSN network to predict the contact points of the desired grasp. As a result, contact points are used as the inverse kinematics’s goal, which generates the grasping pose. Similarly, AdaGrasp adopts 3D convolutional neural networks to extract gripper features, ranks all possible poses from which the gripper should approach the object, and executes the best grasp with a planner. However, all hand-aware methods train and evaluate the gripper encoders only with two- and three-finger grippers, hindering their ability to generalize to unseen grippers or handle unseen scenarios. Critically, these methods solve the final grasp deterministically, yielding similar grasping poses.

Hand-agnostic methods rely on carefully designed sampling strategies . For instance, ContactGrasp leverages the classic grasp planner in GraspIt! to match a selected contact map, and Liu et al. and Turpin et al. sample hand-centric contact points/forces and update the hand pose to minimize the difference between desired contacts and actual ones. All these methods adopt stochastic sampling strategies that are extremely slow to overcome the local minima in the landscape of objective functions. As a result, existing hand-agnostic methods take minutes to generate a new grasp, impractical for real-world applications.

II-B Contact Map

Contact map has been an essential component in modern grasp generation and reconstruction. Initialized by GraspIt! and optimized by DART , ContactGrasp uses thumb-aligned contact maps from ContactDB to retarget grasps to different hands. ContactOpt uses an estimated contact map to improve hand-object interaction reconstruction. NeuralGrasp retrieves grasping poses by finding the nearest neighbors in the latent space projections of contact maps. Wu et al. samples contact points on object surfaces and uses inverse kinematics to solve the grasping pose. Mandikal et al. treats contact maps as object affordance and learns an RL policy that manipulates the object based on the contact maps. dfc simultaneously updates hand-centric contact points and hand poses to sample diverse and physically stable grasping from a manually designed Gibbs distribution. GraspCVAE and Grasp’D use contact maps to improve grasp synthesis: GraspCVAE generates a grasping pose and refines the pose w.r.t. an estimated contact map, whereas Grasp’D generates and refines the expected contact forces while updating the grasping pose. IBS-Grasp learns a grasping policy that takes an interaction bisector surface, a generalized contact map, as the observed state. Compared to prior methods, the proposed GenDexGrasp differs by treating the contact map as the transferable and intermediate representation for hand-agnostic grasping. We use a less restrictive contact map and a more efficient optimization method for faster and more diversified grasp generation; see detailed in Sec. IV-A.

II-C Grasp Datasets

3D dexterous grasping poses are notoriously expensive to collect due to the complexity of hand structures. The industrial standard method of collecting a grasping pose is through kinesthetic demonstration , wherein a human operator manually moves a physical robot towards a grasping pose. While researchers could collect high-quality demonstrations with kinesthetic demonstrations, it is considered too expensive for large-scale datasets. To tackle this challenge, researchers devised various low-cost data collection methods.

The straightforward idea is to replace kinesthetic demonstration with a motion capture system. Recent works have leveraged optical and visual MoCap systems to collect human demonstrations. Another stream of work collects the contact map on objects by capturing the heat residual on the object surfaces after each human demonstration and using the contact map as a proxy for physical grasping hand pose . Despite the differences in data collection pipelines, these prior arts collect human demonstrations within a limited setting, between pick-up and use. Such settings fail to cover the long-tail and complex nature of human grasping poses as depicted in the grasping taxonomy and grasp landscape . As a result, the collected grasping poses are similar to each other and can be represented by a few principal components . We observe the same problem in programmatically generated datasets using GraspIt! .

III Dataset Collection

To learn a versatile and hand-agnostic contact map generator, the grasp dataset ought to contain diverse grasping poses and corresponding contact maps for different objects and robotic hands with various morphologies.

We selected 58 daily objects from the YCB dataset and ContactDB , together with 5 robotic hands (EZGripper, Barrett Hand, Robotiq-3F, Allegro, and Shadowhand) ranging from two to five fingers. We split our dataset into 48 training objects and 10 test objects. We show a random subset of the collected dataset in Fig. 1.

Given an object OO, a kinematics model of a robotic hand HH with pose qHq_{H} and surface H\mathcal{H}, and a group of nn hand-centric contact points X⊂HX\subset\mathcal{H}, we define the differentiable force closure estimator dfc as:

dfc describes the total wrench when each contact point applies equal forces, and friction forces are neglectable. As established in Liu et al. , dfc is a strong estimator of the classical force closure metric.

Next, we define the prior and penetration energy as

where qH↑{q_{H}}_{\uparrow} and qH↓{q_{H}}_{\downarrow} are the upper and lower limits of the robotic hand parameters, respectively. δ(x,O)\delta(x,O) gives the signed distance from xx to OO, where the distance is positive if xx is outside OO and is negative if inside.

We use a Metropolis-adjusted Langevin algorithm (mala) to simultaneously sample the grasping poses and contact points. We run the malaalgorithm on an NVIDIA A100 80GB with a batch size of 1024 for each hand-object pair and obtain 436,000 valid grasping poses. It takes about 1,400 GPU hours to synthesize the entire dataset.

III-B Contact Map Synthesis

Given the grasping poses, we first compute the object-centric contact map Ω\Omega as a set of normalized distances from each object surface point to the hand surface. Instead of using Euclidean distance, we propose an aligned distance to measure the distance between the object’s surface point and the hand surface. Given the object OO and the hand HH with optimized grasp pose qHq_{H}, we define O\mathcal{O} as the surface of OO and H\mathcal{H} as the surface of HH. The aligned distance D\mathcal{D} between an object surface point vo∈Ov_{o}\in\mathcal{O} and H\mathcal{H} is defined as:

where ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle denotes the inner product of two normalized vectors, and non_{o} denotes the object surface normal at vov_{o}. γ\gamma is a scaling factor; we empirically set it to 1. The aligned distance considers directional alignment with the object’s surface normal on the contact point and reduces contact ambiguities on thin-shell objects. Fig. 2 shows that our aligned distance correctly distinguishes contacts from different sides of a thin shell, whereas the Euclidean distance mistakenly labels both sides as contact regions.

Next, we compute the contact value C(vo,H)\mathcal{C}(v_{o},\mathcal{H}) on each object surface point vov_{o} following Jiang et al. :

where C(vo,H)∈(0,1]\mathcal{C}(v_{o},\mathcal{H})\in(0,1] is 11 if vov_{o} is in contact with H\mathcal{H}, and is if it is far away. C≤1\mathcal{C}\leq 1 since D\mathcal{D} is non-negative.

Finally, we define the contact map Ω(O,H)\Omega(\mathcal{O},\mathcal{H}) as

IV GenDexGrasp

Given an object OO and the kinematics model of an arbitrary robotic hand HH with NN joints, we aim to generate a dexterous, diverse, and physically stable grasp pose qHq_{H}.

Generating qHq_{H} directly for unseen HH is challenging due to the sparsity of the observed hands and the non-linearity between qHq_{H} and hand geometry. Inspired by Brahmbhatt et al. , we adopt the object-centric contact map as a hand-agnostic intermediate representation of a grasp. Instead of directly generating qHq_{H}, we first learn a generative model that generates a contact map over the object surface. We then fit the hand to the generated map.

Inspired by the successful applications of generative models in grasping , we adopt CVAE to generate the hand-agnostic contact map. Given the point cloud of an input object and the corresponding pointwise contact values C\mathcal{C}, we use a PointNet encoder to extract the latent distribution N(μ,σ)\mathcal{N}(\mu,\sigma) and sample the latent code z∼N(μ,σ)z\sim\mathcal{N}(\mu,\sigma). When decoding, we extract the object point features with another PointNet, concatenate zz to the per-point features, and use a shared-weight MLP to generate a contact value C^(vo)\hat{\mathcal{C}}(v_{o}) for each vo∈Ov_{o}\in\mathcal{O}, which forms the predicted contact map Ω^(O)={C^(vo)}vo∈O\hat{\Omega}(\mathcal{O})=\{\hat{\mathcal{C}}(v_{o})\}_{v_{o}\in\mathcal{O}}.

We learn the generative model by maximizing the log-likelihood of pθ,φ(Ω∣O)p_{\theta,\varphi}(\Omega\mid O), where θ\theta and ϕ\phi are the learnable parameters of the encoder and decoder, respectively. According to Sohn et al. , we equivalently maximize the ELBO:

where ZZ is the prior distribution of the latent space; we treat ZZ as the standard normal distribution N(0,I)\mathcal{N}(0,I).

We leverage a reconstruction loss to approximate the expectation term of ELBO:

where NoN_{o} is the number of examples. Ωi\Omega^{i} and Ω^i\hat{\Omega}^{i} denote the expected and generated contact map of the ii-th example, respectively.

Of note, since the generated contact map is empirically more ambiguous than the ground-truth contact map, we sharpen the generated contact map with

IV-B Grasp Optimization

Given the generated contact map Ω^^\hat{\hat{\Omega}} on object OO, we optimize the grasping pose qHq_{H} for hand HH. We initialize the optimization by randomly rotating the root link of the hand and translating the hand toward the back of its palm direction. We set the translation distance to the radius of the minimum enclosing sphere of the object.

We compute H\mathcal{H} by differentiable forward kinematics and obtain the current contact map Ω˙\dot{\Omega}. We compute the optimization objective EE as

Since the computation of the objective function is fully differentiable, we use the Adam optimizer to minimize EE by updating qHq_{H}. We run a batch of 32 parallel optimizations to keep the best result to avoid bad local minima.

IV-C Implementation Details

V Experiment

We quantitatively evaluate GenDexGrasp in terms of success rate, diversity, and inference speed.

Diversity

We measure the diversity of the generated grasps as the standard deviation of the joint angles of the generated grasps that pass the simulation test.

Inference Speed

We measure the time it takes for the entire inference pipeline to run.

We compare GenDexGrasp with dfc , GraspCVAE (GC), and UniGrasp (UniG.) in Tab. I. The columns represent method names, whether the method is generalizable, success rate, diversity, and inference speed. We evaluate all methods with the test split of the ShadowHand data in MultiDex. We trained our method with the training split of EZGripper, Robotiq-3F, Barrett, and Allegro. Since GraspCVAE is designed for one specific hand structure, we train GraspCVAE on the training split of the ShadowHand data and keep the result before and after test-time adaptation (TTA). We evaluate UniGrasp with its pre-trained weights.

Of note, since the UniGrasp model only produces three contact points, we align them to the thumb, index, and middle finger of the ShadowHand for inverse kinematics. In addition, UniGrasp yields zero diversity since it produces the top-1 contact point selection for each object. To evaluate its diversity, we include top-8, top-32, and top-64 contact point selections. We observe that dfc achieves the best success rate and diversity but is overwhelmingly slow. GraspCVAE can generate diverse grasping poses but suffers a low success rate and cannot generalize to unseen hands. We attribute the low success rate to our dataset’s large diversity of grasping poses. The original GraspCVAE was trained on HO3D , where grasp poses are similar since six principal components can summarize most grasping poses. UniGrasp can generalize to unseen hands and achieve a high success rate. However, it fails to balance success rate and diversity.

Our method achieves a slightly lower success rate than dfc and UniGrasp top-1 but can generate diverse grasping poses in a short period of time, achieving an excellent three-way trade-off among quality, diversity, and speed.

We examine the efficacy of the proposed aligned distance in Tab. II. Specifically, we evaluate the success rate and diversity of the full model (full) and the full model with Euclidean distance contact maps (-align). The experiment is repeated on EZGripper, Barrett, and ShadowHand to show efficacy across hands. In all three cases, we observe that using the Euclidean distance lowers the success rate significantly while improving the diversity slightly. Such differences meet our expectations, as contact maps based on Euclidean distances are more ambiguous than those based on aligned distances. During the evaluation, such ambiguities bring more uncertainties, which are treated as diversities using our current metrics. We also observe that the model performs worse on the EZGripper due to the ambiguities in aligning two-finger grippers to multi-finger contact maps.

We further compare the performances of GenDexGrasp on seen and unseen hands in Tab. III. We train two versions of GenDexGrasp for each hand. The in-domain version is trained on all five hands and evaluated on the selected hand. The out-of-domain version is trained on all four hands except the selected hand and evaluated on the selected hand. Our result shows that our method is robust for various hand structures in out-of-domain scenarios.

The qualitative results in Fig. 4 show the diversity and quality of grasps generated by GenDexGrasp. The generated grasps cover diverse grasping types, including wraps, pinches, tripods, quadpods, hooks, etc. We also show failure cases in Fig. 5, where the first three columns show failures from our full model, and the last column shows failures specific to the -align ablation version. The most common failure types are penetrations and floatations caused by imperfect optimization. We observe an interesting failure case in the first example in the bottom row, where the algorithm tries to grasp the apple by squeezing it between the palm and the base. While the example fails to pass the simulation test, it shows the level of diversity that our method provides.

Finally, we demonstrate that our approach can be applied to tabletop objects after proper training; see Fig. 6.

VI Conclusion

This paper introduces GenDexGrasp, a versatile dexterous grasping method that can generalize to unseen hands. By leveraging the contact map representation as the intermediate representation, a novel aligned distance for measuring hand-to-point distance, and a novel grasping algorithm, GenDexGrasp can generate diverse and high-quality grasping poses in reasonable inference time. The quantitative experiment suggests that our method is the first generalizable grasping algorithm to properly balance among quality, diversity, and speed. In addition, we contribute MultiDex, a large-scale synthetic dexterous grasping dataset. MultiDex features diverse grasping poses, a wide range of household objects, and five robotic hands with diverse kinematic structures.

Acknowledgement: This work is supported in part by the National Key R&D Program of China (2021ZD0150200), the Beijing Municipal Science & Technology Commission (Z221100003422004), and the Beijing Nova Program.

References