DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Objects Based on Simulation

Ruicheng Wang, Jialiang Zhang, Jiayi Chen, Yinzhen Xu, Puhao Li, Tengyu Liu, He Wang

I INTRODUCTION

Robotic object grasping is an important technology for many robotic systems. Recent years have witnessed great success in developing vision-based grasping methods and large-scale datasets for parallel-jaw grippers, e.g., synthetic object-centric dataset, ACRONYM , and real-world dataset of grasping in clutter, GraspNet .

Although simple and effective for pick-and-place, parallel-jaw grippers show certain limitations in dexterous object manipulation, e.g., using scissors, due to their low DoFs. On the contrary, multi-fingered robotic hands, e.g., ShadowHand , are human-like, designed with very high DoFs (26 for ShadowHand), and can attain more diverse grasp types. Those dexterous hands can support many complex and diverse manipulations, e.g., solving Rubik’s cube , and can be used in task-specific grasping .

Arguably, dexterous grasping is the first step to dexterous manipulation. However, dexterous grasping is highly under-explored, compared to parallel grasping. One major obstacle is the lack of large-scale robotic dexterous grasping datasets required by learning-based methods. Up to now, the only dataset is provided by Liu et al. (Deep Differentiable Grasp, referred to as DDG), which contains only 6.9K grasps and 565 objects and is much smaller than the grasp datasets for parallel grippers, e.g., GraspNet , ACRONYM . Considering the high-DoF nature of the dexterous hand, dexterous grasping datasets need to be significantly larger and more diverse for the sake of generalization.

In this work, we propose DexGraspNet, a large-scale simulated dataset for robotic dexterous grasping. This dataset contains 1.32 million dexterous grasps for ShadowHand on 5355 objects, with more than 200 diverse grasps for each object instance. The objects are from more than 133 hand-scale object categories and collected from various synthetic and scanned object datasets. In addition to the scale, our dataset also features high diversity and high physical stability. All grasps have been examined by force closure and validated by Isaac Gym physics simulator, enabling further tasks in both real-world and simulation environments.

Note that synthesizing diverse high-quality dexterous grasps at scale is known to be very challenging. For dexterous grasping data synthesis, previous works, e.g., DDG, mainly use GraspIt! , which lacks diversity in grasping poses due to its naive search strategy. A recent work proposes a novel method to address this diversity issue. This work devises a differentiable energy term to approximate force closure and then uses it to synthesize diverse and stable grasps via optimization. However, suffers from low yield, slow convergence, and strict constraints on object meshes, making it infeasible for us to use for synthesizing a large-scale dataset.

To achieve our desired diversity, quality, and scale, we propose several critical improvements to , making it much more efficient and robust. First, we design a better hand pose initialization strategy and carefully select contact candidates to boost yield. For synthesizing 10000 valid grasps, we speed up from 400 GPU hours to 7 GPU hours. Second, we propose an alternative way to compute penetration energy and signed distances, which enables us to handle object meshes of much lower quality, and also highly simplifies their preprocessing procedures. Third, we introduce energy terms that punish self-penetration and out-of-limit joint angles to further improve grasp quality. Additionally, with simple modifications, the entire pipeline can be applied to other dexterous hands, such as MANO and Allegro.

To verify the advantage of our dataset over the one from DDG, we train two dexterous grasping algorithms on our dataset and DDG. The cross-dataset experiments confirm that training on our dataset yields better grasping quality and higher diversity. Also, the great diversity of the hand grasps from our dataset leaves huge improvement space for future dexterous grasping algorithms.

II RELATED WORK

Researches in grasping can be broadly categorized by the types of end effectors involved. The most thoroughly studied ones are the suction cup and parallel jaw grippers, whose grasp pose can be defined by a 7D vector at most, including 3D for translation, 3D for rotation, and 1D for the width between the two fingers. Dexterous robotic hands with three or more fingers such as ShadowHand and humanoid hands such as MANO require more complex descriptors, sometimes up to 24DoF as in ShadowHand . In this paper, we are dedicated to researches on the latter type. To bridge the gap between humanoid hands and robotic hands, numerous researches have shown the efficacy of retargeting humanoid hand poses to dexterous robotic hands .

Early researches in dexterous grasping focus on optimizing grasping poses to form force closure that can resist external forces and torques .

Due to the complexity of computing hand kinematics and testing force closure, many works were devoted to simplifying the search space . As a result, these methods were applicable to restricted settings and can only produce limited types of grasping poses. Another stream of work looks for simplifying the optimization process with an auxiliary function. proposed to use a differentiable estimator of the force closure metric to synthesize diverse grasping poses for arbitrary hands.

II-B Data-Driven Grasping

Recent works shift their focus to data-driven methods. Given an object, the most straightforward approach is to directly generate the pose vectors of the grasping hand . A refinement step is usually implemented in these methods to remove inconsistencies such as penetration.

Other methods take an indirect approach that involves generating an intermediate representation first. Existing methods use contact points , contact maps , and occupancy fields as the intermediate representations. The methods then obtain the grasping poses via optimization , planning , RL policies , or another generative model .

Compared to most analytical methods, data-driven methods show improved inference speed and diversity of generated grasping poses. However, the diversity is still limited by the training data.

II-C Dexterous Grasp Datasets

Dexterous grasping is impossibly difficult to annotate for its overwhelming degrees of freedom. Most existing works are trained on programmatically synthesized grasping poses using the GraspIt! planner. The planner first searches the eigengrasp space for pregrasp poses that cross a threshold. Then, the planner squeezes all fingers in the selected pregrasp poses to construct a firm grasp. Since the initial search is performed in the low dimensional eigengrasp space, the resulting data follows a narrow distribution and cannot cover the full dexterity of multi-finger hands.

More recent works leverage the increasing capacity of computer vision to collect human hand poses when interacting with the object. HO3D computes the ground truth 3D hand pose for images from 2D hand keypoint annotations. The method resolves ambiguities by considering physics constraints in hand-object interactions and hand-hand interactions. DexYCB and ContactPose solve the 3D hand shape from multi-view RGBD camera recordings. Latest datasets use optical motion capture systems to track hand and object shapes during interactions. While these methods produce natural and smooth interaction demonstrations, the data is restricted to humanoid hand structures and daily hand poses.

In addition, ContactDB and ContactPose leverage IR cameras to collect contact maps on object surfaces.

III DATASET GENERATION METHOD

We collect object models from various datasets. ShapeNetCore and ShapeNetSem contain Computer-Aided-Design (CAD) models with category labels, from which we select 3980 objects in 133 categories. YCB , BigBIRD , Grasp , KIT , and Google’s scanned Object Dataset are scanned model repositories without category labels, from which we select 1375 objects. They are not labeled with categories in our dataset.

Since not all objects from these datasets are aligned to sizes in the real world, we choose to normalize all models into a unit sphere and then augment each object by uniformly scaling them with 5 fixed sizes. Then we remesh them into manifolds to make closed figures. Finally, for simulation purposes, we create collision meshes for every object mesh through convex decomposition using CoACD .

III-B Grasp Generation

III-B2 Review of Differentiable Force Closure

Our dexterous grasp generation method is mainly built upon the original work , in which they propose a novel differentiable force closure estimator as an energy term and use optimization to synthesize grasps. The proposed differentiable force closure term, which encourages a set of contact points to form force closure, can be expressed as

A pair of attraction and repulsion energy functions are introduced to ensure contact and prevent penetration:

where S(H)S(H) is the surface point cloud of the hand mesh HH, d(⋅,⋅)d(\cdot,\cdot) is the point-to-mesh distance and [v∈O]=1[v\in O]=1 if point vv is inside object mesh OO. They also use an energy function EpriorE_{\rm prior} to keep the hand in a natural state. The complete energy function is as follows:

They design a modified MALA optimization algorithm to minimize the energy EE over the augmented grasp tuple g′=(T,R,θ,x)g^{\prime}=(T,R,\theta,x). The algorithm takes an initial hand pose g0′g^{\prime}_{0}, which is randomly initialized. Then, in each iteration, T,R,θT,R,\theta are updated according to Langevin dynamics, and contact points xx are randomly re-sampled with a small probability. The update is accepted or rejected stochastically by the Metropolis-Hastings rule. The optimization ends after 10000 steps. For more details, please refer to .

III-B3 Our Method

Though takes a great step forward, it is still quite hard to obtain a large-scale grasp dataset directly using their method due to three reasons: 1. the algorithm suffers from a low success rate and slow convergence; 2. most object meshes we use have no thickness due to the poor quality of the object dataset, making it impossible to compute penetration energy; 3. due to random initialization, some generated grasping poses may look twisted. To overcome these issues, we propose several ways to improve the efficiency, effectiveness, and robustness of the original algorithm:

First, we propose an initialization strategy that can greatly improve the success rate and speed up the convergence. We find that their optimization algorithm’s success rate drops dramatically when the initial hand is closed. This motivates us to introduce two constraints to the initialization strategy: 1. the five fingers should be opened to form a space for grasping; 2. the palm should face the object.

More specifically, we manually choose a canonical hand pose θref\theta_{\rm ref}, as shown in Figure 2, then jitter each joint angle within its limit using the truncated normal distribution. Then, on each object mesh, we first take its convex hull, then push every vertex of the hull away from the origin by 0.2m0.2m to obtain the inflated convex hull. Next, we sample a random point pp on the surface of the inflated convex hull, and compute the direction vector from pp to its nearest point on the original object mesh, then jitter this direction vector within a cone and get n⃗\vec{n}. Finally, the hand is moved to pp, and rotated to face the same direction as n⃗\vec{n}, then push away from the object mesh along n⃗\vec{n} by a random distance, and rotated around n⃗\vec{n} randomly.

The resulting initial hand pose (T0,R0,θ0)(T_{0},R_{0},\theta_{0}) can be easily optimized to a grasp pose, thus raising the success rate. Moreover, for each object, if we sample enough initial hands, they can surround the object densely and evenly, so we can generate diverse data. Also, we found that using our strategy, the final grasp poses look more natural.

Second, we propose an alternative way to compute penetration energy that can make the algorithm more robust to thin object meshes of low quality. In practice, the original algorithm will fail completely when the object mesh has no thickness because the penetration energy will always be zero. Therefore, instead of taking the point cloud from the hand, we take it from the object and compute each point’s distance to the hand mesh. We call this the reverse penetration energy. It does not require the object mesh to have any thickness at all, allowing us to process far more object CAD models.

Third, we modify EpriorE_{\rm prior}, inspired by to penalize self penetration and out-of-limit joint angles:

where wdis=100,wpen=100,wspen=10,wjoints=1w_{\rm dis}=100,w_{\rm pen}=100,w_{\rm spen}=10,w_{\rm joints}=1.

Another minor difference between our implementation and lies in the optimization algorithms. Because our initialization strategy has already raised the success rate of the algorithm to an acceptable level, we simplify MALA and use simple gradient descent to update T,R,θT,R,\theta during each optimization step. In our settings, the optimization process converges in less than 6000 iterations, which reduces almost half of the original iterations.

Finally, we use a modified version of Kaolin instead of DeepSDF to compute point-to-mesh signed distances, which eliminates the need for pretraining category level DeepSDF networks, and significantly reduces the memory cost when optimizing grasps.

III-C Grasp Validation

To filter out those bad results after the optimization converges, we validate all of the grasps in a physical simulator Isaac Gym with PhysX as the basic physics engine. We first initialize the gripper using the final grasp parameters. Then, in order to apply active forces on the object, we slightly move each contacting link of the gripper along the normal of its contact point, and set the moved pose as target positions for position control. Finally, gravity with a magnitude of 9.8m/s29.8m/s^{2} is added to the scene. A grasp is considered successful if the gripper is still in contact with the object after 100 simulation steps under all 6 axis-aligned directions of gravity. The distribution of the object number with respect to the average success rate for each object is shown in Fig. 3. Moreover, if the max penetration depth exceeds 0.1cm, we also consider the grasp as a failure. We only save those grasps who pass both the simulation validation and the penetration validation in our dataset.

IV Dataset Analysis and Comparison

With our improved pipeline, we generate more than 200 grasps per object, which sum up to 1.32 million grasps in total, forming the largest grasping dataset for ShadowHand. Some visual qualitative results are shown in Fig. 5. Additionally, this pipeline can be stably applied to other dexterous hands. Fig. 4 shows some synthesized results for human hand (MANO) and Allegro along with ShadowHand.

Compared to the original algorithm , our improved pipeline achieves a significant speed-up. On NVIDIA A100 with 19.49 TFLOPS, our algorithm takes 74min to optimize 10000 grasps for 6000 steps, out of which about 18% are considered valid under our settings. The original algorithm of takes 37min on NVIDIA 3090 with 35.58 TFLOPS to optimize 512 grasps for 10000 steps, out of which about 3% are considered valid. It took us 950 GPU hours on A100 to generate 1.32 million valid grasps, which would have taken the original algorithm 50000 GPU hours. This speed-up is contributed by faster convergence, smaller memory (which leads to bigger batch size), and a higher success rate.

We demonstrate the quality of our dataset by comparing DexGraspNet with the dataset proposed in DDG (in short DDGdata), a grasping dataset for ShadowHand generated by GraspIt! , on two aspects below.

First, we conclude that DexGraspNet is more diverse. As shown in Fig. 6, the planner in GraspIt! can only clench each finger in a fixed direction, so the root joints of every finger lose one DoF each. Moreover, the angles of many joints in DDGdata often collapse to their upper or lower limits due to their simple generation strategy. These phenomena contribute to a serious loss of diversity. In contrast, our optimization method can generate diverse grasps with much higher dexterity, as shown in Fig. 5. We further use the mean entropy to model the diversity quantitatively. To evaluate this metric, we first discretize each joint’s motion range into 100100 bins, then use samples from each dataset to estimate a probability distribution, calculate the entropy of these distributions, and take the mean over all joints. Results are shown in Table II.

Second, we show that DexGraspNet is more stable by comparing the Q1Q_{1} metric , which is intuitively the norm of the smallest wrench that can destabilize the grasp:

where {wi}\{w_{i}\} are contact friction cone wrenches. We choose 1mm1{\rm mm} as the contact threshold, and allow at most one contact point for each link to save computational time. The results are shown in Table II. We can find that the average value of our dataset is significantly better than that of DDGdata. It is worth noting that our entire generation pipeline does not explicitly optimize these metrics.

V Benchmarks

We benchmark two methods of dexterous grasp synthesis, DDG and GraspTTA, on our dataset, and compare them with the same methods trained on DDGdata.

DDG designs a differentiable Q1Q_{1} metric, which generalizes the standard Q1Q_{1} metric to the case when the gripper is not in contact with objects. With this generalized Q1Q_{1} metric, they are able to supervise the neural network to predict fine grasp end-to-end. Their network takes 5 depth images of the object as input and directly regresses 6D pose and joint angles of the ShadowHand. To ease learning, they divide the training process into two stages. In the first stage, they only use the loss of the grasp poses, and in the second stage, they fine-tune the network with differentiable Q1Q_{1} loss and other losses to avoid penetration and pull the hand closer to the object. We follow their data pre-process pipeline to generate BVH representations and depth images of our objects, and train the network with official settings on each dataset.

Another work GraspTTA (short for Test-Time Adaptation) proposes to synthesize high-quality grasps by ensuring the contact consistency between the hand and the object. They design two networks, one is a CVAE to synthesize grasps, and the other is a contact net to predict contact regions of the object. During training, those two networks are trained separately. During testing, a grasp is synthesized in a two-stage process. First, the CVAE takes the object point cloud as condition, samples a latent code, and then decodes the hand 6D global pose and joint angles, which can be further transformed into the hand point cloud through forward kinematics. Second, the contact net takes both the object and the hand point cloud to predict a target contact map, and optimizes the hand parameters to minimize the difference between the current contact map and the target contact map. We re-implement GraspTTA on ShadowHand, process the data from DexGraspNet and DDGdata in the same way as in , and train the networks on each dataset for the same number of iterations.

V-B Experiments and Results

We report the following metrics for evaluation. 1) Simulation success rate(%) in Isaac Gym. We adopt an easier criterion (the criterion in Sec. III-C is too strict for the baselines): a grasp is considered valid if it can withstand at least one of the six gravity directions and has a maximal penetration less than 5mm5{\rm mm}. 2) Mean Q1Q_{1} , which is introduced in Sec. IV. Since these methods cannot guarantee exact contact, we relax the contact threshold to 1cm1{\rm cm}. Particularly, if the penetration depth is greater than 5mm5{\rm mm}, the Q1Q_{1} metric is not well defined, so we manually set Q1Q_{1} of these results to 0. 3) Maximal penetration depth(cm). This is defined as the maximal penetration depth from the object point cloud to hand meshes.

The main results are presented in Table III. Comparing models trained on DexGraspNet with models trained on DDGdata, we observe that no matter which baseline, test set, or metric we use, the former always scores higher. We thus conclude that learning-based grasping methods achieve higher performance when they are trained under our dataset. Table III also shows that the output of DDG has higher quality than GraspTTA most of the time. More specifically, GraspTTA suffers severely from penetration.

Apart from grasp quality, we use joint angle entropy (the same as in Section IV) to evaluate the diversity of the grasps generated by the two methods. Table IV shows the joint angle entropy of models trained on DexGraspNet always have higher means and lower standard deviations than models trained on DDGdata, which means DexGraspNet improves the diversity of grasping methods. We also find that GraspTTA has higher diversity than DDG, partly due to the test time optimization used in GraspTTA that can generate many variations. To compare the joint angle entropy of models trained on DexGraspNet and the original joint angle entropy of DexGraspNet, we further find that: 1) DDG’s entropy is lower than DexGraspNet’s entropy, meaning DDG cannot fully recover DexGraspNet’s diversity; 2) although GraspTTA yields an entropy higher than DexGraspNet, this does not necessarily mean that GraspTTA learns diverse grasping, given its success rate is very low. We interpret this as the trade-offs that DDG and GraspTTA individually make, given that stability and diversity are contradictory to some degree. More importantly, this status quo shows that none of the existing grasping methods can fully learn the highly diverse grasp poses of DexGraspNet while keeping a reasonable success rate at the same time.

VI LIMITATIONS

By comparing grasps in our dataset with the taxonomy from , we notice that our dataset cannot cover every grasping type described. Since the optimization step tends to pull every candidate point closer to the object, the final grasps are always contact-rich, or power grasps. Therefore, precision grasps hardly appear, which represents the dexterity of multi-finger robotic hands. Additionally, our method lacks semantic guidance, which makes it hard to generate functional grasps, e.g. picking up the mug by its handle. Precision grasps and functional grasps remain important issues for us to explore.

VII CONCLUSIONS

In this paper, we present a large-scale synthetic dexterous grasping dataset, DexGraspNet, synthesized via our proposed deeply-accelerated optimization-based method. This dataset has a much larger scale, better grasp quality, and higher diversity than previous datasets. Trained on DexGraspNet, previous grasp synthesis methods can achieve consistent improvements in both quality and diversity. However, none of the existing methods can perform well on both metrics. Compared to grasping using parallel grippers, we argue that dexterous grasping has a larger room for research. We release DexGraspNet and hope that its scale, quality, and diversity can help future methods tackle the task of dexterous grasping, and exploit more potential of dexterous grippers.

VIII Acknowledgements

This work is supported in part by the National Key R&D Program of China (2022ZD0114900) and the Beijing Municipal Science & Technology Commission (Z221100003422004).

References