Deep Differentiable Grasp Planner for High-DOF Grippers

Min Liu, Zherong Pan, Kai Xu, Kanishka Ganguly, Dinesh Manocha

I Introduction

Robot grasping of unknown objects is an important problem and an essential component of various applications, including robot object packing and dexterous manipulation . Earlier methods could generate grasp poses for an arbitrary gripper or target object, but they ignored the uncertainty of real world situations. Recent learning-based methods have demonstrated improved robustness in terms of handling sensor noise. Instead of directly inferring the grasp poses, these methods propose learning various intermediary information such as grasp quality measures or reconstructed 3D object shapes and then use this information to help infer grasp poses. On the positive side, it has been shown that learning this kind of information can improve both the data-efficacy of training and the success rate of predicted grasp poses. On the negative side, however, this intermediary information complicates the training procedure, hyper-parameter search, and data preparation .

Ideally, a learning-based grasp planner should infer the grasp poses directly from raw sensor inputs such as RGB-D images. Such approaches have been developed by many researchers . However, recent methods show that it is preferable to first learn a grasp quality metric function and then optimize the metric at runtime for an unknown target object using sampling-based optimization algorithms, such as multi-armed bandits . Such optimization can be very efficient for low-DOF parallel jaw grippers but less efficient for high-DOF anthropomorphic grippers due to their high-dimensional configuration spaces. In addition, it is possible for the sampling algorithm to generate samples at any point in the configuration space, and the learned metric function has to return accurate values for all these samples. To achieve high accuracy, a large amount of training data is needed, as shown in the 6.7 million ground truth grasps in the dataset used by .

Various techniques have been proposed to improve the robustness and efficiency of grasp planner training. Prior works proposed improving the data-efficiency of training by having the neural network recover the 3D volumetric representation of the target object from 2D observations. A 2D-to-3D reconstruction sub-task allows the model to learn intrinsic features about the object. However, a volumetric representation also incurs higher computational and memory cost. In addition, compared with surface meshes, volumetric representations based on signed distance fields cannot resolve delicate, thin features of complex objects . Demonstrating an alternative method, prior works in show that higher robustness can also be achieved using adversarial training, which in turn introduces additional sub-tasks of training and requires new data.

Main Results: We present a differentiable theory of grasp planning, extending ideas from , an early attempt to formulate grasp planning as a continuous optimization. Our main contribution is a generalized definition of the grasp quality metric that is defined when the gripper is not in contact with the target object. We show that this metric function is locally differentiable and that its gradient can be computed from the sensitivity analysis of the optimality condition in a similar manner to . We also propose a loss function to ensure that grasps are (self-)collision-free in a differentiable manner, which can be computed from only surface meshes of target objects.

Our method can be used as a locally optimal grasp planner similar to simulated annealing , but our method is guided by analytic gradients and can quickly find a locally optimal solution. More importantly, our method can be used to improve the quality of learned grasp poses using a simple neural network architecture. Specifically, we use a network that takes as input a set of multi-view depth images of the target object and directly predicts a grasp pose for a high-DOF gripper. This design choice is preferable to those in prior work because it leads to a higher performance during runtime, as there is no need to optimize a learned grasp metric and we can obtain the grasp pose by a single forward propagation through the neural network.

By adding our differentiable loss, we show that the simple neural network architecture can predict high-quality grasps for the Shadow Hand (Figure 1) after training on a dataset of only 400 objects and 40K ground truth grasps. When compared with the supervised learning baseline , our method achieves a 22%22\% higher success rate on physical hardware and a 0.120.12 higher value in the Q1Q_{1} grasp quality metric . Our learning architecture is illustrated in Figure 2.

II Related Work

In this section, we review related works in grasp planning that uses either model-based or learning-based methods.

Model-Based Grasp Planners assume perfect sensing of environment geometries and target object shapes. Given the geometric information, a grasp planner searches for a grasp pose that maximizes a certain grasp quality metric; many techniques have been proposed for defining reasonable grasp quality metrics and designing efficient search algorithms . These methods can be applied to both low- and high-DOF grippers and can be classified into discrete sampling-based techniques and continuous optimization techniques . Sampling-based methods allow virtually any grasp quality metric to be used as the objective function, while continuous methods require the metric to be differentiable with respect to the configuration of the gripper. In practice, continuous optimization techniques are more efficient in terms of finding the (locally) optimal grasp poses. Some works plan grasps by optimizing differentiable losses. However, these losses do not directly measure grasp qualities.

Some planning methods only compute optimal grasp points, while others compute both the grasp points and the gripper poses. When a gripper pose is needed, the planner uses a two-stage approach: a set of grasp points is first selected on the surface of the target object and then the pose of the gripper is found by inverse kinematics. Based on the idea of numerically optimizing the grasp quality metric, we extend the definition of a grasp quality metric to be well-defined in the ambient space, i.e. when the gripper is not in contact with the target object, thereby unifying grasp points selection and gripper pose computation.

Learning-Based Grasp Planners can predict grasp points or gripper poses given noisy observations of the environment. Most early works in this direction assume that a parallel-jaw gripper is designed for the target object and that the input is a single depth image of the target object. In this case, the grasp problem boils down to selecting the gripper’s initial direction and orientation, which can be solved using an analytic method . A noteworthy success in this problem is achieved by DexNet , which uses deep convolutional neural networks to learn object similarity functions and grasp quality functions. DexNet can robustly pick a large number of unknown objects using a dataset of tens of thousands of target objects and millions of ground truth grasp poses.

More recent techniques aim to improve the data efficiency of learning-based planners and also make the planner robust in challenging settings involving high-dimensional visual observation of the environment , arbitrary approaching directions , more general gripper types , and model discrepancies . It has been shown in , among others, that the grasp planning task can be divided into two sub-tasks, object reconstruction and gripper pose prediction, and that learning these two sub-tasks can improve the rate of success. It is shown in that adversarial training can also improve the robustness of the learned model. However, these methods either perform extensive data generation or require delicate parameter tuning for the adversarial training.

A common drawback of prior works is that they learn a grasp quality metric function or grasping success predictor, which requires an additional sampling-based optimizer to search for gripper poses. This requirement limits these methods to low-DOF grippers, since the high-dimensional configuration space of high-DOF grippers makes the sampling-based optimization computationally costly. Some recent methods overcame this difficulty by directly predicting a nominal gripper pose from an observation of the object. However, the predicted gripper pose is not directly usable and needs to be post-processed. In comparison, our method predicts robust, usable gripper poses using a simple neural-network architecture and uses a smaller dataset for training. In addition, our method can be combined with previous learning-based methods to improve their results.

As an alternative to supervised learning, reinforcement learning allows a learned grasp planner to discover useful grasp poses through exploration. Learned grasp planners have been successfully applied to grasping and other manipulation problems . However, the number of state transition data needed in a typical training is on the level of millions , while we show that robust gripper poses can be predicted by supervised learning on a dataset with 400 example objects using 40K ground truth grasp poses.

III Learning Grasp Poses for High-DOF Grippers

However, when the gripper is high-DOF, the maximization of Nˉ\bar{\mathcal{N}} becomes a search in a high-DOF configuration space, which is time-consuming. As a result, we choose to learn N\mathcal{N} instead of Nˉ\bar{\mathcal{N}}. The major challenge in learning N\mathcal{N} is to resolve the ambiguity in grasp poses, because infinitely many grasp poses can have the same grasp quality for a target object but our neural network N\mathcal{N} can only predict one pose. In order to resolve this ambiguity in grasp poses, we need the dataset to be consistent. A consistent grasp pose dataset is one where all the ground truth grasp poses can be represented by a single neural network. To enforce consistency, one prior work attempted to train N\mathcal{N} by precomputing multiple grasp poses for each target object and used a Chamfer loss to have N\mathcal{N} pick the most consistent pose. However, the learned gripper poses cannot be used directly due to their low quality, and post-processing is needed to deploy the learned poses on physical hardware.

We aim to further improve the quality of the learned function N\mathcal{N} without increasing the complexity of training in terms of either the amount of data or the network architecture. Instead, we are inspired by earlier works , which formulate grasp planning as a continuous optimization. We incorporate all the criteria of good grasps as additional loss functions in terms of stochastic optimization. It has recently been shown that gradients can be brought through complex numerical algorithms to provide additional guidance. These domain-specific differentiable models can significantly improve the convergence rate of neural-network training and reduce the amount of data needed. However, we need to overcome several difficulties when using these approaches for grasp planning:

All the existing grasp quality metrics have discontinuities , so we have to modify them for differentiability.

A grasp quality metric is only defined when the gripper and the target object have exact contact, which is generally not the case when gripper poses are being stochastically updated by the training algorithm.

Our differentiable loss function is defined for a target object represented using triangle meshes. These triangle meshes come from well-known 3D shape datasets , some of which are of low quality. If gradient computation becomes unreliable on low-quality meshes (with nearly degenerate triangles), training will be misled.

We present our design of loss functions and discuss how to address the three challenging problems in the next section.

IV Differentiable Grasp Planner

Our loss function L\mathcal{L} is comprised of three terms: e−Q1\mathbf{e}^{-Q_{1}}, Lguide\mathcal{L}_{guide}, and Lcoll,self\mathcal{L}_{coll,self}. The first term is a generalized Q1Q_{1} grasp metric that measures the quality of a grasp using physics-based rules. However, when force closure is not satisfied, both the metric value and its gradients are zero. In this degenerate case, we add a second, heuristic term Lguide\mathcal{L}_{guide} that always provides a non-vanishing gradient. Our third term Lcoll,self\mathcal{L}_{coll,self} penalizes both self-collision and collisions between the gripper and the target object.

Throughout the paper, we assume that a target object is defined by a watertight triangle mesh T\mathcal{T}. As illustrated in Figure 3, given a point p\mathbf{p} in the workspace, we can also define the signed distance to T\mathcal{T} as d(p)\mathbf{d}(\mathbf{p}) and the outward normal with respect to T\mathcal{T} as n(p)\mathbf{n}(\mathbf{p}). In addition, we also define ng(p)\mathbf{n}_{g}(\mathbf{p}) as the gripper normal, i.e. the outward normal direction on the gripper mesh. We further assume that the target object’s center-of-mass coincides with the origin of the Cartesian coordinates. During grasping, the object will be under an external wrench \mathbf{w}=\left(\begin{array}[]{cc}{\mathbf{f}}^{T},&{\boldsymbol{\tau}}^{T}\end{array}\right)^{T}.

For a set of grasp points p1,⋯ ,N\mathbf{p}_{1,\cdots,N} satisfying d(pi)=0\mathbf{d}(\mathbf{p}_{i})=0, with respective grasping forces fi\mathbf{f}_{i}, the quality of a grasp pose is defined by the Q1Q_{1} metric as follows:

In practice, it is infeasible to assume that a grasp metric can be computed in its original form, i.e. Equation 1. This is because a learning system will generally not produce grasping points that lie exactly on the surface of the target object. It is well known that incorporating hard constraints into neural networks is difficult . When a stochastic training scheme is used and neural network parameters are randomly perturbed, exact constraint satisfaction will be lost. As a result, we have to deal with cases where d(pi)≠0\mathbf{d}(\mathbf{p}_{i})\neq 0. Taking these cases into account, we derive a generalized version of Q1Q_{1} by modifying the first condition of admissible wrenches in Equation 1 as follows:

which essentially extends Q1Q_{1} to the ambient space by an exponential weight function with two terms. The first term ∥d(pi)∥\|\mathbf{d}(\mathbf{p}_{i})\| ensures that our generalized Q1Q_{1} attains larger values when grasp points are closer to the surface of the target object. The second term β(1+n(pi)Tng(pi))\beta(1+\mathbf{n}(\mathbf{p}_{i})^{T}\mathbf{n}_{g}(\mathbf{p}_{i})) ensures that our generalized Q1Q_{1} attains larger values when the normal direction on the gripper and the normal direction on the target object align. Finally, it is obvious that Equation 7 converges to Equation 1 as α,β→∞\alpha,\beta\to\infty. Like previous works on generalized contact-implicit models, our generalized metric allows a learning algorithm to determine the number of contact points and their positions.

To train neural networks using the generalized Q1Q_{1} metric, we need to compute its sub-gradient with respect to pi\mathbf{p}_{i} efficiently. Unfortunately, the exact computation of the Q1Q_{1} metric is difficult because the optimization in Equation 1 is non-convex; several approximations have been proposed in . We present two different techniques for computing Q1Q_{1} and ∂Q1/∂pi\partial{Q_{1}}/\partial{\mathbf{p}_{i}}. The first method computes an upper bound of generalized Q1Q_{1}, which is cheaper to compute but creates zero entries in the gradient vector. The second method computes a smooth, lower bound of generalized Q1Q_{1}, which propagates non-zero gradient information but costlier to compute.

Our first technique adopts , which approximates Q1Q_{1} by assuming that w\mathbf{w} must be along one of a discrete set of directions: s1,⋯ ,D\mathbf{s}_{1,\cdots,D}. This assumption results in a tractable upper bound of Q1Q_{1} and can be extend to our generalized Q1Q_{1} metric as follows:

In this form, each operation for computing our generalized Q1Q_{1} can be implemented as a standard math operation with derivatives that can be computed using automatic differentiation tools such as .

We have shown that computing an upper bound of Q1Q_{1} reduces to a series of simple operations with well-defined sub-gradients. However, due to the max  \underset{}{\mathbf{max}}\; function in the computation of the Q1Q_{1} upper bound, the sub-gradient is non-zero for only one of the contact points, which is less efficient for training. To resolve this problem, it has been shown in that sum-of-squares (SOS) optimization can be used to compute a lower bound of Q1Q_{1}. This theory can be extended to compute our generalized Q1Q_{1} metric. If we define t1,⋯ ,T(pi)\mathbf{t}_{1,\cdots,T}(\mathbf{p}_{i}) as a set of directions on the tangent plane, then the generalized Q1Q_{1} can be found by solving the following SOS optimization problem:

where we have extended the definition of Vjk\mathcal{V}_{jk} to account for our generalization (Equation 7). Equation 9 can be reduced to a semidefinite programming (SDP) problem, and its gradients can be computed via the chain rule:

While the second term in the chain rule above can be computed directly via automatic differentiation, the first term ∂Q1/∂Vik{\partial{Q_{1}}}/{\partial{\mathcal{V}_{ik}}} requires a sensitivity analysis of an SDP problem, as shown in (see supplementary material for more details). Since SDP is a smooth approximation of a non-smooth optimization, the derivatives are generally non-zero on all the contact points. As a result, each neural network update can adjust all the fingers of the gripper to generate better grasp poses, which is more efficient than the case with an upper bound on Q1Q_{1}. On the other hand, the cost of solving Equation 9 is also higher than that of solving Equation 8 because Equation 9 involves an SDP solve. Note that a similar analysis for quadratic programming (QP) problems has been previously exploited for training neural networks in .

IV-C Geometry Related Loss Functions

In this section, we show that geometric terms such as d(p)\mathbf{d}(\mathbf{p}) can be computed robustly from a triangle mesh. We also formulate the collision-free requirement as a novel loss term. Geometric terms arise in many places in a grasping system. To compute the Q1Q_{1} metric, we need to evaluate d(pi)\mathbf{d}(\mathbf{p}_{i}) and n(pi)\mathbf{n}(\mathbf{p}_{i}). In addition, we need to avoid penetrations between grippers and the target objects. To perform these computations, we can introduce a monotonic loss function:

where β\beta is the weight of loss. To provide sub-gradients for all these terms, we need to plug ∂d(pi)∂pi\frac{\partial{\mathbf{d}(\mathbf{p}_{i})}}{\partial{\mathbf{p}_{i}}} into the chain rule. In this section, we show a robust method to compute ∂d(pi)∂pi\frac{\partial{\mathbf{d}(\mathbf{p}_{i})}}{\partial{\mathbf{p}_{i}}} for complex, watertight, triangle meshes of the target objects, which can be accelerated with the help of a bounding volume hierarchy (BVH). Note that it is easy to compute d(pi)\mathbf{d}(\mathbf{p}_{i}) and its gradients from a signed distance field (SDF) , but we choose to use triangle meshes for two reasons. First, most existing 3D shape datasets such as use triangle meshes, and converting them to SDFs is time- and memory- consuming. Second, for very complex meshes, low-resolution SDFs cannot represent thin geometric features and determining an appropriate resolution of SDF is difficult.

Let’s assume that a triangle mesh T\mathcal{T} consists of a set of triangles Tj\mathcal{T}_{j}. Then the distance between pi\mathbf{p}_{i} and Tj\mathcal{T}_{j} is the solution of the following QP problem:

where vk,j\mathbf{v}_{k,j} is the kkth vertex of Tj\mathcal{T}_{j}. Finally, the signed distance d(pi)\mathbf{d}(\mathbf{p}_{i}) is defined as:

where nj\mathbf{n}_{j} is the outward normal of Tj\mathcal{T}_{j} and sgn\mathbf{sgn} is the sign function. Similarly, we can define the outward normal of n(pi)\mathbf{n}(\mathbf{p}_{i}) to be:

In these formulations, the sign function and the argminj  \underset{j}{\mathbf{argmin}}\; operator define a disjoint convex set with well-defined sub-gradients. The gradient of d(pi,T)\mathbf{d}(\mathbf{p}_{i},\mathcal{T}) is:

Also, the gradient of n(pi)\mathbf{n}(\mathbf{p}_{i}) can be computed from the dirichlet features on the triangle mesh to which pi\mathbf{p}_{i} belongs, as illustrated in Figure 4. If the closest feature to pi\mathbf{p}_{i} is an edge e\mathbf{e}, then we have:

If the closest feature to pi\mathbf{p}_{i} is a vertex v\mathbf{v}, then we have:

If the closest feature to pi\mathbf{p}_{i} is inside a triangle, then ∂n(pi)/∂pi=0{\partial{\mathbf{n}(\mathbf{p}_{i})}}/{\partial{\mathbf{p}_{i}}}=0.

Finally, Equation 10 and Equation 11 involve a loop over all triangles to find the one with smallest distance, which can be accelerated by building a BVH and quickly rejecting nodes where the bounding volume is further from pi\mathbf{p}_{i} than the current best distance .

In our experiments, the technique described above is computationally efficient but prone to floating-point’s truncation error. If a point is close to the triangle’s plane, finite-precision floating point arithmetics have difficulty deciding whether the point lies inside the triangle mesh or not. To solve this problem, we use exact rational arithmetics implemented in to perform all the computations in this section and convert the results back to inexact, finite precision floating point numbers at the end of the computation.

IV-D Self-Collision of the Gripper

IV-E Defending Against Degenerate Cases and Local Minima

Our generalized Q1Q_{1} metric is similar to the standard Q1Q_{1} metric in that it implies force closure. However, if an initial guess for the gripper pose has no force closure, then Q1=0Q_{1}=0 and no gradient information is available. In this case, we add the following heuristic term to guide the optimization to compute a force-closed pose with a high probability:

by ensuring that all the grasp points are as close to the object as possible. In addition, our generalized Q1Q_{1} has many local minima due to nonlinearity and complex geometries of objects. To defend our neural network against these sub-optimal solutions, we add a data loss to guide the training. We use Chamfer loss for our data term:

following previous works , where N∗\mathcal{N}^{*} is the ground truth grasp pose and Ldata\mathcal{L}_{data} is the Chamfer distance measure in the gripper’s configuration space. In other words, we precompute many ground truth grasp poses for each target object and let the neural network pick the grasp pose that leads to the minimal distance.

IV-F Forward Kinematics

Our neural network predicts N\mathcal{N}, which consists of the global rigid transformation and the joint angles to define the pose of a gripper. Further, the gradient with respect to the grasp points pi\mathbf{p}_{i} is propagated backward to N\mathcal{N} via a forward kinematics layer denoted as FK\mathbf{FK}, similar to . We make a minor modification to account for joint limits with non-vanishing gradients. If N\mathcal{N} has joint limits in range [l,u][\mathbf{l},\mathbf{u}], then we transform N\mathcal{N} as follows:

which is guaranteed to satisfy the constraints and has non-vanishing gradients compared with the min,max\mathbf{min},\mathbf{max} functions.

In summary, our learning system uses the following compound loss function:

V Experimental Setup

Data Preparation: Following , we prepare a small dataset of 500 watertight objects by combining existing grasping datasets . We split the dataset into an 80%80\% (400) training set and a 20%20\% (100) test set. It is known that predicting a single grasp pose from a single target object is an ambiguous problem because many grasp poses are equally effective . Therefore, we use to precompute a set of 100100 grasp poses for each target object and then use Chamfer data loss to let the neural network pick which grasp pose is the most representable (details can be found in ). This gives a dataset of 4040K grasps, from which our neural network will select 400400 as ground truth. For our 24-DOF gripper, collecting these data requires about 150 CPU hours of computation on a cluster using a sampling-based grasp planner . Finally, we assume that the neural network observes objects from a set of 55 multi-view depth cameras of resolution 224×224224\times 224. These images are obtained by using Blender 2.79 to render the triangle mesh of each target object into the depth channel . As a result, each sample in our dataset is a <D1,⋯ ,5,T,N∗><D_{1,\cdots,5},\mathcal{T},\mathcal{N}^{*}>-tuple of depth images, triangle mesh, and ground truth grasp poses. After collecting our dataset, we augment it by rotating each target object and gripper for 8 times along 8 symmetric axes.

Gripper Setup: In all our simulated and real-world experiments, we use a (6+18)-DOF Shadow Hand as our gripper, as shown in Figure 5, which is mounted onto a UR10 arm. However, during the training phase, the DOFs of the arm are not predicted by our neural network. These DOFs are computed at runtime using a conventional motion planner. We use the SrArmCommander to move the UR10 arm to the target poses and use the SrHandCommander to move the Shadow Hand fingers to the target joint states. During the training phase, we manually label NN=45 potential grasp points on the gripper and, to detect self-collisions, we use a denser sample of 15,555 potential contact points using Poisson disk sampling, as illustrated in Figure 5.

Neural Network: We deploy a pre-trained ResNet-50 from the TORCHVISION.MODELS offered by PyTorch as a feature extractor for multi-view depth images. We then fine-tune it with depth images. For each depth image, we duplicate it to 3 channels to meet the input requirement of ResNet-50. A shared ResNet-50 takes multi-view depth images as input and outputs 2,048 dimensional vectors. These vectors are used with max-pooling and connected with a fully-connected layer, of which the output dimension is equal to the gripper’s DOF (6+18 for Shadow Hand). Outputs of the fully-connected layer are the predicted gripper configurations.

Training configurations: We use the parameters listed in Table I for L\mathcal{L} in both settings. Our neural network is trained using the ADAM algorithm with a batch size of 16. The initial learning rate is set to be 1ee-4 and decayed by 0.9 every 20 epochs. All experiments are carried out on a desktop with 2 Intel® Xeon Silver 4208 CPUs, 32 GB RAM, and 2 NVIDIA® RTX 2080 GPUs.

VI Experimental Results

In this section, we evaluate the performance of different settings for high-DOF grasp planning. Our method can be used either as a standalone grasp planner or as a method to train grasp predicting neural networks.

The differentiable grasp metric and collision loss allows our method to be used as a standalone, locally optimal grasp planner. To setup this experiment, we replace the neural network with a (6+18)(6+18)-DOF optimizable vector of the gripper pose, set cdata=0c_{data}=0, and minimize L\mathcal{L} with respect to N\mathcal{N}. Compared to , our planner only provides local optima, and the computational cost is comparable. An example is illustrated in Figure 6, where we use a trivial initialization shown as the transparent green poses. After 22 minutes of optimization, our optimizer converges to the gray poses. However, without guidance from data, our planner can fall into local minima without force closure (Q1=0Q_{1}=0), as shown in the inset.

VI-B Learning Grasp Poses with Ground Truth

In our second benchmark, we use our method to guide the training of a grasp-pose-predicting neural network. The training is performed in two phases. First, we adopt a pre-training by setting L=Ldata\mathcal{L}=\mathcal{L}_{data}, i.e. excluding our differentiable loss. This step brings the neural network close to nearly optimal values and we run 35 epochs of learning at the first stage. Second, we fine-tune the network by adding our differentiable loss and use weights as summarized in Table I. We run 71 epochs of learning at the second stage. The pre-training takes 4 hours and the fine-tuning takes 36 hours. On average, each forward-backward propagation with our additional loss function takes 0.85s and the one without our loss function takes 0.61s, which shows that our additional loss functions only impose a marginal cost to gradient computation. However, to ensure fine-grained convergence to a good local minimum, we use a small learning rate and more epochs for the second stage, which runs 9×9\times slower than the first stage. After training, we test our neural network on the set of 100 held-out objects, on which the mean Q1Q_{1} metric is 0.226 and the variance of the Q1Q_{1} metric is 0.0045. A set of predicted grasp poses on the test set is shown in Figure 9, from which we observe drastically improved grasp quality when guided by the data term. On a physical platform, however, there might be environmental constraints making our predicted grasps infeasible. In this case, we can randomly perturb object poses to create virtual depth images and predict a set of varied grasps from them, as shown in Figure 8, from which we can pick one feasible grasp.

VI-C Comparison

We have compared the (standard) Q1Q_{1} metric of our method and a sampling-based grasp planner in Figure 7. The results show that the qualities of our grasp poses are on par with those of . We have also compared our approach with prior work , which also trains a grasp-pose predicting neural network on a small dataset of a size similar to ours. However, the algorithm in requires a post-processing step to resolve penetrations and collisions. Instead, the grasp poses predicted using our method can be directly deployed onto a physical hardware without post-processing. As illustrated in Figure 9 and Table II, our method can significantly improve the quality of grasp poses.

VI-D Grasping with the Arm on Physical Hardware

As our final evaluation, we deploy our learned neural network onto our physical platform. Our method does not require RGB input and only uses the depth channel. Therefore, we do not perform any sim-to-real transfer. Our neural network only predicts the gripper pose and does not predict the configuration for the UR10 arm to achieve the predicted position and orientation. These configurations of the arm are computed using a motion planner at runtime. We choose 50 YCB objects from our 100 test objects. All YCB objects are unseen and excluded from the training set. We use the depth images from five ASUS Xtion PRO LIVE cameras with 640x480 resolution as our network input. Our depth cameras are calibrated beforehand to make the camera pose exactly the same as in training. We then crop the real depth images and remove the background. Moreover, we make objects’ poses exactly the same as the poses used for rendering depth images.

To profile the rate of success on the 50 YCB objects, we use two metrics summarized in Table II. First, we record how many times the motion planner can successfully move the gripper to the predicted position (Success-Plan). This metric measures the ability of our method to avoid penetrations and collisions with the desk on which the objects are placed since a pose with penetrations or desk collisions cannot be achieved by a motion planner. Second, we record how many times the grasp planner can successfully lift the object (Success-Grasp). We define our Success-Rate as the percentage of Success-Grasp out of Success-Plan, and we define Overall-Success-Rate as the percentage of Success-Grasp out of the 50 trials. This metric measures the ability of our method to improve the grasp quality. Our method outperforms in terms of both metrics. We observe an 8%8\% improvement in terms of Success Plan and a 22%22\% improvement in terms of Success Grasp. We claim that we have succeeded when the object has no relative motion against the hand for a sufficiently long period. Our neural network failed on 55 objects due to slippage. These 5 objects are: wood_block, power_drill, extra_large_clamp, hammer, and potted_meat_can.

VII Conclusion and limitations

We present a differentiable grasp planner that enables a neural network to be trained with a small dataset and a simplified architecture. Our differentiable loss accounts for various requirements for a good grasp, including high grasp metric values and collision-free gripper poses. We use a generalized definition to allow inexact contact and we show that the sub-gradients of each loss term are well-defined and can be efficiently computed from target object shapes represented using watertight triangle meshes. We show that our method can be used both as a standalone grasp planner and as a neural network training algorithm. Finally, we show that the trained neural network performs robustly on unseen objects and hardware platforms.

Our current implementation suffers from several limitations. First, our method requires the target objects to be watertight and to have a non-zero volume. Although we do not require a signed distance field transformation, our method still computes a signed value of distance, which is impossible when the target object is a thin-shell. A limitation related to this problem is that our method suffers from tunneling. In other words, when the target object is very thin, a stochastic update of our neural network might result in the hand going from one side to the other side of the object, leading to missed solutions. In the future, this problem can be resolved using continuous collision detection . Second, our experimental setup and neural network architecture prevents the neural network from predicting multiple grasp poses for a single object. If there are other constraints in the workspace preventing a grasp pose from being achieved, then our method will lead to failure. However, this problem can be resolved by using adversarial training similar to , where a distribution of grasp poses is learned. We emphasize that more sophisticated learning algorithms are orthogonal to our approach and can be combined with it. Finally, by using the exact Q1Q_{1} metric as our loss function, we can only generate precision grasps, meaning more robust power grasp or caging grasp generation is left as future work.

VIII Acknowledgments

We thank the anonymous reviewers for their valuable comments. This work was supported in part by the National Key Research and Development Program of China (No. 2018AAA0102200), NSFC (61572507, 61532003, 61622212), ARO grant W911NF-18-1-0313, and Intel. Min Liu is supported by the China Scholarship Council.

IX Appendix

In this document, we provide some details on computing the Q1Q_{1} lower bound and its derivatives. First, we derive a slightly different formulation of the Q1Q_{1} lower bound using quadratic frictional cones. As compared with the linearized frictional cones used in , using quadratic frictional cones is more efficient in terms of reducing the problem size of semidefinite programming.

1 Q1 Lower Bound Using Quadratic Frictional Cone

We re-derive the lower bound of Q1Q_{1} using SOS optimization as done in , but using quadratic frictional cones. For a set of points p1,⋯ ,N\mathbf{p}_{1,\cdots,N}, with normals n(pi)\mathbf{n}(\mathbf{p}_{i}) and two tangents being t1,2\mathbf{t}_{1,2}, then the cones KB,W\mathcal{K}_{\mathcal{B},\mathcal{W}} are defined as:

It is easy to find that the dual cones of KB,W\mathcal{K}_{\mathcal{B},\mathcal{W}} are defined as:

Note that Equation 12 will induce an SDP problem with exactly the same order (of polynomials) as the original SDP problem induced in , but with fewer cones and also smaller linear system when performing sensitivity analysis. Finally, we briefly prove the correctness of KW∗\mathcal{K}_{\mathcal{W}}^{*}.

The dual cone of KW\mathcal{K}_{\mathcal{W}} is KW∗\mathcal{K}_{\mathcal{W}}^{*}.

Proof: If \left(\begin{array}[]{cc}{\mathbf{a}}^{T},&{b}^{T}\end{array}\right)^{T}\in\mathcal{K}_{\mathcal{W}}^{*}, then for any \left(\begin{array}[]{cc}{\mathcal{V}}^{T},&{\sum_{i}\lambda_{i}}^{T}\end{array}\right)^{T}\in\mathcal{K}_{\mathcal{W}}, we have:

For the other direction, if there is an ii such that (Vi⊥Ta+b)/μ<(Vi∥1Ta)2+(Vi∥2Ta)2({\mathcal{V}_{i}^{\perp}}^{T}\mathbf{a}+b)/\mu<\sqrt{({\mathcal{V}_{i}^{\parallel 1}}^{T}\mathbf{a})^{2}+({\mathcal{V}_{i}^{\parallel 2}}^{T}\mathbf{a})^{2}}, then we can pick a point in KW\mathcal{K}_{\mathcal{W}} as follows:

2 SDP Sensitivity Analysis for Lower Bound of Q1Q_{1}

In this section, we present an efficient way to perform sensitivity analysis for SOS problems. We use the same notations as those in . A standard SDP takes the form:

where there are j=1,⋯ ,1+KNj=1,\cdots,1+KN PSD cones in our problem (KK equals the number of tangent directions if linearized frictional cones are used and K=2K=2 if quadratic frictional cones are used). The dual variable to the jjth cone is Zj\mathbf{Z}^{j} and we have FjZj=0\mathbf{F}^{j}\mathbf{Z}^{j}=0. We also define the coefficient matrix:

In an SOS problem, the first PSD cone and the KNKN other cones are of two different types. The first PSD cone F1\mathbf{F}^{1} specifies the conditional polynomial positivity condition. The other KNKN cones FKN\mathbf{F}^{KN} specify the positivity of Lagrangian multipliers. We also observe that some variables x1\mathbf{x}^{1} only affect the first PSD cone and other variables xKN\mathbf{x}^{KN} affect the other KNKN PSD cones. Therefore, we can write the matrix F\mathcal{F} in a 2×22\times 2 block form as follows:

where it is trivial to verify that we can choose variables to make the bottom right block of F\mathcal{F} an identity matrix. When SDP is solved using primal-dual interior point method, the set of primal and dual solutions are computed simultaneously, with the dual variables defined as:

where we apply the same decomposition of cones for Z\mathbf{Z}. Next, we apply the optimalty condition of SDP:

where ⊗\otimes is the symmetric kronecker product operator. Apply sensitivity analysis with respect to an arbitrary parameter ϵ\epsilon, we have:

In the following derivation, we assume that FKN\mathbf{F}^{KN} and ZKN\mathbf{Z}^{KN} have strict complementarity. Note that if strict complementarity is not satisfied, then the SDP problem is not differentiable. Prior work showed that F1,KN\mathbf{F}^{1,KN} and Z1,KN\mathbf{Z}^{1,KN} have simultaneous diagonalization, and so does F1,KN⊗I\mathbf{F}^{1,KN}\otimes\mathbf{I} and Z1,KN⊗I\mathbf{Z}^{1,KN}\otimes\mathbf{I}:

By plugging these identities info the sensitivity equation, we get:

where there are 8 equations. For 4,5,7,84,5,7,8th rows, we have:

By plugging Equation 13 into the 11st row, we have:

By plugging Equation 13 into the 22nd row, we have:

The 3,63,6th rows will remain intact and 66th row can be eliminated. Finally, our reduced sensitivity equation is:

The main benefit of the system reduction is that the left-hand-side of this matrix becomes symmetric (after removing ΣZ11\Sigma_{\mathbf{Z}1}^{1} and ΣZ1KN\Sigma_{\mathbf{Z}1}^{KN} from the last two rows, respectively). We solve this reduced sensitivity equation using the rank-revealing PLD(PL)TPLD(PL)^{T} factorization.

References