Graspness Discovery in Clutters for Fast and Accurate Grasp Detection
Chenxi Wang, Hao-Shu Fang, Minghao Gou, Hongjie Fang, Jin Gao, Cewu Lu
Introduction
As a fundamental problem in robotics, robust grasp pose detection for unstructured environment has been fascinating our community for decades. It has a broad spectrum of applications in picking , assembling , home serving , etc. Advancing the generality, accuracy and efficiency is a long pursuit of researchers in this field.
For grasp pose detection in the wild, it can be regarded as a two-stage problem: given a single-view point cloud, we first find locations with high graspability (where stage) and then decide grasp parameters like in-plane rotation, approaching depth, grasp score and gripper width (how stage) for a local region.
Previous methods for 6-DoF grasp pose detection in cluttered scenes mainly focused on improving the quality of grasp parameter prediction, i.e., the how stage, and two lines of research are explored. The first line adopts a sampling-evaluation method, where grasp candidates are uniformly randomly sampled from the scene and evaluated by their model. The second line proposes end-to-end networks to calculate grasp parameters for the whole scene, where point clouds are sampled before or during the forward propagation. For all these methods, the where stage is not explicitly modeled (i.e., they do not perform a filtering procedure in a first stage) and candidate grasp points distribute uniformly in the scene.
However, we find that such uniform sampling strategy greatly hinders the performance of the whole pipeline. There are tremendous points in 3D contiguous space, while positive samples are concentrated in small local regions. Take GraspNet-1Billion , the current largest dataset in grasp pose detection as an example. We statistically find that, even with object masks, the graspable points are less than 10% among all the samples, not to mention the candidate points in the whole scene. Such an imbalance causes a large waste of computing resources and degrades the efficiency.
To tackle the above bottleneck in grasp pose detection, we propose a novel geometrically based quality, graspness, for distinguishing graspable area in cluttered scenes. One might think that we need complex geometric reasoning to obtain such graspness. However, we discover that a simple look-ahead search by exhaustively evaluating possible future grasp poses from a point can well represent its graspness. Statistical results demonstrate the justifiability of our proposed graspness measure, where the local geometry around points with high graspness are distinguished from those with low scores. Fig. 1 gives an illustration of our graspness for a cluttered scene.
Furthermore, we develop a graspness model that approximates the above process in practice. Given a point cloud input, it predicts point-wise graspness score, which is referred to as graspable landscape. Benefiting from the stability of the local geometry structures, our graspness model is object agnostic and robust to variation of viewpoint, scene, sensor, etc., making it a general and transferable module for grasp point sampling. We qualitatively evaluate its robustness and transferability in our analysis. Tremendous improvements in both speed and accuracy for previous sampling-evaluation based methods are witnessed after equipping them with our graspness model.
Based on our graspness model, we also propose Graspness-based Sampling Network (GSNet), an end-to-end two-stage network with a graspness-based sampling strategy. Our network takes a dense scene point cloud as input, which preserves the local geometry cues. The sampling layer firstly selects the points with high graspness. Remaining points are discarded from the forward propagation to improve the computation efficiency. Such two-stage design is beneficial to network convergence and also the final accuracy by providing more positive samples during training.
We conduct extensive experiments to evaluate the effectiveness of our proposed graspness measure, model and the end-to-end network. Several baseline methods equipped with our graspness model outperform their vanilla counterparts by a large margin in both speed and accuracy. Moreover, our GSNet outperforms previous methods to a large extent. Our library has been integrated into AnyGrasp to facilitate research in the robotics community.
Related Work
In this section, we first briefly review previous methods on grasping in cluttered scenes, followed by concluding the common strategies they have used to sample grasp candidates. Finally we surveyed some literature in cognitive science area where graspness recognition is witnessed in human perception.
For cluttered scene grasp pose detection, previous research can be mainly divided into two categories: plannar based grasp detection and 6-DoF based grasp detection. The research in the first category mainly took RGB images or depth images as inputs and output a set of rotated bounding boxes to represent the grasp poses. Due to the limitation of low DoF, their applications were usually restricted. Another line of research aimed to predict full DoF grasp poses. Among them, two different directions were explored. The first direction adopted the sampling-evaluation based two-step policy, where grasp candidates were densely uniformly sampled in the scene and evaluated using a deep quality model. The second direction adopted the end-to-end strategy, where point clouds of the scene were directly processed by end-to-end networks. For each input point, the network attempted to predict the most feasible grasp pose. All the mentioned methods focused on improving the quality of grasp parameters, and the problem of where to grasp was not investigated.
Several kinds of sampling strategies can be concluded from the above methods. The most common used strategy is the uniform sampling, which is adopted by . Specifically, GPD and PointNetGPD uniformly sampled grasp points in the scene point cloud and estimated the rotation by darboux frame. Some end-to-end models down-sampled the input point cloud by voxel grid to avoid memory explosion. A similar strategy, farthest point sampling, is adopted by other end-to-end model . Some optimization based methods are also explored. Ciocarlie et al. and Hang et al. adopted the simulated annealing method, while Mahler et al. proposed cross-entropy methods. In , a grasp sampler network first sampled possible grasp poses on partial object point cloud and conducted iterative refinement by a grasp evaluator based on its gradient. In a recent paper by Clemens et al. , several sampling methods for grasp dataset generation are reviewed. However, all the previous methods ignore the geometric cues for graspable point sampling. In this paper, we propose a novel graspness measure based on local geometry for graspable point sampling, which is much more efficient than previous uniform sampling and optimization based methods.
In cognitive area, researchers have studied the visual attention during grasping for a long period. Many literature demonstrated that human bias the allocation of available perceptual resources, named as affordance attention, towards the region with the highest graspability. And such attention usually precedes the action preparation stage . Such discovery corresponds to our graspness concept and motivates us to apply it in the grasp sampling strategy.
Graspness Discovery
As mentioned above, we decouple the grasp pose detection problem into two stages. Before the common practice in previous research that directly calculates the grasp parameters, we first sample points and views with high graspness. Computational resources will be allocated to these areas thereafter to improve computational efficiency.
To determine the suitable grasp locations and the feasible approach directions with high graspability, we define two kinds of graspness in a high dimensional space to represent parallel attention in point locations and approach directions. Before detailing our graspness measure, we first introduce some basic notations.
For a point sets , we assume approach directions uniformly distributed in a sphere space .
Two kinds of graspness scores are discussed in this paper. The first is the point-wise graspness scores denoted as
where $$ denotes that our graspness for each point ranges from 0 to 1. The second is the view-wise graspness scores denoted as
where denotes -dim graspness ranging in .
In the following section, we illustrate how we measure graspness for both single object and the cluttered scene.
2 Graspness Measure
By doing so, we guarantee that higher graspness value always denote higher possibility of successful grasping.
In practice, such an oracle does not exist, and can contain infinite grasp poses in a continuous space. Thus, we make an approximation to the above process. For view of point , we generate grasp candidates by grid sampling along gripper depths and in-plane rotation angles. For each grasp , we calculate a grasp quality score using a force analytic model . A threshold is manually set to filter out unsuccessful grasps. Then, the relaxation form of Eqn. 1 is:
After defining the object-level graspness, we extend it to cluttered scenes by first discussing the gap between them and then redefining the graspness in cluttered scenes.
A cluttered scene contains multiple objects and the irrelevant background. As shown in Fig. 2(a), the simplest way to compute scene-level graspness is directly projecting the object-level graspness score to the scene by object 6D poses. However, this solution ignores the differences between an object model and a scene cloud captured from RGB-D camera. Firstly, a valid grasp of a single object may collide with background or other objects when placing in cluttered manner and becomes a negative grasp. Secondly, as the depth camera provides single-view partial point clouds, we need to associate the scene point cloud with the projected object point.
To deal with the collision problem, we follow to reconstruct the scene using object 3D models and corresponding 6D poses. Each grasp is evaluated by a collision checking process and assigned a collision label . Our graspness scores are then updated as:
After that, we project the object points to the scene by object 6D poses. For each point in the scene, we obtain its graspness scores by nearest neighbor search and associate it with the nearest projected object point.
Finally, to obtain a coherent representation for the scene-level graspness scores, we perform a normalization for each scene:
where denotes column wise minimum:
and so does . Fig. 2(b) shows an example of scene-level graspness scores.
3 Justification
In order to justify our graspness measure, we analyze the local geometry for regions with different graspness to find out whether they are really distinguishable geometrically. For a single-view point cloud, the cascaded graspness model detailed in Sec. 4.1 is used to extract the local feature vector of each point. The points with graspness more than 0.3 are treated as positive samples, and negative ones of the same size are sampled with graspness less than 0.1. Fig. 3 shows the t-SNE visualization of the encoded local geometry (feature vectors of each point produced by backbone network) for all the scenes in GraspNet-1Billion training/testing set respectively. We can observe that regions with different graspness are quite distinguishable. It demonstrates that our graspness measure is rational and reveals the potential of learning graspness from point cloud.
GSNet Architecture
After defining the graspness measure, we introduce the end-to-end grasp pose detection network, GSNet, where our graspness is learned by an independent module and can be applied to other methods.
Given a dense single view point cloud , graspness model needs to learn two approximations: and .
It is challenging to find a direct mapping from point coordinates to graspness scores due to the large domain gap between these two spaces. Instead, we decompose the whole process into two sub-functions. Consider a high dimensional feature set :
where denotes function composition, and the feature set is shared by both and .
Although and can be learned simultaneously, the computation overhead is quite expensive since is in high dimensional space. Meanwhile, it is not necessary to compute the view-wise graspable landscapes for all points since most of the points are not graspable at point level. Hence, we propose cascaded graspness model to learn , and step by step, where points are sampled by the output of before learning to reduce computation cost.
Approximation of requires a strong backbone network for extraction of both global and local point features. We adopt ResUNet14 built upon MinkowskiEngine because it can flexibly process point sets of any size with sparse convolution and has shown excellent performance in multiple tasks of 3D deep learning . The network can also be replaced by other point-wise networks, such as PointNet , PointCNN and SSCNs .
The network adopts a U-shape architecture with residual blocks, which obtains point features using 3D sparse (transposed) convolutions and skip-connections. For a point cloud of size , it extracts a -channel feature vector set, and outputs a point set of size for graspable sampling and grasp generation.
The modeling for is implemented with a multi-layer perceptron (MLP) network to generate point-wise graspable landscape. Specifically, the output contains a prediction for the graspable landscape of size and a binary objectness classification scores of size , resulting a total output of size . Graspness scores of non-object points are set to 0.
is also modeled by an MLP. We apply it to the sampled seed points and output vectors for view-wise graspable landscapes and residual features for grasp generation. views are sampled from a unit sphere using Fibonacci lattices .
After obtaining the view-wise graspness scores, we select the best view for afterward predictions during inference. For training, we adopt probabilistic view selection (PVS) that normalizes the graspness scores of all views on a seed point to (0,1) and regard them as probability scores, according to which the view is sampled. The seed point-view pairs are then used to estimate grasp scores, gripper widths, approach distances and in-plane rotation angles.
2 Grasp Operation Model
Crop-and-refine has been proven effective in estimating candidate configuration in both 2D and 3D tasks . We crop points in directional cylinder spaces which are generated by seed point-view pairs, transform them to gripper frames and estimate their grasp parameters.
The locations and directions of cylinder spaces are determined by seed point coordinates and view vectors respectively. For each of the point-view pairs, we group and sample points from seed points using the cylinder with fixed height and radius . After aligning the cylinder with gripper frame as , the point coordinates are normalized by cylinder radius and concatenated with feature vectors which are the sum of features output by graspable FPS and graspable PVS. The grouped point sets of size are called grasp candidates, where stands for the number of sampled points in each group.
We use a shared PointNet for grasp generation. Grasp candidates are processed by an MLP network and a max-pooling layer, and be output as feature vectors of size . Finally we get grasp configurations by a new MLP network.
The output of GSNet contains scores and widths for different (in-plane rotation)-(approach depth) combinations. We pick the combination with the highest score as the grasp prediction. The output size is , where denotes the number of in-plane rotation angles, denotes the number of gripper depths and denotes the score and the width.
We use the minimum friction coefficient under which a grasp is antipodal to evaluate the quality of the grasp. Based on this, we define the grasp score as
All scores are normalized to $\mu_{i}q_{i}$ and more probability to succeed.
2.1 Loss Function
Cascaded graspness model and grasp operation model are trained simultaneously with multi-task losses:
where is for objectness classification, , , and are for regressions of point-wise graspable landscape, view-wise graspable landscape, grasp scores and gripper widths respectively. and are calculated only if the related points are on objects, is calculated for views on seed points and is calculated for grasp poses with ground truth scores 0. We use softmax for classification tasks and smooth- loss for regression tasks.
Experiments
GraspNet-1Billion is a large-scale dataset for grasp pose detection, which contains 190 scenes with 256 different views captured by two cameras (RealSense/Kinect). The testing scenes are divided into three splits according to the object categories (seen/similar/novel). A unified evaluation metric is proposed to benchmark both image based methods and point cloud based methods. We adopt this benchmark as it aligns well with real-world grasping.
The point cloud is downsampled with voxel size 0.005m before being fed into the network, and contains only XYZ in camera coordinates. Input clouds are augmented on the fly by random flipping along YZ plane and random rotation around Z axis in .
To obtain graspness for scenes in GraspNet-1Billion, we follow the process illustrated in Sec. 3.2 since it contains abundant grasp pose annotations. For each point, it densely labels grasp quality score for 300 different views and 48 grasps for each view. Thus, our approach directions and grasp candidates per view are set as 300 and 48.
Our model is implemented with PyTorch and trained on Nvidia GTX 1080Ti GPUs for 10 epochs with Adam optimizer and the batch size of 4. The learning rate is 0.001 at the first epoch, and multiplied by 0.95 every one epoch. The network takes about 1 day to converge. During training, we use one GPU for model update and one GPU for label generation. In inference, we only use one GPU for fast prediction.
2 Performance of Cascaded Graspness Model
Cascaded graspness model is proposed to distinguish graspable areas in various scenes, thus the generality and stability across different domains are important for the model. Here we design an experiment to illustrate its generality and stability.
The ranking error is used to quantitatively evaluate the function approximation ability of the model. We divide the range of graspness score into bins uniformly and convert the contiguous scores to discrete ranks. The ranking error is defined as the mean rank distances between predictions and labels:
We conduct three groups of experiments where the dataset is split by object categories, viewpoints and cameras respectively (detailed in Tab. 1). In the first group, we train the model on scene 0-99, and test it on scenes with three object categories (seen, similar and novel). The second group divides the 256 viewpoints into 3 sets, trains the model on viewpoint 0-127, and tests on three viewpoint sets respectively. The third group trains the model on Kinect captured data, and tests the performance on data captured by RealSense.
3 Comparing with Representative Methods
We compare our method with previous representative methods. GG-CNN and Chu et al. are rectangle based methods which take images as input. GPD and Liang et al. classify grasp candidates generated by rule-based point cloud sampling. Fang et al. propose an end-to-end network which predicts grasp poses directly from scene point clouds.
We test our method in three object categories respectively and report the results in Tab. 2. The models for RealSense and Kinect are trained separately. Our method outperforms previous methods by a large margin on both cameras without any post-processing. Compared with Fang et al., the previous state-of-the-art method, GSNet improves the performance by 2x on AP metric . Notably, on the most difficult metric , GSNet still achieves a great relative improvement () on all categories. Fig. 5 presents the qualitative results of our network. The top-1 grasp accuracy on three categories are 78.22/76.49, 62.88/57.64 and 28.97/24.04 for Realsense/Kinect input.
We also report the results after simple collision detection using a parallel-jaw gripper model, where all grasps collided with scene points are removed. The results are improved by 1.42/2.31 AP, 1.06/1.79 AP and 0.33/0.77 AP on the three categories respectively.
4 Boosting with Cascaded Graspness Model
We apply the cascaded graspness model (CGM) to GPD, Liang et al. and Fang et al. directly and compare the results with the original methods. For Fang et al., we simply replace ApproachNet with our module. For GPD and Liang et al., we first determine the grasp candidate points using our predicted point-wise graspable landscape, followed by their post processing of Darboux frame estimation and grasp images/clouds classification.
In the middle of Tab. 2, we show the results after adding the CGM. Both the two-step methods and the end-to-end method achieve significant performance gains, proving the effectiveness of cascaded graspness model. Graspable landscapes can not only improve candidate qualities, but also reduce the huge computation time caused by densely sampling.
5 Analysis
In Sec. 4.1 we use graspable FPS to sample seed points from graspable landscapes, while other sampling methods can also be applied to the network. We compare our sampling method with three alternatives: a) random sampling from the whole point cloud; b) FPS from the whole point cloud; c) random sampling from graspable landscapes. Tab. 3 shows the results of the models trained using different sampling methods. FPS outperforms random sampling by at least 4.98 AP and sampling with graspable landscapes improves the results by over 7 AP for both FPS and random sampling, which proves the effectiveness of graspable FPS.
For view selection, we compare graspable PVS with two methods: a) selecting views by surface normal; b) selecting the view with the highest graspness score during training. The results in Tab. 3 show that our method outperforms both alternative strategies. Graspable PVS dynamically selects approach vectors, which provides richer data for model training than other methods.
In Sec. 3.2 we extend object-level graspness scores to cluttered scenes. Tab. 5 shows that sampling from scene-level graspness performs better than object-level counterpart. The representation for graspness score also has multiple choices. We replace the original definition ratio of feasible grasps with mean and maximum grasp quality scores respectively in the calculation of view-wise graspness scores, and the results in Tab. 5 shows that feasible grasp ratio performs the best.
Tab. 6 shows the inference time of our method. Cascaded graspness model achieves a high speed on RealSense/Kinect data, which can also provide accurate sampling for various grasp detection methods. GPD and PointNetGPD take 1s while ours takes only 0.1s.
6 Real Grasping Experiments
We also conduct grasping experiments for cluttered scenes in the real-world setting. The configuration of our experimental setup is illustrated in supplementary materials. The experiments are conducted on a UR-5 robotic arm with an Intel RealSense D435 camera and a Robotiq two-finger gripper. During experiments, we only keep the points on table workspace for speed up.
We conduct grasping experiments in six cluttered scenes. Each scene contains 6-8 objects selected from GraspNet-1Billion. Objects are put together randomly and we repeat the grasping pipeline until the table are cleaned. The success rate is defined as the ratio of object number and attempt number. Tab. 7 reports the grasping performance, which proves the effectiveness of our method. A comparison with other baselines is detailed in supplementary materials.
Conclusion
In this paper, we propose a novel geometrically based quality named graspness. A look-ahead searching method is adopted as our graspness measure and we statistically demonstrate its effectiveness and rationality. An end-to-end network is developed to incorporate graspness into grasp pose detection problem, wherein an independent model learns the graspable landscapes. We conduct extensive experiments and demonstrate the stability, generality, effectiveness and robustness of our graspness model. Large margin of improvements are witnessed for previous methods after equipping with our graspness model, and our final network sets a high record for both accuracy and speed.
This work is supported in part by the National Key R&D Program of China, No.2017YFA0700800, National Natural Science Foundation of China under Grants 61772332 and Shanghai Qi Zhi Institute, SHEITC(2018-RGZN-02046).
References
Appendix A Video Demo and Library
A video demo is attached in the supplementary fileshttps://openaccess.thecvf.com/content/ICCV2021/supplemental/Wang_Graspness_Discovery_in_ICCV_2021_supplemental.zip for real grasping using the results predicted by GSNet which is trained on GraspNet-1Billion. Watch the video “demo.mp4” for more details. Notably, some objects (chain, mesh bag with marbles, slipper, etc.) in the demo are not collected from GraspNet-1Billion and our model shows robustness on novel objects.
GSNet library has been upgraded and integrated into AnyGrasp . See the project websitehttps://graspnet.net/anygrasp for details and more videos.
Appendix B Grasping Experiment Configuration
Fig. 6 shows the configuration of our grasping experiments. A grasp with a high score output by GSNet is chosen and sent to the robotic arm. The program attempts to grab one object each time, and repeat execution until all the objects are cleaned from the table.
Appendix C Robotic Experiments with Baselines
We compare our methods with other baselines in real experiments. Objects are divided into three sets, each containing 10 objects from . These methods are used to remove all the objects in the workspace with single-view point clouds as input. Four repeated experiments are conducted for each object set. Tab. 8 shows the results of different methods, where success rate is defined in Sec. 5.6. GSNet outperforms other methods on all three sets.
Appendix D Details of Grasp Operation Model
Grasp operation model (GOM) in GSNet is designed based on OperationNet in GraspNet baseline model with several improvements. The main differences between the two components are listed as follows.
In , points are cropped and transformed into a cylinder region for each depth bin, which leads to multiple groups with repeated points on one grasp proposal. GOM replaces them with a single cylinder region, where the height is determined by the maximum depth (0.04m) used in our experiments. Depth classification is moved to final output accordingly.
In , points are transformed without scaling. Since all groups shared the same gripper coordinate frame and the transformed coordinates are relative small in abusolute value ( 0.05m), we scaled the points with the cylinder radius (0.05m). The width prediction is modified accordingly.
For each depth bin on one grasp point, OperationNet in directly samples 64 points with only xyz coordinates from the original input (about 20k points). GOM samples 16 points from the seeds (about 1k points). The xyz coordinates are then concatenated with point features output by cascaded grasp model. This modification helps reduce computing overhead of point sampling.
output grasp scores, in-plane rotation angles and gripper widths for each depth bin, and choose the parameters with the highest angle classification scores. In GOM, grasp scores and gripper widths are predicted for each (in-plane rotation)-(approach depth) combination and output the parameters of the combination with the highest grasp score. In addition, output grasp scores and gripper widths are replaced with relative value from 0 to 1.
Appendix E Visualization of Point-wise Graspness
Point-wise graspness predicted by GSNet is visualized in Fig. 7. Regions with higher graspness are annotated with brighter colors. We can see that graspness is not only decided by the object itself, but also influenced by its position. Most of the low-graspness areas are caused by collision with tables, which have higher graspness in single object. Graspness of an object is also influenced by its neighbours. For example, the knife on the banana (Fig. 7d) cut off the contiguous graspness of the latter. The object size is also an important factor. In Fig. 7c, the box has no areas with high graspness although the shape is relative simple. That is because there are few areas for a gripper with width up to 0.1m to grasp on when the box are lying on the table.
Appendix F t-SNE Visualization of Point Features
We visualize the point features output by GSNet. Fig. 8 shows the t-SNE visualization of point features in different test setting, where points are obtained from GraspNet-1Billion dataset and feature vectors are output by GSNet. Tab. 9 details the experimental settings. The points with high graspness are labeled as positive samples, and other points are labeled as negative samples. We can see that graspable points are quite distinguishable from others, which demonstrates the generality of graspness model across different settings.