Learning Dexterous Grasping with Object-Centric Visual Affordances
Priyanka Mandikal, Kristen Grauman
I Introduction
Robot grasping is a vital prerequisite for complex manipulation tasks. From wielding tools in a mechanics shop to handling appliances in the kitchen, grasping skills are essential to everyday activity. Meanwhile, common objects are designed to be used by human hands (see Fig. 1). Hence, there is increasing interest in dexterous, anthropomorphic robotic hands with multi-jointed fingers . Unlike simpler end effectors such as a parallel-jaw gripper, a dexterous hand has the potential for fine-grained manipulation. Furthermore, because its morphology agrees with that of the human hand, in principle it is readily compatible with the many real-world objects built for people’s use. Of particular interest is functional grasping, where the robot should not merely lift an object, but do so in such a way that it is primed to use that object . For instance, picking up a pan by its base for cooking or gripping a hammer by its head for hammering is contrary to functional use.
Learning to perform functional grasping with a dexterous hand is highly challenging. Typical hand models have 24 degrees of freedom (DoF) across the articulated joints, presenting high-dimensional state and action spaces to master. As a result, a reinforcement learning approach trained purely on robot experience faces daunting sample complexity. Existing methods attempt to control the complexity by concentrating on a single task and object of interest (e.g., Rubik’s cube ) or by incorporating explicit human demonstrations . For example, a human “teacher" wearing a glove instrumented with location and touch sensors can supply trajectories for the agent to imitate . While inspiring, this strategy is limited by its expense in terms of human time, the possible need to wear specialized equipment, and the close coupling between the person’s arm/hand trajectory and the target object of interest, which limits generalization.
Towards overcoming these limitations, we propose a new approach to learning to grasp with a dexterous robotic hand. Our key insight is to shift from person-centric physical demonstrations to object-centric visual affordances. Rather than learn to mimic the sequential states/actions of the human hand as it picks up an object, we learn the regions of objects most amenable to a human interaction, in the form of an image-based affordance prediction model. We embed this visual affordance model (a convolutional neural network) within a deep reinforcement learning framework in which the agent is rewarded for touching the afforded regions. In this way, the agent has a “human prior" for how to approach an object, but is free to discover its exact grasping strategy through closed loop experience. Aside from accelerating learning, a critical advantage of the proposed object-centric design is generalization: the learned policy generalizes to unseen object instances because the image-based module can anticipate their affordance regions (see Fig. 1).
Our main contribution is to learn closed loop dexterous grasping policies with object-centric visual affordances. We demonstrate our idea with the 30 DoF AdroitHand model in the MuJoCo physics simulator . We train the visual affordance model from images annotated for human grasp regions . Importantly, image annotations are a much lighter form of supervision than state-action trajectories from today’s status quo expert demonstrations.
In experiments with 40 objects, we show our approach yields significantly better quality grasps compared to other pure RL models unaware of the human affordance prior, even in the presence of sensor and actuation noise. The learned grasping policies are stable under hostile external forces and robust to changes in the objects’ physical properties (mass, scale). Furthermore, our approach significantly improves the sample efficiency of learning process, for a 3 speed up in training despite having no state-action demonstrations. Finally, we show our agent generalizes to pick up object instances never encountered in training. For example, though trained to pick up a hammer, the model leverages partial visual regularities to pick up an axe. Our results offer a promising step towards agents that learn by watching how people use real-world objects, without requiring information about the human operator’s body.
II Related Work
Grasping with planning Traditional analytical approaches use knowledge of the 3D object pose, shape, gripper configuration, friction coefficients, etc. to determine an optimal grasp . With the advent of deep neural nets, learning-based approaches to grasping have gained traction. A common protocol estimates the 6-DoF object pose, followed by model-based grasp planning . Image modules trained to detect successful grasps by parallel jaw grippers can accelerate the robot’s learning . The above strategies are typically employed for simple pick-up actions (not functional grasps) with simple end-effectors like parallel jaw grippers or suction cups, for which a control policy is easier to codify. Some recent work explores related open-loop strategies with complex controllers, but, unlike our method, assumes access to the full 3D model of the objects .
Reinforcement learning for closed-loop grasping Reinforcement learning (RL) models offer a counterpoint to the planning paradigm. Rather than break the task into two steps—static grasp synthesis followed by motion planning—the idea is to use closed-loop feedback control based on visual and/or contact sensing so the agent can dynamically update its strategy while accumulating new observations . Our proposed model is also closed-loop RL and hence enjoys this advantage. However, unlike prior work, we inject an object-centric affordance prior learned from human grasps. It boosts sample efficiency, particularly important for the complex action space of dexterous robotic hands.
Some impressive RL-based systems for dexterous manipulation tackle a specific task with a specific object, like solving Rubik’s cube , shuffling Baoding balls , or reorienting a cube . In contrast, our focus is on grasping and lifting objects, including novel categories, and again our injection of object-centric human affordances is distinct.
Learning manipulation with imitation To improve sample complexity, imitation learning from expert demonstrations is frequently used, whether for non-dexterous or dexterous end effectors. Though advancing the state of the art in dexterous manipulation, the latter approaches rely on “person-centric" human demonstrations with motion capture gloves. Aside from gloves, demonstrations may be captured via teleoperation and video or paired video and kinesthetic demos . In any case, expert demonstrations can be expensive, are specific to the end effector of the demonstration, and their trajectories need not generalize to novel objects.
In contrast, the proposed object-centric affordances sidestep these issues, at the cost of instead supervising the predictive image model. We use supervision from thermal image “hotspots" where people hold objects to use them , though other annotation modes are possible. ContactGrasp leverages thermal image data to rank GraspIt hand poses for a model-based optimization approach. In contrast, our approach 1) learns a closed-loop RL policy for grasping, and 2) incorporates a predictive image-based affordance model that allows generalization to unseen objects. Furthermore, once trained, our policy runs in real-time on new objects, whereas ContactGrasp takes about 4 hours to sample GraspIt poses for each unseen object.
Visual affordances A few methods infer visual affordances for grasping with simple grippers and explore non-robotics affordances . Traditionally, supervision comes from labeled image examples or a robot’s grasp success/failure , while newer work explores weaker modes of supervision from video . Recent work has shown that visual models can help focus attention for a pick and place robot . All of the prior methods make use of simple grippers in an open-loop control setting . To our knowledge, ours is the first work to demonstrate closed-loop RL policies learned with visual affordances.
III Approach
Our goal is to learn dexterous robotic grasping policies influenced by object-centric grasp affordances from images. Our proposed model, called GRAFF for Grasp-Affordances, consists of two stages (Fig. 3). First, we train a network to predict affordance regions from static images (Sec. III-A). Second, we train a dynamic grasping policy using the learned affordances (Sec. III-B). All of our experiments are conducted on a simulated tabletop environment using a 30 DoF dexterous hand as the robotic manipulator (detailed below). We next detail each of these stages.
We first design a perception model to infer object-centric grasp affordance regions from static images. As discussed above, an object-centric approach has the key advantage of providing human intelligence about how to grasp while forgoing demonstration trajectories. Furthermore, by predicting affordances from images, we open the door to generalizing to new objects the robot has not seen before.
Thermal image contact training data We train the affordance model with images with ground truth functional grasp regions obtained from ContactDB . ContactDB contains 3D scans of 50 household objects along with real-world human contact maps captured using thermal cameras. Participants grasped each object using two different post-grasp functional intents—use and hand-off—and a thermal camera on a turntable recorded the multi-point “hotspots" where the object was touched. Our model could alternatively be trained with manual image annotations. Note that our work infers visual affordances on new images, whereas ContactDB is a dataset of actual grasp measurements.
We consider contact maps corresponding to the use intent and exclude objects having bimanual grasps, which yields 16 total objects. Since each object has thermal maps captured from 50 different participants, we use -medoids clustering to obtain a representative thermal map for each object. Specifically, for a given object, we cluster the XYZ values of mesh points with a contact strength value above 0.5 (following ), then take the medoid of the largest cluster as our representative contact map for that object. We port the 3D models into the MuJoCo physics simulator and render them on a tabletop to create an image training set.
For each object, we obtain a set of image-affordance pairs by rendering the 3D object and the 3D contact map, respectively. See Fig. 2a. We rotate each object randomly within a 0-180 range of the camera viewing angle and augment the dataset with varying camera positions. Finally, we obtain a dataset of 15k training pairs, which we divide into an 80:10:10 train/val/test split.
Image affordance prediction model Let represent the domain of object RGB images, and let be the object-centric grasp affordances. Our goal is to learn a mapping that will infer the grasp affordance regions from an individual image. During training, we have labelled (image, mask) pairs . We pose the affordance learning problem as a segmentation task to predict binary per-pixel labels, and approximate with a convolutional neural network. We adapt the Feature Pyramid Network (FPN) to perform semantic segmentation and use an ImageNet-pretrained ResNet-50 as the backbone. See Fig 3a.
We now have a simple but effective model to infer object-centric grasp affordances from static images, which we will use below to guide a dexterous grasping policy. On the ContactDB test split, the segmentation accuracy averages 80.4% in IoU. Fig. 2b shows sample predictions for both ContactDB and 3DNet ( see Sec. IV for dataset info). Our affordance anticipation model is able to predict meaningful functional affordances for novel objects and viewpoints. For example, it faithfully infers graspable handles of saucepan, axe, and pliers despite not having encountered these categories in the training set.
III-B Dexterous Grasping using Visual Affordances
We want a controller that can intelligently process sensory inputs and execute successful grasps for a variety of objects with diverse geometries. Towards this end, we develop a deep model-free reinforcement learning model for dexterous grasping. Our robot model assumes access to visual sensing and proprioception, as well as 3D point tracking. However, the agent does not have access to world dynamics, full object state, or the reward function. Given the large action and state spaces, sample efficiency is a significant challenge. We show how the visual affordance model streamlines policy exploration to focus on object regions most amenable to grasping. See Fig. 3b.
State space The work space of the robot consists of an object positioned on a table at a random orientation. The state space consists of the visuomotor inputs used to train the control policy: (see Fig. 4). The visual input at time consists of an RGB-D image captured by an egocentric hand-mounted camera that translates with the hand but does not rotate. The affordance input is the binary affordance map inferred from the RGB image before the agent moves its hand in view, . The proprioception input consists of the positions and velocities of each DoF in the hand actuator.
The distance input is the distance between the agent’s hand and the object affordance region. We compute it as the pairwise distance between fixed points on the hand and points sampled from the backprojected affordance map. We obtain the latter by backprojecting to 3D points in the camera coordinate system using the depth map at , then tracking those points throughout the rest of the episode. Hence we do not assume access to the full object state (we do not know the object mesh or mass), but we do assume perfect tracking of the affordance region that was automatically detected in the agent’s first video frame. In experiments, we study the effect of substantial tracking failures to relax this assumption. We leave it as future work to incorporate SoTA visual tracking, e.g., by strengthening the segmentation model in the presence of occlusions.
Action space We use a 30-DoF position-controlled anthropomorphic hand from the Adroit platform as our manipulator. It consists of a 24-DoF five-fingered hand attached to a 6-DoF arm. Hence, our action space consists of 30 continuous position values. Delta angles are predicted by sampling from a multivariate Gaussian of unit variance whose mean is returned by the policy .
Reward function The reward function should not only signal a successful grasp, but also guide the exploration process to focus on graspable object regions. To realize this, we combine two rewards: (positive reward when the object is lifted off the table) and (negative reward denoting the hand-affordance contact distance). is computed as the Chamfer distance between the and points described earlier. We also include an entropy maximization term, , to encourage exploration of the action space . Our total reward function is:
Through , the agent is incentivized to explore areas of the object that lie within the affordance region. The object-centric formulation poses no constraints on the hand pose, and can thus be seen as softer supervision than that employed in imitation learning for manipulation which requires kinesthetic teaching or tele-operation .
Implementation details We implement our approach with the architecture shown in Fig. 4. The affordance network is optimized using Dice loss for 20 epochs with a learning rate of and minibatch size of 8. We preprocess the affordance map by computing its distance transform, which helps densify the affordance input. The CNN encoder consists of three 2D convolutional layers with filters of size , and a bottleneck layer of dimension 512, with ReLU activations between each layer. The proprioception and hand-object distance inputs are processed using a 2-layer fully-connected encoder of dimension . For the hand-object contacts, we use and uniformly sampled points. The CNN and FC embeddings are concatenated and further processed (FCs) before predicting the action values. We optimize the network using the Adam optimizer with a learning rate of . The full network is trained using PPO . We train a single policy for all ContactDB objects for 150M agent steps with an episode length of 200 time steps. With each episode being 2 s long, this amounts to 150 hours of learning experience. The coefficients in the reward function (Eq. 1) are set as: . We train for four random seed initializations. All project code is publicly available on the project website.
IV Experiments
Datasets We validate our approach with two datasets: ContactDB and 3DNet . We train a single policy across all 16 objects from ContactDB with one-hand grasps: apple, cell phone, cup, door knob, flashlight, hammer, knife, light bulb, mouse, mug, pan, scissors, stapler, teapot, toothbrush, toothpaste. First we evaluate grasping on these 16 seen objects. Then, we test on 24 novel object meshes from 3DNet, a CAD model database with multiple meshes per category. We use 24 meshes from 9 categories that roughly align with the objects in ContactDB. Four of the 3DNet categories exist in ContactDB (hammer, knife, mug, scissors), and the other five do not (axe, pencil, pliers, saucepan, wrench), making this a good test of generalization.
Comparisons We first devise two pure RL baselines that lack the proposed affordances: (1) No Prior: uses the lifting success and entropy rewards only. (2) CoM: uses the center of mass as a prior, which may lead to stable grasps , by penalizing the hand-CoM distance for . Both pure RL methods use our same architecture (Fig. 4), allowing apples-to-apples comparisons. (3) DAPG: We also compare to DAPG , a hybrid imitation+RL model that uses motion-glove demonstrations. DAPG is trained with object-specific mocap demonstrations collected by us in VR for grasping each ContactDB object (25 demos per object). We stress that DAPG is a strongly-supervised approach with access to full motion trajectories of expert actions, whereas our approach uses inferred object-centric affordances to guide the policy. We train one policy per object for DAPG, allowing it to specialize to each object’s demonstrations; our GRAFF is a single policy trained on all ContactDB objects. A practical advantage of our method is to replace heavy demonstrations (state-action pairs) with image-based affordances.
We stress that the task at hand is dexterous grasp acquisition with a multi-fingered hand, not end effector pose estimation. Accordingly, we focus our comparisons on closed-loop RL methods to pinpoint exactly where our method has impact. Non-RL methods that evaluate only end-effector pose, e.g., as well as methods that directly regress 6-DoF grasp poses for a parallel-jaw gripper followed by grasp execution at that orientation are not applicable in this domain. In short, our idea is quite different—not only in approach (dynamic RL policy vs. pose estimation), but also in problem domain (dexterous manipulation vs. gripper) and learning signal (human use prior vs. solely geometry or robot experience).
Metrics We use three metrics: (1) Grasp Success: For a given episode, a successful grasp has been executed if the object has been lifted off the table by the hand for at least the last 50 time steps (a quarter of the episode length) to allow time to reach the object and pick it up. (2) Grasp Stability: After an episode completes, we apply perturbing forces of Newton in six orthogonal directions to the object. If the object remains held, the grasp is deemed stable. (3) Functionality: We report the percentage of successful grasps in which the hand lies close to the GT affordance region (measured using Chamfer distance). We execute 100 episodes per object with the objects placed at different orientations ranging from [0,180°], and report mean and std dev of the metrics across all models trained with four random seeds.
Noisy sensing and actuation When deploying trained policies on real robots, we might encounter a number of non-ideal circumstances owing to faulty sensor readings or imperfect robot executions. To better model such realistic settings, we incorporate noise into our training and testing regimes in simulation . Following , we apply additive Gaussian noise on the proprioceptive sensor readings (robot joint angles and angular velocities), object tracking points, and robot actuation. We also apply pixel perturbations in the range and clip all pixel values between . We train all versions of all methods under these noisy conditions (see Fig. 6, discussed below). We additionally test our method with a tracking failure model that freezes the track for 20 frames at random intervals and find that mean success rate is still reasonably high at 54%. These empirical results indicate that GRAFF can remain fairly robust to tracking, sensing, and actuation failures that real robotic systems encounter.
While our lab does not have access to a real dexterous robotic hand, we believe that (like in ) the high quality physics-based simulator together with these noise models offers a meaningful study. We are also encouraged by recent successes translating policies trained only in simulation to real-world dexterous robots using domain randomization .
Grasping seen objects from ContactDB Fig 6 shows the results on ContactDB. We show the mean and std dev over different seeds. GRAFF (our model) outperforms both pure RL baselines consistently on all metrics. Our gains persist with noisy sensing and actuation as well. Fig. 5a shows qualitative examples; please see the video for full episodes and failure cases. GRAFF can successfully grasp the objects at the anticipated affordance regions (handle of pan, mug, teapot, knife, scissors), while the baselines fail to grasp objects with complex geometries (pan, mug, teapot). This shows the effectiveness of the affordance-guided policy in learning stable functional grasps.
Our method also fares well compared to the more intensely supervised imitation+RL method DAPG , outperforming it on all metrics. This is a very encouraging result: not only does our method outperform its RL counterparts, but it is also competitive with a method that leverages expert trajectories. We found that DAPG can be vulnerable to imperfect demonstrations, yet in practice expert demos can be difficult to obtain (e.g., with mocap gloves). We also ran DAPG with the authors’ provided demonstrations for a ball object, which performed slightly worse than the policies trained with the object-specific DAPG policies in Fig. 6.
Robustness to physical properties of the objects To evaluate robustness to changes in object properties, we apply our policy to a range of object masses and scales not encountered during training. Fig. 7 shows 3D plots. Here, and are the mass and scale values used during training. GRAFF remains fairly robust across large variations, which we attribute to GRAFF’s preference for stable human-preferred regions.
Grasping unseen objects from 3DNet Next we push the robustness challenge further by requiring the agent to generalize its grasp behavior to objects it has not encountered before (the 24 3DNet objects). We first render the objects and predict grasp affordances (cf. Fig. 2b). We then apply the policy trained on ContactDB to execute grasps. For DAPG, we use the trained policy of the ContactDB object that is closest in shape to the one in 3DNet. Fig. 6b shows the results. We outperform all three baselines by a large margin in both grasp success and stability. The key factor is our visual affordance idea: the anticipation model generalizes sufficiently to new object shapes so as to provide a useful object-centric prior. Note that we cannot report the functionality score for 3DNet since there are no GT affordances for these objects. Fig. 5b shows sample grasps. GRAFF successfully executes grasps at the anticipated affordance regions (e.g., handle of axe and finger rings on scissors), whereas the baselines may grasp the scissors at its blades or fail to lift the axe.
Ablations The No Prior baseline above (average success rate 16%) is a key ablation for our method. To further tease out elements of our approach, we compare our model with only the hand-object contact term, adding the affordance map, and adding the reward. Average success rates are 39%, 42%, 63%, respectively. Thus our full model is most effective.
Training time Fig. 8 shows the grasp success rate versus number of training samples. Not only does our model learn more successful policies, it also has a sharper learning curve. While the pure RL baselines reach a maximum success rate of 30% in 150M training samples (150 hours of robot experience), our method reaches the same success rate in only 50M samples (50 hours)—a 3 speedup. Recall that 150 hours is a one-time cost: we train a single policy for all ContactDB objects, and simply execute that trained policy when encountering an unseen object. Thus our affordance prior meaningfully improves sample efficiency for dexterous grasping, while convincingly outperforming the other pure RL methods. GRAFF also learns faster and performs better than the more heavily supervised DAPG for the same number of training samples.
V Conclusion
Our approach learns dexterous grasping with object-centric visual affordances. Breaking away from the norm of expert demonstrations, our GRAFF approach uses an image-based affordance model to focus the agent’s attention on “good places to grasp". To our knowledge, ours is the first work to demonstrate closed-loop RL policies learned with visual affordances. The key advantages of our design are its learning speed and ability to generalize policies to unseen (visually related) objects. While there is much more work to do in this direction, we see the results as encouraging evidence for manipulation agents learning faster with more distant human supervision. In future work, we are interested in expanding the visual affordance model, modeling the multi-modal distribution of viable affordance regions, and investigating manipulations beyond grasping (e.g., open, sweep).
Acknowledgements UT Austin is supported in part by ONR PECASE N00014-15-1-2291. Thanks to Samarth Brahmbhatt, Vikash Kumar, and Emo Todorov for helpful discussions.