End-to-End Affordance Learning for Robotic Manipulation

Yiran Geng, Boshi An, Haoran Geng, Yuanpei Chen, Yaodong Yang, Hao Dong

I INTRODUCTION

Learning to manipulate objects is a fundamental problem in RL and robotics. An end-to-end learning approach can explore the reach range of future intelligent robotics. Recently, researchers have shown an increased interest in visual affordance , i.e., a task-specific prior representation of objects. Such representations provide the agents with semantic information of objects, allowing better performance of manipulation.

The existing affordance methods for manipulation have two training stages . For example, VAT-Mart first trains the affordance map with data collected by an RL agent driven by curiosity, and then fine-tunes both the affordance map and the RL agent. In Where2act and many other works , affordance is associated with a corresponding primitive action for each task, such as pushing and pulling. Some recent works learn affordance through human demonstration. A significant drawback of those two-stage methods, which first train the affordance map and then propose action sequence based on the learned affordance, is that the success rate of interaction is highly related to the accuracy of the learned affordance. Any deviations in affordance predictions will significantly reduce the task performance.

In this paper, we investigate learning affordance along with RL in an end-to-end fashion by using contact frequency to represent the affordance. Therefore the affordance is not associated with a specific primitive action but rather the contact information from past manipulation experiences of RL training. In our method, the RL algorithm learns to utilize visual affordance generated from contact information to find the most suitable position for interaction. We also incorporate visual affordance in reward signals to encourage the RL agent to focus on points of higher likelihood. The advantages of end-to-end affordance learning are two-fold: 1) affordance can awaken the agent where to act as an additional observation and be incorporated into reward signals to improve the manipulation policy; 2) learning affordance and manipulation policy simultaneously, without human demonstration or a dedicated data colleting process simplifies the learning pipeline and can migrate to other tasks easily. Additionally, it helps the affordance and manipulation policy to adapt to each other, thus producing a more robust affordance representation.

Using contact information as affordance evidently supports multi-stage tasks, such as picking up an object and then placing it to a proper place, as well as multi-agent tasks, such as pushing a chair with two robotic arms . These two types of tasks are difficult for two-stage affordance methods because they need different pre-defined data collection for each human-defined primitive action. Also, by unifying all interactions as contacts, our method can effectively represent both agent-to-object (A2O) and object-to-object (O2O) interactions, which is hard for other methods.

To test if our method can boost visual-based RL in robotic manipulation, we conducted experiments on eight representative robot tasks, including articulated object manipulation, object pick-and-place and dual arm collaboration tasks. The results showed that our method outperformed all the baselines, including those of RL and the current two-stage affordance methods, and can successfully transfer to the real world. To the best of our knowledge, we are the first to investigate end-to-end affordance learning for robotic manipulation. Our method can be intergrated into visual-based RL to suppport multi-stage tasks and multi-agent tasks without additional annotations or demonstrations.

II Related Work

The recent simulators and benchmarks have boosted the development of manipulation policy learning methods . For rigid object manipulation, there are already robust algorithms handling tasks such as grasping , planar pushing and object hanging . However, it is yet difficult to manipulate articulated objects with multiple parts despite various attempts to approach this problem from different perspectives. For example, UMPNet and VAT-Mart utilized visual observation to directly propose action sequence, while some other studies achieved robust and adaptive control through model prediction. The multi-stage and multi-agent manipulation settings are also challenging for current methods .

II-B Visual Actionable Affordance Learning

Till now, several studies have demonstrated the power of affordance representation on manipulation , grasping , scene classification , scene understanding and object detection . The semantic information in affordance is instructive for manipulation. Some prior affordance learning processes for manipulation, such as Where2Act , VAT-Mart , AdaAfford and VAPO , have two training stages. Specifically, they need to first collect interacting data to pretrain the affordance, and then train the policy based on the affordance. For methods which train affordance and policy simultaneously, however, their affordance learning relies on human demonstration. Unlike them, our method requires neither pre-defined data collection process for different primitive actions / tasks nor any additional human annotations.

II-C Comparison with Related Works

The related works mentioned above studied robotic manipulation in different problem settings, the difference includes observation and annotation. It is hard to compare these works given their distinct settings. Specifically, in the door-opening task, Maniskill utilizes expert demonstrations for imitation learning. However, expert demonstrations are difficult to obtain since they are usually collected by human. Affordance methods, such as Where2Act , VAT-Mart , and VAPO learn the affordance prior to its policy training and are not in an end-to-end fashion. They also output gripper pose as actions which in reality have no guarantee that the gripper can reach the position. Additionally, the related affordance studies mentioned here are designed for single-stage single-agent tasks, such as opening a door or grasping an object, which has no guarantee for multi-stage or multi-agent tasks. However, using contact information for affordance natually allows RL policy to handle multi-stage tasks like picking up an object and placing it to the proper place, and multi-agent tasks like pushing a chair collaboratively by two robotic arms.

Table I compares our method with five representative related works discussed above. Listed works are Where2Act (W2A) , VAT-Mart (VAT) , Maniskill (MSkill) , VAPO and OmniHang (Hang) . "No Demo" means the method does not need any expert demonstrations, such as human-collected trajectories, pre-defined primitive actions and human-designed interaction poses. "No Full Obs" indicates the method does not need state observations of objects such as the coordinate of door handle, since the accurate state of an object is difficult to obtain in real world. "End-to-End" signifies the method trains the policy in an end-to-end fashion, i.e., no multiple training stages are involved and the actions of the policy can be directly applied to agents. "Multi-Stage" infers the method can complete multi-stage tasks which the agent needs to finish multiple dependent tasks sequentially. "Multi-Agent" suggests the method can be adapted to multi-agent tasks where agents need to cooperate one another to finish the work.

III Methods

Visual-based RL is increasingly valued on robotic manipulation tasks, especially those requiring the agent to manipulate different objects with a single policy. Meanwhile, recent studies identified the difficulty of learning observation encoders by RL from high-dimensional inputs such as point clouds and images. In our framework, we tackled this critical problem by exploiting underlying information through a process called "Contact Prediction".

In manipulation settings, contact is the fundamental way humans interact with an object. We believe that physical contact positions during interactions reflect the understanding of crucial semantic information about the object (e.g., a human grasp a handle to open a door because the handle provides the position to apply force).

We proposed a novel end-to-end RL learning framework for manipulating 3D objects. As shown in Fig. 2, our framework is comprised of two parts. 1) Manipulation Module (MAMA Module) is a RL framework which uses the affordance map predicted by a Contact Predictor (CPCP) as an additional observation and reward signal; 2) Visual Affordance Module (VAVA Module) is a per-point scoring network, which uses the contact positions collected from RL training process as the Dynamic Ground Truth (DGTDGT) to indicate the position of interaction.

Concretely, at every time-step tt, the MAMA Module outputs an action ata_{t} based on the robotic arm state sts_{t} (i.e., the angle and angular velocity of each joint) and the affordance map MtM_{t} predicted by the VAVA Module. After each time-step tt, the contact position in RL training is inserted to the Contact Buffer (CBCB). After each kk time-steps, we integrate the data in CBCB to generate the per-point score as the DGTDGT to update the VAVA Module.

III-B Visual Affordance Module: Contact as Prior

During robotic manipulation, physical contacts naturally happen between agent and object, or object and object. As contacts do not relate to any human-defined primitive action such as pull or push, the contact position is a general representation, providing visual prior for manipulation.

The RL training pipeline in Manipulation Module (MAMA Module) continuously interacts with the environment to collect 1) the partial point cloud observation P\mathcal{P}, 2) the contact position under object coordinate. Based on this information, we measure how likely a contact between agent and object (A2O) or between object and object (O2O) is going to happen by the per-point contact frequency as the affordance during the current RL training. The Visual Affordance Module (VAVA Module) then learn to predict the per-point frequency. The training details of VAVA Module is as follow.

Input: Following the prior studies , the input for VAVA Module contains a partial point cloud observation P\mathcal{P}.

Output: The output of the VAVA Module is a per-point affordance map M for each of the point from the input. The map contains A2O affordance and O2O affordance.

Dynamic Ground Truth: To connect the RL pipeline in MAMA Module with the VAVA Module, we use a Contact Buffer CBCB to keep ll record of history contact points, and to compute the DGTDGT. Specifically, each object in the training set has a corresponding CBCB, it records contact positions on the object. To maintain the buffer size, the buffer randomly evicts one record whenever a new record of contact event is inserted. To provide training ground truth for CPCP, we calculate the DGTDGT by first calculating the number of contacts within radius rr from each point on the object point cloud, and then applying normalization to obtain Dynamic Ground Truth DGTDGT. The normalization is as follow:

where DGTtiDGT_{t}^{i} indicates the Dynamic Ground Truth for object ii at time-step tt, CBtiCB_{t}^{i} is the corresponding Contact Buffer.

Training: The CPCP is updated with DGTtiDGT^{i}_{t} as below:

where srtisr_{t}^{i} is the current manipulation success rate on object ii, Pi\mathcal{P}^{i} is the pointcloud of ii-th object and CPt∗\textit{CP}_{t}^{*} is the optimal CP .

III-C Manipulation Module: Affordance as Guidance

Manipulation Module (MAMA Module) is an RL framework able to learn to manipulate objects from scratch. Different from previous methods , our MAMA Module takes advantage of both the reward and observation generated by the VAVA Module.

Input: The input for MAMA Module includes, 1) a point cloud P\mathcal{P} of the real-time environment ; 2) an affordance map MM generated by VAVA Module; 3) the state ss of the robotic arm. The state ss consists of position, velocity and angle of each joint of the robotic arm; 4) a state-based Max-affordance Point Observation (MPOMPO), which indicates the point with the maximum affordance score on P\mathcal{P} .

Output: The output of the MAMA Module is an action aa, which is then executed by the robotic arm. In our setting, the RL policy controls each joint of the robotic arm directly.

Reward from Affordance: We introduce the Max-affordance Point Reward (MPRMPR) into our pipeline, where a point on the point cloud with maximum affordance score predicted by the VAVA Module is selected as the guidance for learning MAMA Module. We use the distance between robot end-effector and this selected point to compute an additional reward in the RL process. We found this reward from affordance could benefit the RL training thus improve the overall performance.

Training: We use Proximal Policy Optimization (PPO) algorithm to train the MAMA Module. To improve the training efficiency by exploiting the high parallelism of our simulator, we deploy kk different objects in the simulator, each object is replicated nn times and given to one or two robotic arms. Hence, there are a total of k×nk\times n environments, each with a robotic arm (or two robotic arms in our multi-agent tasks) interacting with an object, as shown in Fig. 3.

IV Experiment

To evaluate our method, we designed three types of manipulation tasks: single-stage, multi-stage and multi-agent. In all tasks, a robotic arm or two robotic arms are required to complete a specific manipulation task on different objects.

The first type of tasks are single-stage manipulation tasks as follow:

Close Door: A door is initially open to a specific angle. The agent need to close the door completely. We increase the difficulty of this task by applying an additional force on the door attempting to keep the door to the initial position and doubling the friction of the hinge.

Open Door: A door is initially closed. The agent need to open the door to a specific angle. This task can test whether the agent learn to leverage key parts like the handle to open the door, which is challenging.

Push Drawer: A drawer is initially open to a specific distance. similar to close door, the agent need to close the drawer on a cabinet completely.

Pull Drawer: A drawer is initially closed, similar to open door, the agent need to open the drawer to a specific distance.

Push Stapler: A stapler is on the desk, initially open. The agent need to push on the stapler and close it.

Lift Pot Lid: A pot is on the floor with its lid on. The agent need to lift the lid.

To show the agent can learn a policy in a multi-stage task, we use the pick-and-place task as follow:

Pick and Place: An object should be picked up and then placed on a table that already have several random objects on it, both the table and objects are randomly selected from the given datasets. The agent need to place the object stably on the table without collision.

To show our method can be generalized to multi-agent settings, we use the dual-arm-push task as follow:

Dual Arm Push: Two robotic arms need to be controled to push a chair to a specific distance and prevent the chair from falling over.

To make the agent better adapt to the environment, we add a movable base to the arm, allowing the arm to move horizontally within a specific range. The reward designs and other details are listed on our website.

IV-B Dataset and Simulator

We performed our experiments using the Isaac Gym simulator . We used Franka Panda robot arm as the agent for all tasks. Our training and testing data are the subset of the PartNet-Mobility dataset and VAPO dataset . For tasks Close Door and Open Door, we divided the objects with door handles in the StorageFurniture category into four subcategories: one door left, one door right, two door left and two door right. For tasks Pull Drawer and Push Drawer, we divided the objects with door handles in the StorageFurniture category into two subcategories: drawer without door and drawer with door. For tasks Push Stapler and Lift Pot Lid, we chose all Stapler and Pot from PartNet-Mobility dataset. For task Pick and Place, we chose three representative categories of objects from VAPO dataset to pick. We also selected four types of different tables: Round Table, Triangle Table, Square Table and Irregular Table and three daily items from PartNet-Mobility dataset were placed randomly on the table. For task Dual Arm Push, we chose 60 Chairs from PartNet-Mobility dataset.

IV-C Baselines and Ablations

We compared our method with seven baselines:

Where2act : the original method only generates single-stage interaction proposals. To use this method as a baseline in our tasks, we implemented a multi-stage Where2act baseline (up to six steps). The object is gradually altered by pushing or pulling interactions produced by Where2act until the task is completed or the maximum number of steps have been taken. Unlike our own setting, this baseline used a flying gripper instead of a robotic arm.

VAT-Mart : We followed the implementation of paper . Similar to Where2act, we implemented this method in our environment as a baseline with a flying gripper.

RL: we used a point cloud based PPO as our baseline.

RL+Where2act: we replaced the Contact Predictor in our method with a pre-trained Where2act model that can output a per-point actionable score. The parameters in the Where2act model is frozen when training MAMA Module.

RL+O2OAfford and RL+O2OAfford+Where2act: Similar to RL+Where2act, we replaced the O2O affordance map in our method with the map produced by a pre-trained O2OAfford model.

MAPPO: we used a point cloud based multi-agent RL algorithm (MARL): MAPPO as our baseline.

Multi-Task RL : we adapted PPO to the multi-task setting by providing the one-hot task ID as input. To make this method comparable on the test set, both the test set and the training set were used in training process. So this is an oracle baseline.

To further evaluate the importance of different components of our method, we conducted ablation study by comparing our method with five ablations:

Ours w/o MPR: ours without the max-point reward.

Ours w/o MPO: ours without the max-point obversation.

Ours w/o E2E: our method trained by a two-stage procedure. The VAVA Module is trained upon a fixed pretrained MAMA Module. The MAMA Module is then fine-tuned on the freezed VAVA Module.

Ours w/o A2O Map: our method without agent-to-object affordance map in multi-stage tasks.

Ours w/o O2O Map: our method without object-to-object affordance map in multi-stage tasks.

IV-D Evaluation Metrics

For each task, we trained the method (ours, baselines and ablations) on the training set and saved checkpoints every 3200 time-steps within 160,000160,000 total time-steps. After training, we chose the checkpoint with the largest average success rate on training set for comparison, the method was tested on eight different random seeds. We adpoted two metrics to measure the performance:

Average Success Rate (ASR): The ASR is the average of the algorithm’s success rate on all objects in the training / testing dataset.

Master Percentage (MP): We assume a policy is “stable” on an object if it has a success rate of more than 50% on that object. The master percentage is the percentage of objects which the algorithm can success with a probability greater than or equal to 50%50\%. If an algorithm manages to reach a success rate over 50%50\% on a certain object, it is expected to success within two trials.

Due to the page limit, we listed the results of Close Door and Push Drawer and the variance of the reported metric on our website.

IV-E Baseline Comparision and Ablation Study

From Tables II and IV, the results of Where2act and RL show the visual affordance can improve the RL performance. However, our method achieves a more significant improvement over baselines in both training and testing sets. In dual-arm-push, as Table IV shows, our method outperforms both RL and MARL methods.

From all tables, we see the MPO, MPR and E2E components play important roles in our method except that E2E on dual-arm-push. The potential reason is that the predicted max affordance point on the object is changing during object movement, which may influence the RL training. This may be something worth looking into in the future.

Fig. 3 shows the change in affordance maps during end-to-end training and examples of final affordance maps. We can see that as the training proceeds, the affordance map gradually concentrates. More qualitative results can be found on our website.

IV-F Real-world Experiment

We used a digital twin system for real-world experiment: The training process was in simulation, we then used some unseen objects to evaluate our method in real world. The input of the agent has two folds: 1) point cloud input from simulator, 2) agent state input from real world. The actions of the agent were computed upon the combination of the two input sources, and then were applied to the robotic arms both in the simulator and the real world. The experiment settings are shown in Fig.3. Experiments show that our trained model can successfully transfer to the real world. The video and more details can be found on our website https://sites.google.com/view/rlafford/.

V Conclusion

To the best of our knowledge, this the first work that proposes an end-to-end affordance RL framework for robotic manipulation tasks. In RL training, affordance can improve the policy learning by providing additional observation and reward signals. Our framework automatically learns affordance semantics through RL training without human demonstration or other artificial designs dedicated to data collection. The simplicity of our method, together with the superior performance over strong baselines and the wide range of applicable tasks, has demonstrated the effectiveness of learning from contact information. We believe our work could potentially open a new way for future RL-based manipulation developments.

ACKNOWLEDGEMENT

This project was supported by the National Natural Science Foundation of China (No. 62136001). We would like to thank Hongchen Wang, Ruihai Wu, Yan Zhao and Yicheng Qian for the helpful discussion and baseline implementation, and Ruimin Jia for suggestions in paper writing.

References