Benchmarking Structured Policies and Policy Optimization for Real-World Dexterous Object Manipulation
Niklas Funk, Charles Schaff, Rishabh Madan, Takuma Yoneda, Julen Urain De Jesus, Joe Watson, Ethan K. Gordon, Felix Widmaier, Stefan Bauer, Siddhartha S. Srinivasa, Tapomayukh Bhattacharjee, Matthew R. Walter, Jan Peters
I INTRODUCTION
Dexterous manipulation is a challenging problem in robotics that has impactful applications across industrial and domestic settings. Manipulation is challenging due to a combination of environment interaction, high-dimensional control and required exteroception. As a consequence, designing high-performance control algorithms for physical systems remains a challenge. Due to the complexity of the problem, data-driven approaches to dexterous manipulation are a promising direction. However, due to the high cost of collecting data with a physical manipulator and the sample efficiency of current methods, the robot learning community has primarily focused on simulated experiments and benchmarks . While there have been successes on hardware for various manipulation tasks , the hardware and engineering cost of reproducing these experiments can be prohibitive to most researchers.
In this work, we investigate several approaches to dexterous manipulation using the TriFinger platform , an open-source manipulation robot. This research was motivated by the ‘Real Robot Challenge’ (RRC),See https://real-robot-challenge.com. where the community was tasked with designing manipulation agents on a farm of physical TriFinger systems. A common theme among the successful solutions is their use of structured policies, methods that combine elements of classical robotics and modern machine learning to achieve reliability, sample efficiency and high performance. We summarize the solutions here and analyse their performance through ablation studies to understand which aspects are important for real-world manipulation and how these characteristics can be appropriately benchmarked.
The main contributions of our work are as follows. We introduce three independent structured policies for tri-finger object manipulation and two data-driven optimization schemes. We perform a detailed benchmarking and ablation study across policy structures and optimization schemes, with evaluations both in simulation and on several TriFinger robots. The paper is structured as follows: Section II discusses prior work, Section III introduces the TriFinger platform, Section IV describes the structured policies, Section V presents the approaches to structured policy optimization, Section VI details the experiments, and Section VII discusses the findings.
II RELATED WORK
Much progress has been made for manipulation in structured environments with a priori known object models through the use of task-specific methods and programmed motions. However, these approaches typically fail when the environment exhibits limited structure or is not known exactly. The contact-rich nature of manipulation tasks naturally demands compliance, which can be achieved through soft materials on the hardware level . The need for compliant motion has also led to the development of control objectives like force control , hybrid position/force control , and impedance control . Operational space control has been monumental in developing compliant task space controllers that use torque control.
Data-driven Manipulation Given the complexity of object manipulation due to both the hardware and task, data-driven approaches are an attractive means to avoid the need to hand-design controllers, while also providing potential for improved generalizability and robustness. For control, a common approach is to combine learned models with optimal control, such as guided policy search and model predictive control (MPC) . Model-free reinforcement learning (RL) has also been applied to manipulation , including deep RL , which typically requires demonstrations or simulation-based training due to sample complexity. Data-driven methods can also improve grasp synthesis .
Structured Policies Across machine learning, inductive biases provide a means to introduce domain knowledge to improve sample efficiency, interpretability and reliability. In the context of control, inductive biases have been applied to enhance models for model-based RL , or to policies to simplify or improve policy search. Popular structures include options , dynamic movement primitives , autoregressive models , MPC and motion planners . Structure also applies to the action representation, as acting in operational space, joint space or pure torque control affects how much classical robotics can be incorporated into the complete control scheme .
Residual Policy Learning Residual policy learning (RPL) provides a way to enhance a given base control policy by learning additive, corrective actions using RL . This allows well developed tools in robotics—such as motion planning, PID controllers, etc.—to be used as an inductive bias for learned policies in an RL setting. This combination improves sample efficiency and exploration, leading to policies that can outperform classical approaches and pure RL.
Bayesian Optimization for Control Bayesian optimization (BO) is a sample-efficient black-box optimization method that leverages the epistemic uncertainty of a Gaussian process model of the objective function to guide optimization. It can be used for hyperparameter optimization, policy search, sim-to-real transfer and grasp selection .
Benchmarking Manipulation Early work on benchmarking dexterous manipulation was mainly focused on simulation . Lately, there has been an increased interest in real-world setups. Yet, most require large, expensive hardware and complex software solutions, including perception modules . In contrast, this work builds on the RRC, which provides remote access to TriFinger platforms, allowing for an exclusive focus on developing effective manipulation strategies.
III TRI-FINGER OBJECT MANIPULATION
Fig. 1 shows the TriFinger robot . The inexpensive, compact and open-source platform can be easily recreated and serves as the basis for the RRC, which aims to promote state-of-the-art research in dexterous manipulation on real hardware.
About the Robot The robot consists of three identical “fingers” with three degrees of freedom each. Its robust design, together with several safety measures, allow for running learning algorithms directly on the real robot, even if they send unpredictable or random commands. The robot can be controlled with either torque or position commands at a rate of 1 kHz. It provides measurements of angles, velocities and torques of all joints at the same rate. A vision-based tracking system provides the pose of the manipulated object at a frequency of 10 Hz. Users can interact with the platform by submitting experiments and downloading the logged data.
The Real Robot Challenge The Real Robot Challenge 2020 involved three phases using the TriFinger robot: simulation (Phase 1), and the real robot manipulating a cube (Phase 2), and a cuboid (Phase 3). While our methods were used in all three phases, we focus on Phase 2 because it involved a real robot and a simpler object that afforded much more robust and reliable pose estimation, which is crucial for a fair comparison of the approaches. This phase tasks the robot with moving a cube (Fig. 1) from the center of the arena to a desired goal pose. The cube weighs about 94 g, has a side-length of 65 mm, a structured surface to facilitate grasping, and differently colored sides to help vision-based pose estimation.
The phase is subdivided into four difficulty levels (L1–L4). We focus on the final two, which involve reaching a goal position (L3) and pose (L4) sampled from anywhere in the workspace. Thus, for L3 the reward only reflects the position error of the cube and is computed as a normalized weighted sum of error between the actual and goal positions: , with and the range on the x/y-plane and z-axis, respectively. For L4, the orientation error is computed as the normalized magnitude of the rotation (given as quaternion) between actual and goal orientation . Thus, the reward is given by .
This paper benchmarks the solutions of three independent submissions to the challenge. Table I shows their performance.
IV STRUCTURED POLICIES
This section describes the structured controllers considered as baselines. The controllers share a similar high-level structure and can be broken into three main components: cube alignment, grasping, and movement to the goal pose. We will discuss the individual grasp and goal approach strategies, and then briefly mention cube alignment. For a visualization of each grasping strategy, see Fig. 2. For a more detailed discussion of each controller and alignment strategy, please see the reports submitted for the RRC competition: motion planning , Cartesian position control , and Cartesian impedance control . The code is publicly availablehttps://github.com/cbschaff/benchmark-rrc and demo videos are added in the supplementary material.
When attempting to manipulate an object to reach a desired goal pose, the grasp must be carefully selected such that it can be maintained throughout the manipulation. Many grasps that are valid at the current object pose may fail when moving the object. To avoid using such a grasp, we consider several heuristic grasp candidates and attempt to plan a path from the initial object pose to the goal pose under the constraint that the grasp is maintained at all points along the path. Path planning involves first selecting a potential grasp and then running a rapidly exploring random tree (RRT) to generate a plan in task space, using the grasp to determine the fingers’ joint positions. If planning is unsuccessful within a given time frame, we select another grasp and retry. We consider two sets of heuristic grasps, one with fingertips placed at the center of three vertical faces and the other with two fingers on one face and one on the opposite face. In the unlikely event that none of those heuristic grasps admits a path to the goal, we sample random force closure grasps until a plan is found. Throughout this paper we refer to this method for determining a grasp as Planned Grasp (PG).
To move the object to the goal pose, this approach then simply executes the motion plan by following the waypoints for all fingers using a PD position controller without requiring any further cube pose estimates. After execution, to account for errors such as slippage or an inaccurate final pose, we iteratively append waypoints to the plan in a straight line path from the current object position to the goal position.
IV-B Cartesian Position Control (CPC)
We build upon these linear force commands to create position-based motion primitives. Given a target position for each fingertip, we construct a feedback controller with tuned PID gains coupled with some minor adjustments in response to performance changes. This approach works well in simulation, but for the real robot, it results in the fingers getting stuck in intermediate positions in some cases. Based on the limited interaction afforded by remote access to the robots, this could be attributed to static friction causing the motors to stop. We use a simple gain scheduling mechanism that varies the gains exponentially (up to a clipping value) over a specified time interval, which helps to mitigate this degradation in performance by providing the extra force required to keep the motors in motion.
Triangle Grasp (TG) The above controller is combined with a grasp that places fingers on three of the vertical faces of the cube. The fingers are placed such that they form an equilateral triangle . This ensures that the object is in force closure and that the fingers can easily apply forces on the center of mass of the cube in any direction.
IV-C Cartesian Impedance Control (CIC)
Third, we present a Cartesian impedance controller (CIC) . Using CIC enables natural adaptivity with the environment for object manipulation by specifying second-order impedance dynamics. This avoids having to learn the grasping behaviour through extensive experience and eludes complex trajectory optimization that must incorporate contact forces and geometry. Avoiding such complexity results in a controller that has adequate baseline performance on the real system that can then be further optimized.
For the desired Cartesian position of the fingertip , we define to be the error between this tip position and a reference position inside the cube , i.e., . We then define an impedance controller for , a second-order ODE that can be easily interpreted as a mass-spring-damper system with parameters . The damping factor was zeroed for more robustness to the measurement noise that is present in the vision-based estimate of the cube’s pose. Converting this Cartesian space control law back to joint coordinates results in , where denotes the torques to be applied to finger .
To perform cube position control, we follow the ideas proposed by Pfanne et al. and design a proportional control law that perturbs the center of cube based on the goal position : .
Since the above components do not consider the fingers as a whole, they were limited in controlling the orientation of the cube. Contact forces were also passively applied rather than explicitly considered. To incorporate these additional considerations, we superimpose four torques. These include the previously discussed position control and gravity compensation as well as three contact and rotational terms that we describe next, such that .
To also allow one to directly specify the force applied by each finger, we introduce an additional component , where is the force applied by finger . We chose to be in the direction of the surface normal of the face where finger touches the cube (). However, in order to not counteract the impedance controller, the resulting force of this component should be zero.
By solving for , this is ensured. All previous components ensure a stable grasp closure. This is essential for the following orientation control law. Neglecting the exact shape of the cube, we model the moment that is exerted onto the cube as , where denotes the vector pointing from the center of the cube towards the finger position, is the respective skew-symmetric matrix, and an additional force that should lead to the desired rotation. The goal is now to realize a moment proportional to the current rotation errors, which are provided in the form of an axis of rotation and its magnitude . Thus, the control law yields . We achieve by solving for .
Center of Three Grasp (CG) The above controller is combined with a grasp that places the fingers in the center of three of the four faces perpendicular to the ground plane. The three faces are selected based on the task and goal pose with respect to the current object pose. When only a goal position is specified (L1–L3) the face that is closest to the goal location is not assigned any finger, ensuring that the cube can be pushed to the target. For L4 cube pose control, two fingers are placed on opposite faces such that the line connecting them is close to the axis of rotation. To avoid colliding with the ground, the third finger is placed such that an upward movement will yield the desired rotation.
IV-D Cube Alignment
While we will focus our experiments on the above methods, another component was necessary for the teams to perform well in the competition. Moving an object to an arbitrary pose often requires multiple interactions with the object itself. Therefore, teams had to perform some initial alignment of the cube with the goal pose. All three teams independently converged on a sequence of scripted motion primitives to achieve this. These primitives consisted of heuristic grasps and movements to (1) slide the cube to the center of the workspace, (2) perform a degree rotation to change the upward face of the cube, and (3) perform a yaw rotation. Following this sequence, the cube is grasped and moved to the goal pose.
V POLICY OPTIMIZATION
In addition to the approaches described above, the teams experimented with two different optimization schemes: Bayesian optimization and residual policy learning. In this section, we will briefly introduce these methods.
In this work, we learn residual controllers on top of the three structured approaches defined above. To do this, we use soft actor-critic (SAC) , a robust RL algorithm for control in continuous action spaces based on the maximum entropy RL framework. SAC has been successfully used to train complex controllers in robotics .
VI EXPERIMENTS
To provide a thorough benchmark of the above methods, we perform a series of detailed experiments and ablations in which we test the contribution of different components on L3 and L4 of the RRC.
Experiment Setup In our experiments, we will be comparing different combinations of grasp strategies, controllers, and optimization schemes in a PyBullet simulation environment and on the TriFinger platform. For each combination, we report the reward and final pose error averaged over several trials, as well as the fraction of trials the object is dropped. In simulation, we provide each method with the same initial object poses and goal poses. On the TriFinger platform, we cannot directly control the initial pose of the object, so we first move it to the center of the workspace and assign the goal pose as a relative transformation of the initial pose. All methods are tested with the same set of relative goal transformations. To isolate the performance of the grasp strategies and controllers, we initialize the experiments so that no cube alignment primitives are required to solve the task.
Mix and Match The choice of grasp heuristic and control strategy are both crucial to success, but it is hard to know how much each component contributes individually. To test each piece in isolation, we “mix and match” the three grasp heuristics with the three structured controllers and report the performance of all nine combinations. Each combination is tested on L4 and the results are averaged over trials. We note that the speed and accuracy of the initial cube alignments had a large impact on the competition reward. To account for this and to test the robustness of these approaches to different alignment errors, we evaluate each method with three different initial orientation errors degrees. When the MP controller is paired with a grasp strategy other than PG, a motion plan is generated using that grasp.
Bayesian Optimization All of our approaches are highly structured and rely on a few hyperparameters. We investigate whether BO can improve the performance of the controllers by optimizing them for both L3 and L4. For all experiments, we initialize the optimization with four randomly sampled initial sets of parameters and run optimization iterations. We do not explicitly exploit any information from our manually tuned values. Instead, the user only has to specify intervals. This sample-efficient optimization process only takes about hours to complete on the real system. After the optimized parameters are found, each controller is tested over trials on the TriFinger platform with both the manually tuned parameters and the BO parameters. For our proposed approaches, we optimize the following values: CIC: gains and reference position ; CPC: gains, including values for the exponential gain scheduling; and MP: hyperparameters that control the speed of movement to the goal location on the planned path.
Residual Policy Learning In these experiments, we investigate to what extent RPL can be used to improve the performance of our three controllers. To do this, we train a neural network policy to produce joint torques in the range (the maximum allowed joint torque on the system is ), which are then added to the actions of the base controller. The policies are trained on L3 in simulation for 1M timesteps using SAC . The reward function consists of a combination of the competition reward as well as terms for action regularization, maximizing tip force sensor readings, and maintaining a grasp of the object. The policy architecture is as follows: The observation and base action are separately embedded into a -dimensional space. These embeddings are then concatenated and passed to a three-layer feed-forward network that outputs a Gaussian distribution over torque actions. Actions are sampled from this distribution, squashed using a function, and then scaled to fit in the range. We evaluate MP-PG, CPC-TG, and CIC-CG in simulation over trials for L3. We then test their ability to transfer to the real system with another trials.
VII RESULTS
In this section, we describe the results of the experiments described above. In total, we conducted more than 20k experiments on the real system.
We find that CPC drops the cube more frequently than other approaches, and its performance varies significantly depending on the choice of grasp. This is because it drives the fingertip positions to the goal pose without considering whether the grasp can be maintained. Yet, in the cases in which it can retain its grasp, CPC achieves much lower orientation errors than CIC or MP. While CIC is similar to CPC, it is more robust to drops at the expense of accuracy, because it explicitly considers the forces the fingertips need to apply to achieve a desired motion. MP is the most robust against drops and grasp choices, because it attempts to move only to locations in which the selected grasp can be maintained. This comes at the expense of orientation errors, since the planner may have only found a point near the goal for which the grasp is valid.
When comparing CG and TG, we find that TG performs better in terms of both orientation error and drop rate. We hypothesize that this is the case because the triangle shape facilitates applying forces to the cube in all directions. The benefits of the planned grasp (PG) become apparent when the initial orientation error is large, improving the drop rate across all three controllers. This verifies the intuition for grasp planning that, when the required orientation change is large, it helps to carefully select a grasp that is feasible both at the initial pose and near the goal pose.
Bayesian Optimization Table III depicts the results from running BO on the real system to optimize the controllers’ hyperparameters. As can be seen on the left-hand side, for the L3 experiments, the newly obtained hyperparameters significantly improve the policies’ mean reward as well as the mean position errors. Furthermore, although we only averaged across five rollouts during training, the improvements persist when evaluating the policies on newly sampled goal locations. Visually inspecting the rollouts, we conclude that running BO results in higher gains such that the target locations are reached more quickly.
Repeating the same experiment for L4 yields the results presented on the right-hand side of Table III. As shown in the table, only the performance for the CPC control strategy can be improved significantly. Comparing the two sets of parameters, the BO algorithm suggests using lower gain values. This results in a more stable and reliable control policy and increases performance. In general, even though we do not provide the manually obtained parameters as a prior to the BO algorithm, for the other two approaches, the performance of the optimized parameters is still on par with the manually tuned parameters. We reason that for the MP approach, the two hyperparameters might not provide enough flexibility for further improvements, while the CIC controller might have already reached its performance limits. The results indicate that BO is an effective tool to optimize and obtain performant hyperparameters for our approaches and mitigates the need for tedious manual tuning.
Residual Policy Learning Table IV shows the performance of our control strategies with and without RPL. We find that the effectiveness of RPL is limited to improving only the MP controller. When inspecting the results we find that residual control is able to help the MP policy to maintain a tight grasp on the cube, preventing catastrophic errors such as dropping the cube and triggering another round of planning. Surprisingly, the learned policies are able to transfer to the real system without any additional finetuning or algorithms such as domain randomization . For the MP controller, we find improved grasp robustness and a reduction in the drop rate. Additionally, for CPC and CIC, it is surprising that performance was not more adversely affected given the lack of improvement in simulation. We hypothesize that this is the case for two reasons. First, the torque limits on the residual controller are small and may not be able to cause a collapse in performance. Second, the base controllers provide increasing commands as errors are made that work to keep the combined controller close to the training data distribution. For example, the MP controller has a predefined path and the PD controller that follows that path provides increasing commands as deviations occur.
From these experiments, we conclude that RPL may be effective when small changes can be made, such as helping to maintain contact forces, in order to improve the robustness of a controller, and that transferring a residual controller from simulation is much easier than transferring a pure RL policy.
Challenge Retrospective Overall, the newly obtained results (e.g., Table III) differ only slightly from to the scores reported during the competition (Table I). CPC-TG yields the best results on L3, but the team using MP-PG was able to win, in part because they implemented more reliable primitives for cube alignment. Through exploiting insights from this approach, we assume that CPC-TG could perform similarly. Nevertheless, the robustness of MP-PG might still outweigh the gains on the successful runs of the reactive policies, especially with increasing task difficulty. During the third phase, all methods were also evaluated with a small cuboid and are thus not limited to a particular object type, as long as the grasping strategies are adapted accordingly. Yet, smaller objects make the problem of vision-based state estimation considerably more difficult, resulting in even further advantages of the MP-PG approach. Looking broadly at the RRC and our results, it may be worth adjusting the RRC protocol to also provide an evaluation of the individual components (e.g., grasp planning) to yield more informative results.
VIII CONCLUSION AND OUTLOOK
In this work, we present three different approaches to solving the tasks from the RRC. We perform extensive experiments in simulation and on the real platform to compare and benchmark the methods.
We find that using motion planning provides the best trade-off between accuracy and reliability. Compared to motion planning, the two reactive approaches vary in both reliability and accuracy. Concerning grasp selection, the results show that using the triangle grasp yields the best performance.
We further show the effectiveness of running Bayesian optimization for hyperparameter tuning. Especially for L3, the performance can be increased significantly across all approaches. Augmenting the structured methods with a learned residual control policy can improve performance when small changes to a controller are beneficial. Surprisingly, we also find that transferring the learned residual controllers required no finetuning on the real system or other techniques to cross the sim-to-real gap, although applying those techniques is likely to be beneficial.
We hope that our work serves as a benchmark for future competitions and dexterous manipulation research using the TriFinger platform.