In-Hand Object Rotation via Rapid Motor Adaptation

Haozhi Qi, Ashish Kumar, Roberto Calandra, Yi Ma, Jitendra Malik

Introduction

Humans are remarkably good at manipulating objects in-hand – they can even adapt to new objects of different shapes, sizes, mass and materials with no apparent effort. While several works have shown in-hand object rotation with real-world multi-fingered hands for a single or a few objects , truly generalizable in-hand manipulation remains an unsolved challenge of robotics.

In this paper, we demonstrate that it is possible to train an adaptive controller capable of rotating diverse objects over the zz-axis with the fingertips of a multi-fingered robot hand (Figure 1). This task is a simplification of the general in-hand reorientation task, yet still quite challenging for robots since at all times the fingers need to maintain a dynamic or static force closure on the object to prevent it from falling (as it can not make use of any other supporting surface such as the palm).

Our approach is inspired by the recent advances in legged locomotion using reinforcement learning. The core of these works is to learn a compressed representation of different terrain properties (called extrinsics) for walking, which is jointly trained with the control policy. During deployment, the extrinsics is estimated online and the controller can perform rapid adaptation to it. Our key insight is that, despite the diversity of real-world objects, for the task of in-hand object rotation, the important physical properties such as local shape, mass, and size as perceived by the fingertips can be compressed to a compact representation. Once such a compressed representation (extrinsics) of different objects is learned, the controller can estimate it online from proprioception history and use it to adaptively manipulate a diverse set of objects.

Specifically, we encode the object’s intrinsic properties (such as mass and size) to an extrinsics vector, and train an adaptive policy with it as an input. The learned policy can robustly and efficiently rotate different objects in simulation environments. However, we do not have access to the extrinsics when we deploy the policy in the real world. To tackle this problem, we use the rapid motor adaptation to learn an adaptation module which estimate the extrinsic vector, using the discrepancy between observed proprioception history and the commanded actions. This adaptation module can also be trained solely in simulation via supervised learning. The concept of estimating physical properties using proprioceptive history has been widely used in locomotion but has not yet been explored for in-hand manipulation.

Related Work

Classic Control for In-Hand Manipulation. Dexterous in-hand manipulation has been an active research area for decades . Classic control methods usually need an analytical model of the object and robot geometry to perform motion planning for object manipulation. For example, rely on such a model to plan finger movement to rotate objects. assumes objects are piece-wise smooth and use finger-tracking to rotate objects. demonstrate reorientation of different objects in simulation by generating trajectories using optimization. There have been also attempts to deploy systems in the real-world. For example, calculates precise contact locations to plan a sequence of contact locations for twirling objects. plans over a set of predefined grasp strategies to achieve object reorientation using two multi-fingered hands. use throwing or external forces to perturb the object in the air and re-grasp it. Recently, works such as do in-grasp manipulation without breaking the contact. demonstrates complex object reorientation skills using a non-anthropomorphic hand by leveraging the compliance and an accurate pose tracker. The diversity of the objects they can manipulate is still limited due to the intrinsic complexity of the physical world. In contrast to traditional control approaches which may use heuristics or simplified models to solve this task, we instead use model-free reinforcement learning to train an adaptive policy, and use adaptation to achieve generalization.

Reinforcement Learning for In-Hand Manipulation. To get around the need of an accurate object model and physical property measures, in the last few years, there has been a growing interest in using reinforcement learning directly in the real-world for dexterous in-hand manipulation. learns simple in-grasp rolling for cylindrical objects. learns a dynamics model and plan over it for rotating objects on the palm. use human demonstration to accelerate the learning process. However, since reinforcement learning is very sample inefficient, the learned skills are rather simple or have limited object diversity. Although complex skills such as re-orientating a diverse set of objects and tool use can be obtained in simulation, transferring the results to real-world remains challenging. Instead of directly training a policy in the real world, our approach learns the policy entirely in the simulator and aims to directly transfer to the real world.

Sim-to-Real Transfer via Domain Randomization. Several works aim to train reinforcement learning policies using a simulator and directly deploy it in a real-world system. Domain randomization varies the simulation parameters during training to expose the policy to diverse simulation environments so that they can be robustly deployed in the real-world. Representative examples are and . They leverage massive computational resources and large-scale reinforcement learning methods to learn an agile object reorientation skills and solve Rubik’s Cube with a single robot hand. However, they still focus only on manipulating a limited number of objects. learns a finger-gaiting behavior efficiently and transfers to a real robot when hand facing downwards, but they do not only use fingertips and the objects considered are all cubes. Our approach focuses on generalization to a diverse set of objects and can be trained within a few hours.

Sim-to-Real via Adaptation. Instead of relying on domain randomization which is agnostic to current environment parameters, performs system identification via initial calibration, or online adaptive control to estimate the system parameters for Sim-to-Real Transfer. However, learning the exact physical values and alignment between simulation and real-world may be sub-optimal because of the intrinsic inaccuracy of physics simulations. An alternative way is to learn a low dimensional embedding which encodes the environment parameters which is then used by the control policy to act. This paradigm has enabled robust and adaptive locomotion policies. However, it is not straightforward to apply it on the in-hand manipulation task. Our approach demonstrates how to design the reward and training environment to enable a natural and stable controller that can transfer to the real world.

Rapid Motor Adaptation for In-Hand Object Rotation

An overview of our approach is shown in Figure 2. During deployment (Figure 2, bottom), our policy infers a low-dimensional embedding of object’s properties such as size and mass from proprioception and action history, which is then used by our base policy to rotate the object. We first describe how we train a base policy with object property provided by a simulator, then we discuss how to train an adaptation module that is capable of inferring these properties.

where rrot≐max⁡(min⁡(ω⋅k^,rmax⁡),rmin⁡)r_{\rm rot}\doteq\max(\min({\bm{\omega}}\cdot\hat{\mathbf{k}},r_{\max}),r_{\min}) is the rotation reward, rpose≐−∥q−qinit∥22r_{\rm pose}\doteq-\left\lVert\mathbf{q}-\mathbf{q}_{\rm init}\right\rVert_{2}^{2} is the hand pose deviation penalty, rtorque≐−∥τ∥22r_{\rm torque}\doteq-\left\lVert{\bm{\tau}}\right\rVert_{2}^{2} is the torque penalty, rwork≐−τTq˙r_{\rm work}\doteq-{\bm{\tau}^{T}\bm{\dot{q}}} is the energy consumption penalty, and rlinvel≐−∥v∥22r_{\rm linvel}\doteq-\left\lVert{\bm{v}}\right\rVert_{2}^{2} is the object linear velocity penalty. Note that in contrast to which explicitly encourages at least three fingertips to always be in contact with the object, we do not enforce any heuristic finger gaiting behaviour. Instead, a stable finger gaiting behaviour emerges from the energy constraints and the penalty on deviation from the initial pose.

A good training environment has to provide enough variety in simulation to enable generalization in the real world. In this work we find that using cylinders with different aspect ratios and masses provides such variety. We uniformly sample different diameters and side lengths of the cylinder.

We initialize the object and the fingers in a stable precision grasp. Instead of constructing the fingertip positions as in , we simply randomly sample the object position, pose, and robot joint position around a canonical grasp until a stable grasp is achieved. We also randomize the mass, center of mass, and friction of these objects (see appendix for the details).

2 Adaptation Module Training

We cannot directly deploy the learned policy π\pi to the real world because we do not directly observe the vector et\mathbf{e}_{t} and hence we cannot compute the extrinsics zt\mathbf{z}_{t}. Instead, we estimate the extrinsics vector z^t\hat{\mathbf{z}}_{t} from the discrepancy between the proprioception history and the commanded action history via an adaptation module ϕ\phi. This idea is inspired by recent work in locomotion where the proprioception history is used to estimate the terrain properties. We show that this information can also be used to estimate the object properties.

To train this network, we first collect trajectories and privileged information by executing the policy at=π(ot,z^t)\smash{\mathbf{a}_{t}=\pi(\mathbf{o}_{t},\hat{\mathbf{z}}_{t})} with the predicted extrinsic vectors z^t=ϕ(qt−k:t,at−k−1:t−1)\smash{\hat{\mathbf{z}}_{t}=\phi(\mathbf{q}_{t-k:t},\mathbf{a}_{t-k-1:t-1})}. Meanwhile we also store the ground-truth extrinsic vector zt\mathbf{z}_{t} and construct a training set

Experimental Setup and Implementation Details

Baselines. We compare our method to the baselines listed below. We also compare with the policy with access to privileged information (Expert), as the upper bound of our method (Figure 2, top row).

A Robust Policy trained with Domain Randomization (DR): This baseline is trained with the same reward function but without privileged information. This gives a policy which is robust, instead of adaptive, to all the shape and physical property variations .

Online Explicit System Identification (SysID): This baseline predicts the exact system parameters et\mathbf{e}_{t} instead of the extrinsic vector zt\mathbf{z}_{t} during training the adaptation module.

No Online Adaptation (NoAdapt): During deployment, the extrinsics vector z^t\hat{\mathbf{z}}_{t} is estimated at the first time step and stays frozen during the rest of the run. This is to study the importance of online adaptation enabled by the adaptation module ϕ\phi.

Action Replay (Periodic): We record a reference trajectory from the expert policy with privileged information and run it blindly. This is to show our policy adapts to different objects and disturbances instead of periodically executing the same action sequence.

Metrics. We use the following metrics to compare the performance of our method to baselines.

Time-to-Fall (TTF). The average length of the episode before the object falls out of the hand. This value is normalized by the maximum episode length (20s in simulation experiments and 30s in the real world experiments).

Rotation Reward (RotR). This is the average rotation reward (ω⋅k^{\bm{\omega}}\cdot\mathbf{\hat{k}}) of an episode in simulation. Note that we do not train with this reward. Instead, we use a clipped version of this reward during training.

Radians Rotated within Episode (Rotations). Since angular velocity of the object is hard to accurately measure in the real world, we instead measure the net rotation of the object (in radians) achieved by the policy with respect to the world’s z-axis. This metric is only used in the real world experiments.

Object’s Linear Velocity (ObjVel). We measure the magnitude of the linear velocity to measure the stability of the object. This is only measured in simulation. The value is scaled by 100.

Results and Analysis

In this section, we compare the performance of our method to several baselines both in simulation and in real-world deployment. We also analyze what the adaptation module learns and how it changes during policy execution and as the objects change. Finally, we focus on training a policy for rotating an object along the negative z-axis with respect to the world coordinate. We also explore the possibility of training a multi-axis policy (±\rm{\pm} zz-axis) in the appendix. More qualitative results of our method and several ablations of our method are in our Project Website and the supplementary.

Comparison in Simulation. We first compare our method with the baselines mentioned in Section 4 in simulation. We evaluate all methods under two settings: 1) In the Within Training Distribution setting, we use the same object set and randomization setting as in RL training; 2) In the Out-of-Distribution setting, we use objects with a larger range of physical randomization range. We also change 20% of the objects to be spheres and cubes. We compute the average performance over 500K episodes with different parameter randomizations and initialization conditions. We report the average and standard deviation over five models trained with different seeds.

The results in Table 1 show that our method with online adaptation achieves the best performance compared to all the baselines. We see that adaptation to shape and dynamics of the object not only enables a better performance in training, but also gives a much better generalization to out-of-distribution object parameters compared to all the baseline methods. The Periodic baseline (i.e., simply playback of the expert policy) does not give a reasonable performance. Although it can rotate the exact same object with the same initial grasp and dynamic parameters, it fails to generalize beyond this very narrow setup. This baseline helps us understand the difficulty of the problem. The NoAdapt also performs poorly compared to our method which uses continuous online adaptation. The weaker performance of this baseline is explained by the fact that it does not update the extrinsics during the episode. This shows the importance of continuous online adaptation. The DR baseline, although roughly matches or performs better than the other baselines in terms of RotR and TTF, it is worse in other metrics related to object stability and energy efficiency. This is because the DR baseline is unaware of the underlying object properties and needs to learn a single gait, instead of an adaptive one, for all possible objects. The SysID baseline also performs worse than our method in both of the evaluation distributions. This is because it is harder as well as unnecessary to learn the exact values of the shape and dynamics parameters to adapt. This comparison shows the benefit of learning a low-dimensional compact representation which has a coarse relative activation for different physical properties instead of the exact value (see Figure 5 and Figure 6).

We perform the same comparisons on a collection of irregular objects (Figure 4). It contains a container with moving COM, objects with concavity, a cylindrical kiwi fruit, a shuttlecock, a toy with holes, and a cube toy. Rotating the cube is particularly difficult because we only rely on usage of fingertips. Although these variations are challenging for this task and are beyond what is seen during training, our method still performs reasonably well, outperforming all the baselines. We see that the DR baseline can perform stable but slow in-hand rotation for the container and the shuttlecock, but largely fails for the other objects, indicating its difficulty in shape generalization. For the SysID baseline, its stability is significantly lower than the DR baseline as well as our method, despite having a higher angular velocity. The NoAdapt baseline performs similarly to what we observed in Figure 3.

2 Understanding and Analysis

To understand how our method generalizes to a diverse set of objects, we design a few experiments and visualization to help develop some insights.

Extrinsics over Time. We run one continuous evaluation episode in the real world in which we replace the object in the hand every 30s for 6 objects. Note that during training, we never randomize the objects within an episode. During the entire run, we record the estimated extrinsics and plot 2 out of the 8-dim extrinsics vector in Figure 5. The top plot shows the extrinsic value zt,0\mathbf{z}_{t,0} which responds to changes in object diameter. It has a lower value for smaller diameters and higher value for larger diameters. The bottom plot shows that the extrinsic value zt,2\mathbf{z}_{t,2} responds to variations in object mass.

Extrinsic Clustering. Another way to understand the estimated extrinsics z^\hat{\mathbf{z}} is by clustering the extrinsics vector estimated while rotating different objects. In Figure 6, we visualize the estimated extrinsics vector using t-SNE for rotating 6 different objects. We find that objects of different sizes and different weights tend to occupy separate regions as explained in Figure 6.

Emergent Finger Gaits. We find it important to use cylindrical objects for the emergence of a stable and high-clearance gait. On our website, we compare the learned finger gaits when training with cylindrical objects and with pure spherical objects. The latter training scheme leads to a policy with a dynamic gait which works well on balls but fail to generalize to more complex objects.

3 Real World Qualitative Results

Discussion and Limitations

There are different levels of difficulty for general dexterous in-hand manipulation. The task considered in this paper (in-hand object rotation over the zz-axis) is a simplification of the general SO(3) reorientation problem. However, it is not a limiting simplification. With three policies for rotation along three principle axes, we can achieve rotating the object to any target pose. We view this task as an important future extension of our work.

We aim to study the generalization to the real world via adaptation, so we do not utilize real-world experience to improve our policy. Incorporating real-world data to improve our policy (e.g. by using meta-learning) will be an interesting and meaningful next step.

This research was supported as a BAIR Open Research Common Project with Meta. In addition, in their academic roles at UC Berkeley, Haozhi, Ashish, and Jitendra were supported in part by DARPA Machine Common Sense (MCS) and Haozhi and Yi by ONR (N00014-20-1-2002 and N00014-22-1-2102). We thank Tingfan Wu and Mike Lambeta for their generous help on the hardware setup, Xinru Yang for her help on recording real-world videos. We also thank Mike Lambeta, Yu Sun, Tingfan Wu, Huazhe Xu, and Xinru Yang for providing feedback on earlier versions of this project.

References

Appendix A Additional Results and Analysis

In Figure 7, we show an almost linear correlation between object mass and the average torque commanded by our policy. This indicates our policy behave differently for objects with different weights and commands smaller torque when object is lighter for energy efficiency.

In Section 5.3, we show the average performance over multiple challenging objects. To give more insights about the working and failure cases of our approach, we analyze the performance on each of the 33 objects we used in experiments. The results are shown in Figure 8, Figure 9, and Figure 10.

Our policy can achieve almost perfect stable and continuous in-hand rotation in 22 out of 33 objects. This set includes objects with different scale, mass, coefficient of frictions, and shapes. Notably, for objects with very high center of mass (the plastic bottle and paper towel), the policy can still perform stable and dynamic in-hand rotations. For more challenging objects including cubes, heavy object with larger aspect ratio (Pear and Kiwi Fruit), and small objects, our policy can still perform about 10 to 20 seconds in-hand rotations. The failure cases are mostly object falling from the fingertips because of incorrect contact positions.

A.2 Multi-Axis In-Hand Object Rotation

We explore the possibility of using our approach to do multi-axis in-hand object rotation. We qualitatively demonstrate preliminary results in the project website. We find this is a harder task than single-axis training and needs about 1.5×1.5\times training time. Our policy achieves a smooth gait transition between rotation over different axes.

Appendix B Ablation Experiments

We analyze different design choices for our method in simulation.

Training with Different Randomization Parameters. Our approach applies a large range of physical randomization during training. In Table 2, we compare the performance difference with our method but without physical randomization.

The No Rand entry in Table 2 use the same reward as our method. However, we do not apply any physical randomization during training. The object scale, mass, coefficient of friction, and other physical properties are fixed. This method performs much worse than our method with physical randomization, especially in the out-of-distribution case.

Length of Proprioceptive History. In Table 3, we compare and report the performance with different proprioceptive history length during training the adaptation module. We report the performance a history length 10, 20, and 30. We observe the performance keeps increasing when using a longer history, but saturates at about 30. Considering the efficiency during inference time, we decide to use T=30 in final experiments.

Additional Baseline for Robust Domain Randomization (DR). In the main text, the Robust Domain Randomization (DR) baseline only receives ot\mathbf{o}_{t} as the input, which contains the robot joint positions and actions in the past three timesteps. In Table 4, we study the effect of including longer robot observations.

To fuse robot observations from multiple steps, we consider two different architectures. In MLP, we flatten and concatenate the observations over the time dimension and directly feed it into the policy network. We also explore to use an LSTM network to capture the temporal correlation of input observations. The input will first pass through an LSTM network and then feed into the policy network. The results are shown in Table 4.

We find that the MLP network is hard to optimize via PPO when temporal length is higher. We see a clear trend of decreasing performance when we increase the temporal length from 3 to 30. LSTM performs better than MLP when the temporal length is larger but still suffer the optimization difficulty.

We also consider a temporal convolution network but find it cannot be successfully optimize and only produce random policies.

Appendix C Implementation Details

Physical Randomization Parameter. We apply domain randomization during training the base policy as well as the adaptation module. The parameters are listed in Table 5 (Train Range). During simulation evaluation, we use a larger randomization range to test the performance of out-of-distribution generalization.

At lease two fingers are in contact with the object.

In practise, we discretized (each region is separated by 0.2) the scales specified in Table 5 and pre-sampled 50,00050,000 grasping poses for the robot hand and the object for each scale.

rrot≐max⁡(min⁡(ω⋅k^,rmax⁡),rmin⁡)r_{\rm rot}\doteq\max(\min({\bm{\omega}}\cdot\hat{\mathbf{k}},r_{\max}),r_{\min}) is the rotation reward. k^\hat{\mathbf{k}} is a desired rotation axis in the world coordinate. In the experiment, we use k^=⊤\hat{\mathbf{k}}=^{\top}. We use rmax⁡=0.5r_{\max}=0.5 and rmin⁡=−0.5r_{\min}=-0.5 to clip the rotation reward.

rpose≐−∥q−qinit∥22r_{\rm pose}\doteq-\left\lVert\mathbf{q}-\mathbf{q}_{\rm init}\right\rVert_{2}^{2} is the hand pose penalty. λpose=−0.3\lambda_{\rm pose}=-0.3.

rtorque≐−∥τ∥22r_{\rm torque}\doteq-\left\lVert{\bm{\tau}}\right\rVert_{2}^{2} is the torque penalty. λtorque=−0.1\lambda_{\rm torque}=-0.1.

rwork≐−τ⊤q˙r_{\rm work}\doteq-{\bm{\tau}^{\top}\bm{\dot{q}}} is the energy consumption penalty. λwork=−2.0\lambda_{\rm work}=-2.0.

rlinvel≐−∥v∥22r_{\rm linvel}\doteq-\left\lVert{\bm{v}}\right\rVert_{2}^{2} is a object linear velocity penalty. λlinvel=−0.3\lambda_{\rm linvel}=-0.3.

Our final reward function is then defined as