General In-Hand Object Rotation with Vision and Touch
Haozhi Qi, Brent Yi, Sudharshan Suresh, Mike Lambeta, Yi Ma, Roberto Calandra, Jitendra Malik
Introduction
Despite recent progress on in-hand manipulation for a single or a few objects , generalizable object manipulation remains a challenge. In this paper, we present a model that integrates visual, tactile, and proprioceptive sensory inputs and achieves fingertip-based in-hand object rotation over multiple different axes. This continuous rotation task is important for achieving large-angle in-hand re-orientation skill and is challenging because it requires simultaneously maintaining stable force closure for objects with diverse geometries.
An overview of our method, RotateIt, is shown in Figure 2. Our approach draws inspiration from recent advances in training reinforcement learning policies with privileged information , more specifically rapid motor adaptation . We first train an oracle policy that is conditioned on a representation of the privileged information (called extrinsics, denoted as , as shown in Figure 2), which contains ground-truth physical properties and shapes of the objects. With access to this representation, the oracle policy is able to efficiently and stably manipulate diverse objects over multiple axes in the simulator.
The key challenge for real-world deployment lies in estimating the extrinsics encoding when privileged information is inaccessible. To address this challenge, we use multimodal sensing from vision and touch, just as humans do . We implement this by designing a visuotactile transformer which operates on a history of multimodal proprioceptive, visual, and tactile inputs to infer . Concretely, during training, we rollout the oracle policy in simulation and collect the foreground object depth, contact locations on the fingertips, proprioception, and action history. Then we feed these multimodal streams into a transformer to produce an estimate of the , denoted as . The visuotactile transformer is trained to minimize the difference between the predicted and estimated encodings of the privileged information. In the real-world, we get the foreground objects using Segment Anything , which enables RotateIt to be robust to cluttered backgrounds. We use tactile images from an omnidirectional vision-based touch sensor to retrieve these contact locations.
We demonstrate RotateIt can perform multi-axis object rotation using only its fingertips. In simulation, we quantitatively study the performance of rotating skills over three principal axes in the hand-centric frame and the impact of incorporating vision and touch at various stages (Section 5.1 and Section 5.2). To further understand what is learned by the policy, we investigate how accurately the latent representation of the policy captures objects’ shapes by using it to recover 3D shapes (Section 5.3). Finally, we deploy the learned policies to rotate multiple different objects over multiple axes in the real world (Section 5.4). On the website, we show our policy can rotate objects, including but not limited to, the three canonical axes. Our work highlights the importance of both visual and tactile sensing in manipulation and presenting a step towards general dexterous in-hand manipulation.
Related Work
Classic Control Methods. Classical methods typically minimize a desired cost function using planning methods and a simplified system model . State-of-the-art systems in this category include full reorientation using a compliance-enabled hand and an accurate pose tracker . In contrast, our work does not rely on an object model: we instead combine multi-sensory inputs with a learning-based policy that is trained on a large set of objects.
Real-World Learning. In-hand manipulation skills can be learned directly in the real-world, either by reinforcement learning or imitation learning . However, reinforcement learning methods usually suffer from sample efficiency and real-world environment cannot provide enough variation. In contrast, our policy is trained using reinforcement learning via GPU-accelerated simulators, and does not need any human demonstrations.
Sim-to-Real Methods. OpenAI et al. first transferred dexterous in-hand manipulation policies to the real-world. Similarly, Sievers et al. , Pitz et al. uses a torque-controlled hand for cube rotation and reorientation when the hand facing downwards. However, they focus on manipulating one single object. Although generalizable in-hand manipulation of diverse objects can be learned in simulation , transferring it to the real world remains a challenge.
Among sim-to-real methods, learning with privileged information is shown to be effective for legged locomotion and manipulation . Recently, several works study generalizable in-hand object rotation using sim-to-real and reinforcement learning. In this paper, we use the same robot hand (Allegro ) as Qi et al. and a similar variety of rotated objects with the significant advance being that the use of visual and tactile information now enables to rotate the object about an arbitrary axis, not just the -axis. Other previous work shared this limitation of only demonstrating rotation about the -axis. Compared to , our task is more challenging as it does not utilize a supporting surface, which allows constant tactile feedback on fingertips and enables a natural finger-gaiting to emerge. While we study continuous rotation (as many revolutions as possible), Chen et al. study the task of object reorientation to an arbitrary pose (no limitation to -axis rotation) and obtain impressive results. There are significant differences to our approach in oracle policy training, hand hardware, and goal specification. In addition to proprioception, they only use vision to sense the object, while our visuotactile policy utilizes both vision and touch. In the experiment section, we show that both these components improve the performance of our system.
Visuotactile Sensing and Learning. Tactile sensors such as GelSight , TacTip , DIGIT , DTact , GelTip , ArrayBot , and AllSight have been used for numerous applications including grasping , playing the piano , 3D reconstruction and localization , and cup unstacking and bottle opening . Previous work explores the usage of vision and touch for manipulation but not for in-hand manipulation. To the best of our knowledge, RotateIt is the first work that intersects visuotactile sensing and learning to achieve general in-hand object rotation with a dexterous hand.
Transformers in Robotics. The transformer architecture was originally proposed for machine translation and later used in computer vision . In robotics, there are growing attempts to incorporate it in imitation learning or reinforcement learning . Chen et al. also use the transformer on multimodal data but their tactile refers to force/torque sensing on robot joints. In contrast, our method uses the transformer for temporal modeling of multimodal proprioceptive, visual, and tactile information.
General In-hand Object Rotation with Vision and Touch
An overview of our method is shown in Figure 2. Our policy training consists of two stages: First, we train an oracle policy with privileged information. Next, we train a visuotactile policy with realistic yet noisy observations. Both of these stages happen in simulation. In this paper, we consider privileged information to be the object physical properties and object shape information. Real-world observations comprise of a stream of proprioceptive, visual, and tactile inputs. Our method trains one policy for each rotation axis and we show how to distill them into one general policy in Section 5.5.
Privileged Information. For the object shape information, we sample points from the object’s mesh and encode it to a feature vector with dimensions using PointNet . One key difference from previous works is that we explicitly encode object shape into the oracle policy, which we find to be critical especially for complex objects that are harder to manipulate.
The physics property contains object’s mass, center of mass, coefficient of friction, scale, and restitution, resulting in a 7-dimensional vector. The pose contains object’s position, orientation (as a quaternion), and angular velocity, resulting a 10-dimensional vector. These vectors are concatenated together and projected to an 8-dim encoding vector . Our final privileged encoding is concatenated from the shape encoding and physical property encoding .
Reward Function. Our reward function is modified from with an additional penalty on undesired angular velocities component:
The object rotation task is defined as where is the object’s angular velocity and is the desired rotation axis in the hand-centric axis. Naively applying this reward will result in unstable behaviors when rotating over and -axis. To alleviate this problem, we add a rotation penalty term . To make the policy stable, smooth, and energy efficient , we use a few penalty terms: is the hand pose deviation penalty, is the torque penalty, is the energy consumption penalty, and is the object linear velocity penalty, where be the starting robot configuration, be the commanded torques at each timestep, and is the object’s linear velocity.
Policy Optimization. We use PPO to optimize the oracle policy. The weights between the policy and the critic network are shared, with an extra linear projection layer to estimate the value function. During training, each environment is assigned to an object with randomized physical properties and a stable initial grasp. We curate a list of hundreds of objects for training as shown in Figure 3.
2 Visuotactile Policy Training with Transformers.
We find robust and adaptive finger-gaiting emerges from the oracle policy training. However, it is assumed to know full object physical properties, pose, and shape as the input. To deploy it in the real-world, we need to use real-world observations to infer (representations of) these properties. Qi et al. uses proprioceptive states to estimate such information. In this work, we augment it to include vision and touch and study their important roles in improving manipulation performance.
Touch (Figure 4). To reduce the sim-to-real gap for tactile sensors, we choose to use the discretized contact location projected on D plane as the proxy of tactile information. In simulation, we directly parse the contact position provided by the simulator, project it onto a D plane in fingertip frame, and discretize it to locations. Specifically, the touch observation is a dimensional array, where is the number of contact at each timestep. For each contact, it contains the discretized contact location (8-dimension) and the index of the finger. During training, since the number of contact points across timesteps are not the same, we use an MLP to each contact information and take an average of different contact point features. In the real-world, we use four omnidirectional vision-based touch sensors at the fingertips. We track the deformation of the highest intensity pixel on each sensor, which serves as a proxy for contact position (Figure 4). This 2D keypoint from vision-based touch, similar in spirit to Sodhi et al. , is directly fed into the policy.
Vision (Figure 5). We use object depth as the vision representation since 1) it is a general representation and does not require human labeling in the real-world and 2) it is hard to realistically simulate RGB images whereas depth is a good abstraction of object shape . In real-world deployment, instead of using the raw depth from a RGBD camera, we use Segment-Anything to segment out the objects to reduce the sim-to-real gap. Formally, given an object depth image , we encode it 3-layer ConvNet to output . An overview of the vision pipeline is shown in Figure 5. We also randomize the camera position and orientation during training, to make the policy robust to minor viewpoint changes.
Visuotactile Transformer. The goal of our visuotactile policy is to accurately infer the learned representation of privileged information. To tackle these challenges, we use a transformer architecture to model these multimodal sensory stream. We concatenate the encoded depth image , encoded tactile contact points , joint positions , and action at the previous timestep to form the feature vector . We feed a sequence of features as input to the transformer. The transformer outputs as the predicted extrinsic vector.
Evaluation Setup
Simulation Setup. We use the IsaacGym simulator. Each environment contains a simulated AllegroHand and a sampled object from our curated object datasets (Figure 3). Each object is of different physical properties (the exact parameters are in the supplementary material) and a random initial pose. For depth and viewpoint consistency between the real and simulated cameras, we measure the camera-robot extrinsics with an ArUco tag placed on the palm of the real-world Allegro. In IsaacGym, we use this transformation augmented with random pose noise, and further apply realistic depth noise on the resultant images .
Object Set. We create a curated dataset for objects used in our experiments from EGAD , Google Scanned Objects , YCB , and ContactDB . We select objects with width/depth/height (w/d/h) aspect ratio less than 2.0 (see Figure 3 for a visualization).
Evaluation Metric. We use the metrics defined in to evaluate our method both in simulation and in the real-world. In addition, we also evaluate undesired rotation penalties in simulation. We find this metric is particularly important for rotation over and axis.
Rotation Reward (RotR). This is the average rotation reward of an episode in simulation.
Rotation Penalty (RotP). This is the average rotation penalty per timestep .
Radians Rotated (Rotations). The rotation (in radians) achieved by the policy with respect to the desired axis. This metric is only used in the real world experiments.
Results and Analysis
In this section, we first quantitatively study our method in simulation. In particular, we study the importance of using object shape information for policy training (Section 5.1), as well as the importance of vision and touch in the visuotactile policies (Section 5.2). Then, we use a shape prediction task to study the information recovered by estimated extrinsic vectors. We show our visuotactile policy learns object shape representation by predicting the 3D shape of objects using (Section 5.3). We also evaluate our method on a real-world robot (Section 5.4) and finally show how to train a single policy to rotate over six principle axes.
The performance is shown in Table 1. We compare RotateIt with previous work and our method without the usage of point cloud while still using the quaternion. Experiments show that using point-cloud significantly improves the performance on all of the metrics and for all rotation axis.
To get more insights, we further plot the relative improvements on varies objects shape for -axis rotation, shown in Figure 7 (the “stage1” row). We find that point-cloud gives the largest improvement on objects with non-uniform w/d/h (width/depth/height) ratios and objects with irregular shapes such as the bunny and light bulb. The improvements on regular objects are smaller but still over 40%. In addition, we also evaluate the oracle policy on 15 held-out challenging objects (Figure 8 (b)). We show that not using point cloud results in a 22% decrease in generalization gap while using point-cloud can improve it to only 8% drop.
Point-cloud as an input is also used in Qin et al. and Bao et al. but they do not explore how to use it for in-hand manipulation. Note that our design is different from Chen et al. , which uses only pose for the oracle policy and uses object shape information only in the student policy. In our setting using object pose is not sufficient to achieve good enough performance.
2 Visuotactile Transformer
The oracle policy evaluated in Section 5.1 cannot be transferred to the real-world because it needs access to a manipulated object shape and physical properties. We instead learn to infer this representation during execution from proprioceptive, visual, and tactile history. In Figure 6, we show that using either vision or touch alone gives a significant performance improvements compared to proprioceptive inputs. We also find using a combination of vision and touch sensing can further improve the performance. By integrating visuotactile sensing and temporal transformer, our method can match the performance of the oracle policy. In appendix Table 4, we also show transformer has better sequence modeling ability compared to temporal convolutions used in previous work .
Similar to what we find in the oracle policy training, we observe the visuotactile policy has larger improvements on irregular and non-uniform objects (Figure 7, “stage2” row). In Figure 8 (c), we show the visuotactile information are critical for OOD generalization. Using proprioception only will lead to a 41% performance drop while using vision and touch can improve it to 15% drop.
Importance of Finer Tactile Sensing. In contrast to prior work , we find in Table 2 that binary contact does not provide benefits. In contrast, contact locations are vital for improving performance in RotateIt. We speculate that this discrepancy is because Khandate et al. does not use proprioceptive and action history.
3 Representation Learned in the Latent Space
Next, we study the information that is encoded into and . After we finish training the four policies in Figure 9, we freeze the network and then we run our policy on 20 objects in our object dataset (16 for training, 4 for testing). This gives us an extrinsic vector dataset for each policy. On each of the datasets, we train one decoder whose input is a sub-sequence of extrinsic vectors and output is the voxel grid. After training this decoder, we run it on the 4 held-out testing objects.
In Figure 9, we visualize predicted shapes averaged over 100 randomly selected subsequences from rollouts on novel test objects for four policies: the stage 1 oracle policy with and without shape (mesh) conditioning, and the stage 2 policy with and without visuotactile sensory inputs. The results suggest that shape information is preserved and useful for our oracle policy even though the only learning signal is the reward function. We also find policies without object shape will consider all the objects as spherical or cuboid objects, which explains the huge improvement on objects with large w/d/h ratios Figure 7. Next, our results also highlight both the capabilities and limits of proprioception, which we see can robustly distinguish between spherical (beige) and cuboidal (green) objects. Shape understanding for more irregular objects like the pear (blue), however, requires additional sensors. This supports the increased benefit of vision and touch for more complex objects that we observe in Section 5.1.
4 Real-world Evaluations
Finally, we quantitatively compare RotateIt and Hora in the real-world on rotating different objects over the -axis. We find that without vision and touch, Hora cannot finish this task. It only learns in-grasp movement with thumb slowly moving to the bottom of the object. It is also not able to maintain stability; the object quickly falls down. In contrast, RotateIt can successfully manipulate multiple objects with different geometries such as cubes, spheres, or cylinders by 2 radians within 20 seconds. Note that many real-world objects are outside our training set such as the box, Cocoon, Squishy, and Stego. The real-world physics is also different from the simulated physics. Having a successful sim-to-real transfer is a strong evidence of generalization. We show qualitative results on rotation around and beyond the three canonical axes on our website. In the video, we also test a policy trained with the assumption that one of the touch sensors is off. The policy performs similarly to the full policy, demonstrating the robustness of the algorithm.
5 Multi-axis Training
In previous sections, each oracle policy is trained with a fixed rotation axis . In this section, we demonstrate it is also feasible to train a single network to perform multi-axis object rotation. To achieve this, we augment the observation space with and train it with the reward defined in Section 3.1 and the imitation learning objective with the corresponding single-axis oracles.
We show the episode rotation reward for both the single-axis oracle policy and the multi-axis policy in Table 3. We empirically find the distilled multi-axis policy performs on par with the single task oracles. We also observe the policy does not converge when training with only reinforcement learning.
Limitations and Future Work
In this paper, we show the feasibility of training policies that can rotate many objects over multiple axes. We view this capability as an important step towards general-purpose in-hand manipulation.
We assume the objects are not too long (e.g. a pencil or a screwdriver) and are within the mechanical limit of the robot hand. Our method is not able to utilize real-world experiences during deployment since it is frozen after training. Lifelong learning in the real-world with cross-modal supervision is a valuable future direction. There are also various ways to improve the touch processing system since we only use the low-dimensional contact location as the input and do not utilize the full information output by the omnidirectional image-based tactile sensor. In addition, we can also improve our vision system by techniques such as visual pre-training.
This research was supported as a BAIR Open Research Common Project with Meta. In their academic roles at UC Berkeley, Haozhi Qi and Jitendra Malik are supported in part by DARPA Machine Common Sense (MCS), Brent Yi is supported by the NSF Graduate Research Fellowship Program under Grant DGE 2146752, and Haozhi Qi, Brent Yi, and Yi Ma are partially supported by ONR N00014-22-1-2102 and the InnoHK HKCRC grant. Roberto Calandra is funded by the German Research Foundation (DFG, Deutsche Forschungsgemeinschaft) as part of Germany’s Excellence Strategy – EXC 2050/1 – Project ID 390696704 – Cluster of Excellence “Centre for Tactile Internet with Human-in-the-Loop” (CeTI) of Technische Universität Dresden. We thank Shubham Goel, Eric Wallace, and Angjoo Kanazawa, Raunaq Bhirangi for their feedback. We thank Austin Wang and Tingfan Wu for their help on hardware. We thank Xinru Yang for her help on real-world videos.
References
Appendix A Additional Experiments
Detailed Comparison of Using Vision and Touch. Table 4 shows the detailed comparison of using vision, touch, and the transformer architecture. We show each of the component can significantly improve over the baseline and are also complement with each other.
Randomization of Simulated Vision Sensing. During training, we apply various randomizations to the vision sensing to make it robust. We evaluate our model under different noise setting in simulation. The results are shown in Table 5.
We add gaussian noise to camera positions and orientations. Cam Pos stands for the value for each different setting (in meters). Cam RPY stands for the extend we randomize the camera rotation (in roll/pitch/yaw values, in radius). The camera field-of-view (fov) is also randomized. The values are set to a uniform distribution according to the Cam FOV column. We also simulate segmentation noise (for each pixel, with probability , the mask is flipped) and segmentation failure (for each timestep, with probability , the mask is completely 0) to simulate segmentation errors in the real-world.
We find that the model behaves robustly under training randomization and slightly out-of-distribution noises. However, too large noise will still impact the performance, highlighting the importance of proper camera calibration.
Appendix B Implementation Details
At lease two fingers are in contact with the object.
In practice, we discretized (each region is separated by 0.2) the scales specified in Table 7 and pre-sampled grasping poses for each object and for each scale.
Reward Hyperparameter. We use , , , , , and . We also find that if we apply at the start of training, the policy will only learn to stably hold the objects. Therefore we set this coefficient to be at the beginning and then linearly decrease it to using curriculum learning .
The visuotactile transformer takes object depth feature, touch feature, proprioception feature, and action history as input. The object depth image is of size and is first be passed to a four layer ConvNet and then a global average pooling layer to produce the feature of dimension . The contact location is a -dimension vector and is first passed to an MLP with hidden unit dimension . The contact feature is aggregated using average pooling. For the robot joint position and actions, we first encode them into a 32-dimensional representations for each timestep via a two-layer MLP (with hidden unit dimension ). The feature dimension of our transformer is and with depth . The self-attention module has parallel head.