Holo-Dex: Teaching Dexterity with Immersive Mixed Reality

Sridhar Pandian Arunachalam, Irmak Güzey, Soumith Chintala, Lerrel Pinto

I Introduction

Learning-based methods have had a transformational effect in robotics on a wide range of domains from manipulation , locomotion , and aerial robotics . Such methods often produce policies that input raw sensory observations and output robot actions. This circumvents challenges in developing state-estimation modules, modeling object properties and tuning controller gains, which requires significant domain expertise. Even with the steep progress in robot learning, we are still long way off from dexterous robots that can solve arbitrary robot tasks akin to methods in game play , text generation or few-shot vision .

To understand what might be missing in robot learning, we need to ask a central question: How do we collect training data for our robots? One option is to collect data on the robot through self-supervised data collection strategies. While this results in robust behaviors , they often require extensive real-world interactions in the order of thousands of hours even for relatively simple manipulation tasks . An alternate option is to train on simulated data and then transfer to the real robot (Sim2Real). This allows for learning complex robotic behaviors multiple orders of magnitude faster than on-robot learning . However, setting up simulated robot environments and specifying simulator parameters often requires extensive domain expertise .

A third, more practical option to collect data is by asking human teachers to provide demonstrations . Robots can then be trained to quickly imitate the demonstrated data. Such imitation methods have recently shown promise in a variety of challenging dexterous manipulation problems . However, there lies a fundamental limitation in most of these works – collecting high-quality demonstration data for dexterous robots is hard! They either require expensive gloves , extensive calibration , or suffer from monocular occlusions .

In this work, we present Holo-Dex, a new framework to collect demonstration data and train dexterous robots. It uses VR headsets (e.g. Quest 2) to put human teachers in an immersive virtual world. In this virtual world, the teacher can view a robotic scene from the eyes of a robot, and control it using their hands through inbuilt pose detectors. Holo-Dex allows humans to seamlessly provide robots with high-quality demonstration data through a low-latency observational feedback system. Holo-Dex offers three benefits: (a) Compared to self-supervised data collection methods, it allows for rapid training without reward specification as it is built on powerful imitation learning techniques; (b) Compared to Sim2Real approaches, our learned policies are directly executable on real robots since they are trained on real data; (c) Compared to other imitation approaches, it significantly reduces the need for domain expertise since even untrained humans can operate VR devices.

We experimentally evaluate Holo-Dex on six dexterous manipulation tasks that require performing complex, contact-rich behavior. These tasks range from in-hand object manipulation to single-handed bottle opening. Across our tasks, we find that a teacher can provide demonstrations at an average of 60s60s per demonstration using Holo-Dex, which is 1.8×1.8\times faster than prior work in single-image teleoperation . On 4/64/6 tasks, Holo-Dex can learn policies that achieve >90%>90\% success rates. Surprisingly, we find that the dexterous policies learned through Holo-Dex can generalize on new, previously unseen objects.

In summary, this work presents Holo-Dex, a new framework for dexterous imitation learning with the following contributions. First, we demonstrate that high-quality teleoperation can be achieved by immersing human teachers in mixed reality through inexpensive VR headsets. Second, we experimentally show that the demonstrations collected by Holo-Dex can be used to train effective, and general-purpose dexterous manipulation behaviors. Third, we analyze and ablate Holo-Dex over various decisions such as the choice of hand tracker and imitation learning methods. Finally, we will release the mixed reality API, demonstrations collected, and training code associated with Holo-Dex on https://holo-dex.github.io/.

II Related Work

Our framework builds upon several important works in robot learning, imitation learning, teleoperation and dexterous manipulation. In this section, we briefly describe prior research that is most relevant to ours.

There are several approaches one can take to teach robots. Reinforcement Learning (RL) can train policies to maximize rewards while collecting data in an automated manner. This process often requires a roboticist to specify the reward function along with ensuring safety during self-supervised data collection . Furthermore, such approaches are often sample-inefficient and might require extensive simulation training for optimizing complex skills.

Simulation to Real (Sim2Real) approaches focus on training RL policies in simulation, followed by transferring to the real robot . Such a methodology of robot training has received significant success owing to the improvements in modern robot simulators. Sim2Real still requires significant human involvement as every task needs to be carefully modeled in the simulator. Moreover, even during training special techniques are required to ensure that the resulting policies can transfer to the real robot .

Imitation learning approaches focus on training policies from demonstrations provided by an expert. Behavior Cloning (BC) is an offline technique that trains a policy to imitate the expert behavior in a supervised manner . Recently, non-parametric imitation approaches have shown promise in learning from fewer demonstrations . Another set of imitation learning is Inverse Reinforcement Learning (IRL) . Here, a reward function is inferred from demonstrations, followed by using RL to optimize the inferred reward. While Holo-Dex is geared towards offline imitation, the demonstrations we collect are compatible with IRL approaches as well.

II-B Dexterous Teleoperation Frameworks

To effectively use imitation learning for dexterous manipulation we need to obtain accurate hand poses from a human teacher. There are several approaches to gather demonstrations for dexterous tasks. Using a custom glove to measure a user’s hand movements such as CyberGlove or Shadow Dexterous Glove has been a popular solution. However, although such gloves have high accuracy, they can be expensive and require significant calibration effort. Vision-based hand pose detectors have shown promise for dexterous tasks. Some examples include using multiple RGBD , single depth , RGB , and RGBD images. However, such methods either require custom calibration procedures or suffer from occlusion-related issues when using single cameras . Recently, a new generation of VR headsets has enabled advanced multi-camera hand pose detection that gave promising results in . This enhancement provides a robust solution that is significantly cheaper compared to CyberGlove and requires little calibration. While VR tools have been used to collect demonstrations for low-dimensional end-effector control, Holo-Dex shows that the VR headsets can be used for high-dimensional control in augmented reality. Concurrent to our work, Radosavovic et al. also show that hand tracking from VR can be used to teleoperate robot hands albeit without using mixed reality.

II-C Dexterous Manipulation

Due to its high-dimensional action space, learning complex skills with dexterous multi-fingered robot hand has been a longstanding challenge . Model-based RL and control approaches have demonstrated significant success on tasks such as spinning objects and in-hand manipulation . Similarly, model-free RL approaches have shown that Sim2Real can enable impressive skills such as in-hand cube rotation and Rubik’s cube face turning . However, both learning approaches requires hand-designing reward functions along with system identification or task-specific training procedures . Coupled with long training times, often requiring weeks , they make dexterous manipulation difficult to scale for general tasks.

To address the poor sample efficiency of prior learning-based methods, several works have looked at imitation learning . Here, given a handful of demonstrations, simulated policies can be trained in a few hours. More recently, such imitation-based approaches have shown success on real robot hands . Holo-Dex takes this idea further by improving the teaching process and demonstrating its utility on a variety of in-hand manipulation tasks.

III Background on Visual Imitation Learning

To understand the imitation learning framework used in Holo-Dex, we first formalize and describe important background work in self-supervised learning and non-parametric imitation. Together, these enable efficient imitation learning from high-dimensional visual observations.

Self-Supervised Learning (SSL) focuses on obtaining low-dimensional embeddings zz from high-dimensional observations oo . Operationally, the observations (e.g. RGB images) are fed into an encoder fθf_{\theta}, where θ\theta denotes the weights of a parametric deep network. While there are several methods to train fθf_{\theta}, the central principle in many works is to predict one ‘view’ of the observation given a different ‘view’ of the same observation. One example of such a learning scheme is data-augmented SSL. Here, the observation oo is augmented by applying visual augmentations such as color jitter or random grayscale. Given two augmented views of this observation o1o^{1} and o2o^{2}, the corresponding embeddings would be z1≡fθ(o1)z^{1}\equiv f_{\theta}(o^{1}) and z2≡fθ(o2)z^{2}\equiv f_{\theta}(o^{2}). The training objective for fθf_{\theta} amounts to maximizing the mutual information between the two embeddings I(z1,z2)I(z^{1},z^{2}).

To optimize this objective, we use the BYOL training scheme, which amounts to predicting z2←gϕ(z1)z^{2}\leftarrow g_{\phi}(z^{1}) through a small deep model gϕg_{\phi} called the ‘projector’. This scheme for learning embeddings has had significant success in a variety of domains ranging from computer vision, audio processing, and robotics. Given its simplicity, we use BYOL to obtain concise embeddings from our demonstrated data.

III-B Non-Parametric Imitation Learning

In our framework for imitation learning we have access to expert demonstrations in the form of DE≡{(ot′E,st′E,at′E)}\mathcal{D}^{E}\equiv\{(o^{E}_{t^{\prime}},s^{E}_{t^{\prime}},a^{E}_{t^{\prime}})\}, where ot′Eo^{E}_{t^{\prime}} represents the sensory observation at time t′{t^{\prime}}, st′Es^{E}_{t^{\prime}} represents the robot state, and at′Ea^{E}_{t^{\prime}} denotes the robot action taken. Note that st′Es^{E}_{t^{\prime}} does not contain information about the object that is being manipulated. Hence, object information needs to be inferred from observations ot′Eo^{E}_{t^{\prime}}. Given these demonstrations, we would like to learn a policy π(at∣ot)\pi(a_{t}|o_{t}) that follows the expert behavior DE\mathcal{D}^{E}. While there are several strategies for optimizing π\pi, we resort to non-parametric approaches given their superior performance in low-data regimes .

Our non-parametric control framework follows VINN , where given the observations oEo^{E} from the expert demonstration dataset DE\mathcal{D}^{E}, a BYOL encoder fθf_{\theta} is trained. Next, the observations in the dataset are all converted to embeddings, i.e. {ot′E}→fθ{zt′E}\{o^{E}_{t^{\prime}}\}\xrightarrow[]{f_{\theta}}\{z^{E}_{t^{\prime}}\}. During run time, when the robot receives an observation oto_{t} it is embedded to ztz_{t}. Then the Nearest-Neighbor (NN) example in {(zt′E,st′E,at′E)}\{(z^{E}_{t^{\prime}},s^{E}_{t^{\prime}},a^{E}_{t^{\prime}})\} is selected to be imitated. We denote this NN example as {(zt∗E,st∗E,at∗E)}\{(z^{E}_{t^{*}},s^{E}_{t^{*}},a^{E}_{t^{*}})\}. Given a small dataset DE\mathcal{D}^{E}, which is often the case in robotic applications, this NN-based imitation learning provides effective learning compared to parametric approaches such as BC.

IV Holo-Dex

As seen in Fig. LABEL:fig:intro, Holo-Dex operates in two phases. In the first phase, a human teacher uses a Virtual Reality (VR) headset to provide demonstrations to a robot. This phase consists of creating a virtual world for teaching, estimating hand poses from the teacher, retargeting the teacher’s hand pose to the robot’s hand and finally controlling the robot hand. After a handful of demonstrations are collected in phase one, the second phase of Holo-Dex learns visual policies to solve the demonstrated tasks. In this section, we will describe each sub-component in detail.

We use the Meta Quest 2 VR headset to place human teachers in a virtual world. The headset surrounds the human in a virtual environment at a resolution of 1832×19201832\times 1920 and a refresh rate of 7272 Hz. The base version of this headset is affordable at $399399 and is relatively light at 503503g. These features allow for comfortable operation by the teacher. Importantly, the API interface of the Quest 2 allows for creating custom mixed reality worlds that visualizes the robotic system along with diagnostic panels in VR. Examples of virtual scenes are depicted in Fig. 2 and Fig. 3.

IV-B Hand Pose Estimation with VR Headsets

In contrast to prior work on dexterous teleoperation, using VR headsets provides three benefits with regard to hand pose estimation of the human teacher. First, since the Quest 2 uses 4 monochrome cameras, its hand-pose estimator is significantly more robust compared to single camera estimators . Second, since the cameras are internally calibrated, they do not require specialized calibration routines that are needed in prior multi-camera teleoperation frameworks . Third, since the hand pose estimator is integrated into the device, it can stream real-time poses at 72Hz. As noted in prior work , a significant challenge in dexterous teleoperation is obtaining hand poses at both high accuracy and a high frequency. Holo-Dex significantly simplifies this problem by using commercial-grade VR headsets.

IV-C Human to Robot Hand Pose Retargeting

Once we have extracted the teacher’s hand pose from VR, we will need to retarget it to the robotic hand. This is done by first computing the individual hand joint angles in the teacher’s hand. Given these joint angles, a straightforward method of retargeting is to directly command the robot’s joints to the corresponding angle. In practice, this works well for all fingers except the thumb. The thumb presents a unique challenge to our Allegro robot hand since its morphology does not match a human’s hand. To address this, we map the spatial coordinates of the teacher’s thumb fingertip to the robot’s thumb fingertip. The joint angles of the thumb are then computed through an inverse kinematics solver. Since the Allegro hand does not have a pinky finger, we ignore the teacher’s pinky joints.

The overall pose retargeting procedure does not require any calibration or user-specific tuning to collect demonstrations. However, we find that thumb retargeting can be improved by finding user-specific maps from their thumb to the robot’s thumb. This entire procedure is computationally inexpensive and can stream desired robot hand poses at 60 Hz.

IV-D Robot Hand Control

Our Allegro Hand is controlled asynchronously over a ROS communication framework. Given desired robot joint positions that were computed from the retargeting procedure, we use a PD controller to output desired torques at 300Hz. To reduce steady-state error, we use a gravity compensation module to compute offset torques. On latency tests, we find that when the VR headset is on the same local network as the robot hand, we achieve latency under 100 milliseconds. Having a low error and latency is crucial for Holo-Dex since it allows for intuitive teleoperation of the robot hand by the human teacher.

As the human teacher controls the robot hand, they can see the robot change in real time (60Hz). This allows the teacher to correct execution errors in the robot. During the teaching process, we record observational data from three RGBD cameras and the action information of the robot at 5Hz. We had to reduce the recording frequency due to the large data footprint and associated bandwidth required from recording multiple cameras.

IV-E Imitation Learning with Holo-Dex Data

Once data is collected in the first phase of Holo-Dex, we now proceed to the second phase, where visual policies are trained on top of this data. We employ the Imitation with Nearest Neighbors (INN) algorithm for learning. Background details of INN is present in Section III-B. In prior work, INN was shown to produce state-based dexterous policies on the Allegro hand . Holo-Dex takes this a few steps further and demonstrates that these visual policies can generalize to novel objects in a variety of dexterous manipulation tasks.

To select the learning algorithm for obtaining low-dimensional embeddings (see Section III-A), we experiment with several state-of-the-art self-supervised learning algorithms and find that BYOL provides the best nearest neighbour results. Hence we select BYOL as our base self-supervised learning method. Once BYOL is trained on the collected demonstrations, we can run dexterous manipulation policies on the robot by performing nearest neighbor action retrieval (see Section III-B) to get the closest example in the trainset (zt∗E,at∗E)(z^{E}_{t^{*}},a^{E}_{t^{*}}). To account for the slower rate of demonstration collection, we set the action to the difference between the succeeding state to the closest neighbor and the current state, i.e. at=(st∗+kE)−(stE)a_{t}=\mathcal{(}s^{E}_{t^{*}+k})-(s^{E}_{t}). Here kk represents the number of states skipped during our recording of demonstrations. Note that directly commanding at∗Ea^{E}_{t^{*}} would fail due to our asynchronous data storage framework.

V Experimental Evaluation

Our experiments and tasks are designed to answer the following questions:

How long does it take Holo-Dex to collect demonstrations?

How successful are policies trained by Holo-Dex?

How general are the skills learned by Holo-Dex?

How many demonstrations from Holo-Dex are required to successfully solve dexterous tasks?

We study six dexterous manipulation tasks that require contact-rich, multi-fingered control for successful completion. Details of these tasks are described below.

Planar Rotation: Given an object placed at a random position on the palm of the robot hand, the goal is to rotate the object in the counter-clockwise direction along the palm normal vector. Solving this task requires the robot to make multi-fingered contacts to both rotate and correct for deviations of the object from the center of the hand. The task is considered a success if the robot is able to rotate the object by 90∘90^{\circ} under a minute.

Object Flipping: Given an object placed at a random position on the palm of the robot, the goal is to flip the object in the hand’s direction. Solving this task requires the robot to make multi-fingered contacts for grasping the top face of the object and to correct the object when it deviates from the center of the hand. The task is considered a success if the robot is able to flip the object by 90∘90^{\circ} within a minute.

Can Spinning: Given a large can placed horizontally on the palm of the robot hand, the task is to spin the can in the counter-clockwise direction along the palm normal vector. Solving this task requires the robot to make synchronized multi-fingered contacts to apply controlled torques on the curved sides of the can. The task is considered a success if the robot is able to spin the can by 90∘90^{\circ} under 30 seconds.

Bottle Opening: Given a bottle on a desk, the task is to single-handedly grab the bottle and turn its lid open using the index finger. Solving this task requires the robot to use the index finger to grip the bottle cap and turn it while using the non-index fingers to hold the bottle firmly. The task is considered a success when the robot rotates the bottle cap by 360∘360^{\circ} under 180 seconds.

Card Sliding: Given a card on a desk in front of the hand, the task is to grab the card by sliding it off the table and picking it up. To solve the task the robot requires to use the thumb finger to slide the card to the edge of the desk and the other non-thumb fingers to grab the card from the edge. The task is considered a success when the robot is able to lift the card off and stably grasp it from the table under 120 seconds.

PostIt Note Sliding: This task is similar to card sliding, but instead of a card we use a thicker post-it note pad as the object to pick from the desk. To solve the task, the robot needs to apply a firmer torque on the post-it note pad using the thumb finger since the object is heavier. The task is considered a success when the robot is able to lift the card off and stably grasp it from the table under 120 seconds.

For the Planar Rotation task we collect 120120 expert demonstrations, while for all the other tasks we collect 3030 demonstrations. Additional demonstrations for Planar Rotation are collected to account for the relative difficulty of this task and experimentation with different dataset sizes.

V-B How long does it take to collect demonstrations?

The closest dexterous teleoperation work to ours that uses commodity sensors is single-image teleoperation (e.g. DIME ), where hand poses are detected via RGB images to get robot joint angles. Despite its simplicity, such single-image pose estimation suffers from hand occlusions, which results in poor teleoperation performance on challenging manipulation tasks . In Table I we show that Holo-Dex can collect successful demonstrations 1.8×1.8\times faster compared to DIME. For 3/63/6 tasks that require precise 3D movements, we find that single-image teleoperation is insufficient to collect even a single demonstration.

To demonstrate the versatility of Holo-Dex we ask five untrained users to collect demonstrations for each of our tasks. Unlike prior work , no user-specific calibration was done for this evaluation. In Table I, we see that these users are successfully able to solve 4/64/6 tasks on their first try, failing only on more difficult sliding tasks. We also find that training on this system is quite important as it yields a nearly 2.6×2.6\times speedup in demonstration collection.

V-C How successful are policies trained by Holo-Dex?

We examine the performance of various imitation learning policies on all dexterous tasks. The imitation learning algorithms include Behavior Cloning (BC) , Behavior Cloning from pretrained representations (BC-Rep) and VINN . Table II shows the success rates of each task with different policies. We find that VINN outperforms both Behavior Cloning algorithms on all tasks. This is in line with prior work and showcases the effectiveness of non-parametric imitation with few demonstrations. However, we find that for the two tasks that involve sliding and picking, the performance of VINN is quite low at 30%30\%. We believe this is due to our robot’s inability to sense touch, which limits our vision-only model from performing precise actions.

V-D How general are the policies learned by Holo-Dex?

To understand the generalization capabilities of our models, we analyse the quality of the embeddings we get for a given visual input. Interestingly, we observed that the Planar Rotation policy’s encoder was able to generalize to the other two in-hand manipulation tasks. We reason that this encoder was able to exhibit this behavior since it was trained on abundant Planar Rotation data, where the object used while collecting demonstrations had different colors on it’s each face and was placed on various locations.

Since our policies are vision-based and do not require explicitly estimating the states of objects, they are compatible with objects not seen in training. We evaluate our in-hand manipulation policies which were trained for performing Planar Rotation, Object Flipping, and Can Spinning tasks on 10 visually and geometrically distinct objects each, with 5 rollouts for every object in different initial positions. Results for this experiment are in Table II and are visualized in Fig. 5. Surprisingly, for all three in-hand manipulation tasks, we find high success rates without any additional demonstration collection or training. We observed that the Planar Rotation policy was able to generalize on 7 out of 10 objects, whereas the Object Flipping and Can Spinning policies were able to succeed at performing the task on 10 and 9 unseen objects respectively. We believe that the policy fails to generalize on some objects because of their visual features (object color and shape) being very different from that of the object in the demonstrations. This means that although we collect demonstrations from Holo-Dex on a single object, the learned policies can generalize in a zero-shot manner.

V-E How many demonstrations are needed to solve our tasks?

In Fig. 6 we visualize the performance on four of our tasks across different dataset sizes. To decouple the effects of representation learning with action prediction, we use the same encoder (trained on all task data) for different dataset splits. We find that for Can Spinning and Bottle Opening, a single demonstration is sufficient to achieve high performance, while for Planar Rotation and Object Flipping we see steady gains in performance as we increase the amount of demonstration data.

VI Limitations and Discussion

We have presented Holo-Dex, a framework that takes some of the first steps towards immersive teaching of dexterous robots through VR. There are currently two limitations of this work. First, we find that for the harder manipulation tasks, such as sliding a card our learned policies achieve poor performance. Integrating tactile sensing to Holo-Dex could remedy this issue. Second, our retargeting procedure only applies to robots that can map to human joints. This limits its applicability to robots with different morphologies (e.g. aerial robots, quadrupeds, etc.). Future research on UX design and retargeting mechanisms can enable mapping VR control to more complex end-effectors.

VII Acknowledgements

We thank Ankur Handa, Ilija Radosavovic, Josh Merel, David Brandfonbrener, Ben Evans, Mahi Shaffiulah, Jeff Cui, Jyo Pari and Siddhanth Haldar for feedback and discussions. This work was supported by a Honda award and ONR award N000142112758.

References