DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to Reality

Ankur Handa, Arthur Allshire, Viktor Makoviychuk, Aleksei Petrenko, Ritvik Singh, Jingzhou Liu, Denys Makoviichuk, Karl Van Wyk, Alexander Zhurkevich, Balakumar Sundaralingam, Yashraj Narang, Jean-Francois Lafleche, Dieter Fox, Gavriel State

Introduction

Multi-fingered robotic hands offer an exciting platform to develop and enable human-level dexterity. Not only do they provide kinematic redundancy for stable grasps, but they also enable the repertoire of skills needed to interact with a wide range of day-to-day objects. However, controlling such high-DoF end-effectors has remained challenging. Even in research, most robotic systems today use parallel-jaw grippers.

In 2018, OpenAI et al. showed for the first time that multi-fingered hands with a purely end-to-end deep-RL based approach could endow robots with unprecedented capabilities for challenging contact-rich in-hand manipulation. However, due to the complexity of their training architecture, and the sui generis nature of their work on sim-to-real transfer, reproducing and building upon their success has proven to be a challenge for the community. Recent advancements in in-hand manipulation with RL have made progress with multiple objects and an anthropomorphic hand , but those results have only been in simulation.

While the NLP and computer vision communities have reproduced and extended the successes of large-scale models like GPT-3 and DALL-E respectively, similar efforts have remained elusive in robotics due to hardware and infrastructure challenges. Using large-scale data from simulations may provide avenues to unlock a similar step function in robotics capabilities.

This paper builds on top of the prior work in . We use a comparatively affordable Allegro Hand with a locked wrist and four fingers, using only position encoders on servo motors; the Shadow Hand used in OpenAI’s experiments costs an order of magnitude more than the Allegro Hand. We also develop a simple vision system that requires no specialised tracking or infrastructure on the hand; the system works on three off-the-shelf RGB cameras compared to OpenAI’s expensive marker-based setup, making our system easily accessible for everyone. Furthermore, we use the GPU-based Isaac Gym physics simulator as opposed to the CPU-based MuJoCo , which allows us to reduce the amount of computational resources used and the complexity of the training infrastructure. Our best models required only 8 NVIDIA A40 GPUs to train, as opposed to OpenAI’s use of a CPU cluster composed of 400 servers with 32 CPU-cores each, as well as 32 NVIDIA V100 GPUs (compute requirements for block reorientation). Our more affordable hand, in combination with the simple vision system architecture and accessible compute, dramatically simplifies the process of developing and deploying agile and dexterous manipulation. We summarise our contributions below:

We demonstrate a system for learning-based dexterous in-hand manipulation that uses low-cost hardware (one order of magnitude less expensive than ), uses a purely vision-based pipeline, sets more diverse pose targets, uses orders-of-magnitude cheaper compute, and offers further insights into this problem with detailed ablations.

We develop a highly robust pose estimator trained entirely in simulation which works through heavy occlusions and in a variety of robotic settings e.g. https://www.youtube.com/watch?v=-MTsm0Uh_5o.

While not directly comparable to due to different hardware, our purely vision-based state estimation results not only outperform their best vision-based results, but also fare comparably to their marker-based results.

We will also release both our vision and RL pipelines for reproducibility. We seek to provide a much broader segment of the research community with access to a novel state-of-the-art in-hand manipulation system in hopes of catalyzing further studies and advances.

Method

We propose a method for performing object reorientation on an anthropomorphic hand. Initially the object to be manipulated is placed on the palm of the hand and a random target orientation is sampled in SO(3)SO(3)In contrast to previous works , which limited it to configurations with flat faces pointing upwards.. The policy then orchestrates the motion of the fingers so as to bring the object to its desired target orientation. Similar to OpenAI et al. , if the object orientation is within a specified threshold of 0.4 radians of the target orientation, we sample a new target orientation. The fingers continue from the current configuration and aim to move the object to its new target orientation. The success criterion is the number of consecutive target orientations achieved without dropping the object or having the object stuck in the same configuration for more than 80 seconds. Importantly, each consecutive success becomes increasingly harder to achieve as the fingers have to keep the object in the hand without dropping, hence testing the policy’s ability to model the dynamics on the go.

For a quick and high level understanding of this work, we encourage readers to watch the video https://www.youtube.com/watch?v=TAUiaYAVkfI. Figure 2 provides a quick overview of different components involved in the system.

2 Hardware

Our hardware setup (see Fig 1) consists of an Allegro Hand rigidly mounted at the wrist. We use 3 Intel D415 cameras for object tracking with RGB frames i.e. no depth images were used. The cameras are extrinsically calibrated relative to the palm link of the hand. Our object tracking is done entirely using a vision-based system, and in contrast to , we do not use any marker-based system to track the cube or fingertip states.

The object we learn to manipulate is a 6.5 cm6.5~{}\text{cm} cuboid with coloured and lettered stickers on the sides. These stickers allow the vision system to distinguish between different faces (see Sec. 2.7). The pose of the cube is represented with respect to the palm of the robot hand. The camera-camera extrinsics and camera-robot calibration allow us to transform the cube pose from the canonical reference frame of a camera to the palm. Since the cube is represented locally in the palm reference frame, the policy performance is not dependent on the physical location of the setup, enabling us to move the setup freely whenever desired.

3 Policy Learning with RL

RL Formulation: The task of manipulating the cube to the desired orientation is modelled as a sequential decision making problem where the agent interacts with the environment in order to maximise the sum of discounted rewards. In our case, we formulate it as a discrete-time, partially observable Markov Decision Process (POMDP). We use Proximal Policy Optimisation (PPO) to learn a parametric stochastic policy πθ\pi_{\theta} (actor), mapping from observations o∈Oo\in\mathcal{O} to actions a∈Aa\in\mathcal{A}. PPO additionally learns a function Vϕπ(s,o)V^{\pi}_{\phi}(s,o) (critic) to approximate the on-policy value function. Following Pinto et al. , the critic does not take in the same observations as the actor, but receives additional observations including states s∈Ss\in\mathcal{S} in the POMDP. The actor and critic observations are detailed in Table 1.

We use a high-performance PPO implementation from rl-games with the following hyper-parameters: discount factor γ\gamma=0.998 We found that following and setting the higher discount of 0.998 (as opposed to 0.99 as used with MLPs ) was essential to allowing us to train LSTMs., clipping ϵ\epsilon=0.2. While in some experiments the learning rate was updated adaptively based on a fixed KL threshold 0.016, our best result was obtained using linear scheduling of the learning rate for the policy (start value lr=1e−4lr=1e{-4}) and a fixed learning rate for the value function (lr=5e−5lr=5e{-5}). Our best policy πθ:O×H→A\pi_{\theta}:\mathcal{O}\times\mathcal{H}\to\mathcal{A} was a Long Short-Term Memory (LSTM) network taking in environment observations oo and previous hidden state h∈Hh\in\mathcal{H}. We use an LSTM backpropagation through time (BPTT) truncation length of 16. The LSTM has 1024 hidden units with layer normalization and is followed by 2 multilayer perceptron (MLP) layers with sizes 512 and ELU activation . The action space A\mathcal{A} of our policy is the PD controller target for each of the 16 joints on the robot hand. The value function LSTM layer has 2048 hidden units with layer normalization, followed by 2 MLP layers with 1024 and 512 units respectively with ELU activation. The output of the policy is low-pass filtered with an exponential moving average (EMA) smoothing factor. During training this factor is annealed from 0.2 to 0.15. Our best results in the real world were obtained with an EMA of 0.1, which provided a balance between agility and stability of the motion, preventing the hardware from breaking or motor cables from burning.

4 Reward Formulation

The reward formulation is inspired by the Shadow hand environment in Isaac Gym, and described and justified in Table 2.

5 Simulation

Our aim in this paper is to learn dexterous manipulation behaviours. Current on-policy learning algorithms can struggle to accomplish this on real robots due to the number of samples required. Hence, we learn our behaviours entirely in simulation. We use the GPU-based Isaac Gym physics simulator , which models contacts differently than MuJoCo’s soft-contact model used in . Isaac Gym gives the advantage of being able to simulate thousands of robots in parallel on a single GPU, mitigating the need for large amounts of CPU resources.

6 Domain Randomisation

It is widely known that there is a "sim-to-real" gap between physics simulators and real life . Compounding this is the fact that the robot as a system can change from day to day (e.g., due to wear-and-tear) and even from timestep to timestep (e.g., stochastic noise). To help overcome this, we introduce various kinds of randomisations into the simulated environment as listed in Table 3.

Vectorised Automatic Domain Randomisation: In our best policies, we set the parameters of the domain randomisations via a vectorised implementation of Automatic Domain Randomisation (ADR, introduced in ). ADR automatically adjusts the range of domain randomisations to keep them as wide as possible while keeping the policy performance above a certain threshold. This allows us to train policies with less randomisation earlier in training (enabling behaviour exploration) while producing final policies that are robust to the largest range of environment conditions possible at the end of training, with the aim of improving sim-to-real transfer by learning policies which are robust and adaptive to a range of environment randomisation parameters. Using the parallelisation provided by Isaac Gym, we implement a vectorised variant of the algorithm, which we call Vectorised Automatic Domain Randomisation (VADR, see Algorithm 1).

The range of randomisations for each environment parameter in ADR is modelled as a uniform distribution dn∼U(p2n,p2n+1)d^{n}\sim U(p^{2n},p^{2n+1}), where p2np^{2n} and p2n+1p^{2n+1} are the current lower and upper randomisation boundaries, and n∈0,…D−1n\in{0,\dots D-1} is the parameter index for each of the DD ADR dimensions. Each dimension starts with initial values from system calibration or best-guesses for the randomisation bounds, pinit2np^{2n}_{init} and pinit2n+1p^{2n+1}_{init} for the lower and upper bounds of parameter nn, respectively. Optional minimum and maximum bounds on the randomisations may also be specified, pminnp^{n}_{min} and pmaxnp^{n}_{max}. Unlike in , we choose the size of step Δn\Delta^{n} separately for each parameter. This trades off more tuning work for more stable training and the mitigation of the need for custom, secondary distributions on top of certain randomisation dimensions, as were used in that work.

Evaluation proceeds as follows: environments sample a value for each randomisation dimension uniformly between the upper and lower bounds. A fraction (40%) of the vectorised environments are dedicated to evaluation. In these environments, one of the ADR randomisation dimensions is fixed to the current lower or upper boundary (the rest of the dimensions are sampled from the aforementioned uniform distribution set by ADR). The episode proceeds to roll out; the number of consecutive successes is recorded at the end of the episode. This figure is added to a queue for the boundary of maximum length N=256N=256. Then, if mean consecutive sucesses on the boundary is above a certain threshold, tH=20t_{H}=20, the range is widened, and if the performance is below a lower threshold tL=5t_{L}=5, then the range is tightened on that bound. This is depicted in Figure 3. If on a particular step the value of a bound changes, the queue is cleared (as the previous performance data then becomes invalid). In this way, ADR will discover the maximum ranges over which a policy can perform, including for example discovering the limits of parameters impacting physics stability. Our vectorised ADR (VADR) full algorithm is described below in Algorithm 1 and Algorithm 2.

When training on multiple GPUs (8 for our best policies), we ran VADR separately on each one. This was done for two reasons: firstly, to avoid additional synchronisation overhead of buffers. Secondly, to partially mitigate the disadvantage of ADR noted below in Section 3.3 caused by the failure to model the joint distribution; having multiple independent parameter sets to some extent will allow multiple sets of extreme parameters. All randomisations are set by ADR.Except for mass and scale, which are randomised within a fixed range, as collision morphologies cannot be changed at run-time. In the following, we describe the physics and non-physics randomisations in more detail.

We apply physics randomisations to account for both changing real-world dynamics and the inevitable gaps between physics in simulation and reality. These include basic properties such as mass, friction and restitution of the hand and object. We also randomly scale the hand and object to avoid over-reliance on exact morphology. On the hand, joint stiffness, damping, and limits are randomised. Furthermore, we add random forces to the cube in a similar fashion to .

Joint Stiffness, Joint Damping, and Effort are scaled using the value sampled directly from the ADR-given uniform distribution.

Mass and Scale are randomised within a fixed range due to API limitations currently. However, we did not observe this as a significant limitation for our experiments, and our policies nevertheless achieved rollouts with high consecutive successes in the real world.

Gravity cannot be randomised per-environment in Isaac Gym currently, but a new gravity value is sampled every 720720 concurrent simulation steps for all environments.

6.2 Non-physics Randomisations

In addition to normal physics randomisations, Table 3 lists action and observation randomisations, which we found to be critical to achieving good real-world performance. To make our policies more robust to the changing inference frequency and jitter resulting from our ROS-based inference system, we add stochastic delays to cube pose and action delivery time as well as fixed-for-an-episode action latency. To the actions and observations, we add correlated and uncorrelated additive Gaussian noise. To account for unmodelled dynamics, we use a Random Network Adversary (RNA, see below).

We apply Gaussian noise to the observations and actions with the noise function

Where δ\delta and ϵ\epsilon are sampled from Gaussian distributions parameterised by the ADR values pi,pjp^{i},p^{j}, δ∼N(⋅;0,var(pi))\delta\sim\mathcal{N}(\cdot;0,\text{var}(p^{i})), ϵ∼N(⋅;0,var(pj))\epsilon\sim\mathcal{N}(\cdot;0,\text{var}(p^{j})) where var(a)=exp⁡[a2]−1\text{var}(a)=\exp\left[a^{2}\right]-1

For δ\delta, this sampling happens once per episode at the beginning of the episode, corresponding to correlated noise. For ϵ\epsilon, sampling happens at every timestep. Note that the formula for var has a cutoff at 0 noise. This allows ADR to set a certain fraction of environments to have 0 noise, which we found an important case that is not covered in previous works when setting fixed or above-zero cutoff variance (since during inference, zero white noise is added).

We apply three forms of delay. The first is an exponential delay, where the chance of applying a delay each step is pip^{i} and is given by f(x;xlast)=xlast⋅d+x⋅(1−d)f(x;x_{last})=x_{last}\cdot\mathsf{d}+x\cdot(1-\mathsf{d}) and d∼Bern(⋅;pi)\mathsf{d}\sim\text{Bern}(\cdot;p^{i}) is the Bernoulli distribution parametrised by the ii-th ADR variable, pi∈[0,1)p^{i}\in[0,1). This delay case, applied to both observations of cube pose and actions, mimics random jitter in latency times.

The second form of delay is action latency, where the action from n timesteps ago is executed. For this parameter, we slightly modify the vanilla ADR formulation to allow smooth increase in delay with ADR value despite the discretisation of timesteps. The bounds are still continuously modified, but the sampling from the range is done from a categorical distribution. Specifically, let ϵ∼U(0,b)+U(−0.5,0.5)\epsilon\sim U(0,b)+U(-0.5,0.5) be the sampled ADR value (plus random noise used to allow probabilistic blending of delay steps when sampling on the ADR boundary). Then the delay k is k=round(ϵ)k=\text{round}(\epsilon).

A third form of delay, this time on observation, is that caused by the refresh rate of the cameras in the real world. To compensate for this, we have randomisation on the refresh rate. Similarly to the aforementioned action latency, we use ADR to sample a categorical action delay d∈{1,…,delaymax}d\in\{1,\dots,delay_{max}\}. We then only update the cube pose observation if (t+r)mod  d=0(t+r)\mod d=0, effectively mimicing a pose estimation frequency of d⋅Δtd\cdot\Delta t (where r is a randomly sampled alignment variable to offset updates from the beginning of the episode randomly).

We noticed that due to heavy occlusion and caging from the fingers, our cube pose estimator exhibited occasional jumps. To ensure that the policy performance did not deteriorate and LSTM hidden state become corrupted by this, we occasionally inject completely random cube poses into the network. At the start of each episode for each environment, we sample a probability p∈U(0,0.3)p\in U(0,0.3). Then each step we sample a variable m∼Bern(⋅;p)m\sim\text{Bern}(\cdot;p), and the cube pose becomes: pose_obs=pose⋅(1−m)+random_pose⋅m\mathsf{pose\_obs}=\mathsf{pose}\cdot(1-m)+\mathsf{random\_pose}\cdot m.

Random Network Adversary, introduced in , uses a randomly-generated neural network each episode to introduce much more structured, state-varying noise patterns into the environment, in contrast to normal Gaussian noise. As we are doing simulation on GPU rather than CPU, instead of using a new network per environment-episode and wasting memory on thousands of individual MLPs, we generate a single network across all environments and use a unique and periodically refreshed dropout pattern per environment. Actions from the RNA network are blended with those from the policy by a=α⋅aRNA+(1−α)⋅apolicy\mathbf{a}=\alpha\cdot\mathbf{a}_{\text{RNA}}+(1-\alpha)\cdot\mathbf{a}_{\text{policy}}, where α\alpha is controlled by ADR.

6.3 Measuring ADR Performance in Training and in the Real World

Nats per Dimension (npd) was a metric developed by OpenAI in to measure the amount of randomisation through the average entropy across the ADR distributions. While it does not directly capture the difficulty of the environment (since each dimension is not normalised for difficulty), it provides a rough proxy for how much randomisation there is in the environment. The formula for nats per dimension is given by:

Currently, we measure real-world performance based on the number of consecutive successes. An avenue for future work is directly exploring how real-world policy performance corresponds to ADR randomisation levels in total and across different dimensions.

7 Pose Estimation

Data Generation and Processing: We use NVIDIA Omniverse Isaac Sim with ReplicatorSee https://developer.nvidia.com/isaac-sim to generate 5M images of the cube in hand in just under a day. Each image is 320×\times240 in resolution and contains visual domain randomisations as summarised in Table 5 to add variety to the training set. Such visual domain randomisations allow the network to be robust to different camera parameters and visual effects that may be present in the scene. In addition, we apply data augmentations during training on a batch in order to add even more variety to the training set. As such, a single rendered image from the dataset can provide multiple training examples with different data augmentation settings, thereby saving both rendering time and storage space. For instance, motion blur (important in our case where we have a fast-moving object to track) can be especially time-consuming at rendering time. Instead we generate it on the fly via motion-blur data augmentation by smearing the image with a Gaussian kernel with the blur direction and extent chosen randomly. The data augmentations used on the images are listed in Table 5. Each augmentation is applied with a fixed probability to ensure that the batch consists of a combination of the original as well as augmented images.

We also collect configurations of the cube in-hand generated by our policies running in the real world and play them back in Isaac Sim to render data where the pose estimates are not fully reliable i.e. the pose estimator is not accurate all the time. This happens due to the sim-to-real gap as a result of either sparse sampling or insufficient lighting randomisations for those configurations. Playback in simulation enables dense sampling of pose and lighting around these configurations with more randomisations. This allows us to generate larger datasets that improve the reliability of the pose estimator in the real world i.e. closing the perception loop between real and sim by mining configurations from the current best policy in the real world and using them in sim to render more images. We use the same intrinsics of the cameras in the real world, but randomise extrinsics when rendering data.

Training setup and inference: We use a torchvision Mask-RCNN -inspired network We also tried U-Net based direct regression of the keypoints from the image but found it to be somewhat unreliable. Sometimes the keypoints were detected at entirely different locations in the image where the cube was not even present. Having a network that first localises the cube and then detects the keypoints within the bounding box prevents such misdetections. that regresses to the bounding box, segmentation, and keypoints located at the 8 corners of the cube. The bounding box localises the cube in the image, and the keypoint head regresses to the positions of the 8 corners of the cube within the bounding box. The networks are trained with cross-entropy loss for segmentation and keypoint location regression, and smooth L1 loss for bounding box regression. We use the ADAM optimiser with a learning rate of 1e-4. The network runs on three cameras at an inference rate of 20Hz on an NVIDIA RTX 3090 GPU and a 32-core AMD Ryzen Threadripper CPU. However, because the policy was trained with a control frequency of 30Hz in simulation, the pose estimator was locked to run at 15Hz to ensure that the policy receives pose observations at a constant integer interval of once every two control steps. To make the pose estimate reliable for the downstream policy, we first perform classic PnP on each of the three cameras independently and then filter out the ones where the projected keypoints from the PnP pose do not match the inferred keypoints from the network. We triangulate the keypoints from the filtered cameras and register them against the model of the cube to obtain the pose in a canonical reference frame. We use the OpenCV implementation of PnP and roma for registering keypoints against the model of the cube. We benchmark the pose on a test set consisting of 50K images and provide results in Table 6. Since we do not use any marker-based system in the real world, we can only precisely evaluate the performance of the pose estimator in simulation. Our ablation studies in Section 3.2 do test the strength of the pose estimator for manipulation in the real world. One important difference between our approach and OpenAI et al. is that our pose estimation is not done end-to-end. Since we detect keypoints in the image and use geometric computer vision to obtain the pose, our pose estimator is not tied to a fixed camera setup, unlike .

Results

In the following section, we present the results we achieved in object reorientation in the simulations and then real world using the methods described in Section 2. We then follow it up with tests of policy robustness in reality and simulation.

For all of our experiments, we use a simulation dtdt of 160s\frac{1}{60}s and a control dtdt of 130s\frac{1}{30}s. We train with 16384 agents per GPU and use a goal-reaching orientation threshold of 0.1 rad but test with 0.4 rad as in for all experiments both in simulation and the real world. All policies are trained with randomisations described in Table 3. Most importantly, our ADR policies using the same compute resources — 8 NVIDIA A40s — achieved the best performance in the real world after training for only 2.5 days in contrast to that trained for 2 weeks to months for the task of block reorientationAlthough focused on the Rubik’s cube, they also trained for block reorientation (pp. 20, Table 3) with the same infrastructure to reproduce their results from .. Various training curves for our task a) with manual DR, b) with automatic domain randomisation (ADR), and c) ADR parameter evolution are presented in Figure 6. We note that due to differences in physics engines and hand morphology, our simulation average consecutive successes are not directly comparable, but we achieve performance on par with .

Training with manual DR takes roughly 32 hours to converge on 8 NVIDIA A40s generating a combined (across all GPUs) frame rate of 700K frames/sec. With a dt=160dt=\frac{1}{60}, this amounts to 32×70000060×24×365\frac{32\times 700000}{60\times 24\times 365} which is ∼\sim42 years of real-world experience.

2 Real-World Policy Performance

It is worth noting that, while in simulations, state information is derived directly from physics buffers, in all real-world experiments we use the pose estimator described in Section 2.7 to obtain the pose of the cube and provide it as input to the policy. The qualitative results of this are best illustrated by the accompanying videos at https://dextreme.org.

The deployment pipeline of the system is shown in Figure 7. We use three separate machines to run various components. The main machine has an NVIDIA RTX 3090, which runs both the policy as well as the pose estimator. We also do live visualisation of the real-world manipulation in Omniverse but disable the physics.

Similar to , we observe a large range of different behaviours in the policies deployed on the real robot. Our real-world quantitative results measuring average consecutive successes are illustrated in Table 7. We collect 10 trials for each policy to obtain the average consecutive successes and also collect different sets of trials across different days to understand the inter-day variability that may arise due to different dynamics, temperature, and lighting conditions. We believe such inter-day variations are important to benchmark in robotics and have endeavoured to highlight this specifically in this challenging task. We find that our policies do not show a dramatic drop in average performance, indicating that they are mostly robust to inter-day variations.

We benchmark both ADR and non-ADR (manually-tuned DR ranges) policies in Table 7 and like find that the policies trained with ADR perform the best, suggesting that the sheer diversity of data gleaned from the simulator endows the policies with the extreme robustness needed in the real world. Importantly, we observed that policies trained with non-ADR exhibited ‘stuck’ behaviours (as shown in Figure 8), which ADR-based policies were able to overcome due to increased diversity in training data. We also find that on an average, the trials with ADR achieve more consecutive successes than non-ADR policies. Table 9 puts our results in perspective alongside the previous works of and . We demonstrate performance which significantly improves upon the best vision policies from and ADR (XL)XL and XXL denote the degree of randomisations. policies given high-quality state information from a motion capture system in . Our policies do not achieve the average successes seen in with ADR (XXL) with state information. We hypothesise that this maybe due to (a) better accuracy of the state information from the PhaseSpace markers (b) higher frame-rate of state observations with PhaseSpace markers (c) increased diversity of data with ADR (XXL). However, our best vision-based policy generated ∼\sim2.5×\times higher peak consecutive successes and ∼\sim1.5×\times higher mean consecutive successes than the vision-based policies in as shown in Table 9.

Additionally, we also benchmark our results beyond the basic experiment of goal reaching by making the policy hold the cube at a target orientation NN frames in a row (we use N=10N=10), i.e., we count the number of times the cube orientation is within the threshold of 0.4 rad in a row (as described in 6), refresh the goal only if the frame counter reaches NN, and reset the counter to zero every time it goes outside the threshold. In the basic experiment of goal reaching without the hold, the cube may shoot past the target, making it difficult to tell if the target was achieved merely due to noise in the pose estimation or if the cube orientation was indeed accurately estimated. Therefore, holding the cube for NN frames in a row ensures that the goal was not achieved by chance, highlighting the robustness of the pose estimator. We conduct this experiment only with our ADR policies as shown in (Table 7, 2nd row). It is worth noting that the policy was not trained explicitly to hold the cube and that the experiment is meant to be a test of accuracy of poseTo fully hold the cube stationary in hand for a target orientation means zero velocities at the target, which requires changing the reward function.. We chose N=10N=10 based on our simulation experiments and found that setting NN too high led to a dramatic drop in the performance because the LSTM was not trained for such scenarios (see Table 8). On the other hand, setting it too low did not change the simulation performance; N=10N=10 was a good balance between performance and LSTM stability. This also lets us separate the drop in performance due to LSTM instability from pose estimation errors in the real world. From the trials, we find that while there is a noticeable drop in performance i.e. the maximum consecutive successes are only 45 as opposed to 112 in the other case, the average consecutive successes have not dropped as dramatically suggesting both the robustness of the policy and the pose estimator.

3 Quirks, Problems, and Surprises

Throughout our work, we experienced a variety of recurring problems and observations, which we disclose here in the interest of transparency and providing avenues for future research. Due to complexities involved in training and slow turnaround time in results because of real hardware involved in the loop, we do not have rigorous experiments to prove these. However, they constitute "tricks of the trade" that will hopefully be of use to other researchers.

Even though we are not training on the real robot hand and our policies are very gentle due to the low-pass filters we apply, the hand is quite fragile and requires regular health checks and cable replacements e.g. the ribbon cables connecting joints on the fingers can break or burn, disabling the corresponding joints (see Figure 9). We also apply hot glue at the either ends of the ribbon cables going into connectors to tighten the connection and prevent them from coming out.

We had applied 300lse tape on fingers and palm of the hand for another unrelated task of grasping, but it added significant friction, preventing the cube from sliding and moving smoothly on the palm. Higher friction also meant that the fingers tried aggressively to manipulate the cube when it was stuck in the hand. Therefore, we removed the tape from the palm. However, we kept the tape on the sides of the fingers to allow better grips on the object.

We found that deployment in the real world was quite variable between policies trained with different seeds, even with the same domain randomisation parameters. This is an issue that was also noted in . We suspect that this is because, despite the extreme levels of randomisation we do, there is a "null space" of possible policies which perform similarly in simulation but differently in the real world.

Automatic Domain Randomisation , which was used to achieve our best results, does not model the joint distribution and relationships between performance with different parameters. This means it can provide different results depending on the order in which randomisation ranges increase, and furthermore it may not explore the Pareto-limits of performance on any particular dimension. This can lead to cases where ADR disproportionately increases the performance of the policy on dimension A at the expense of dimension B.

Our best policy trained in a non-ADR setting exhibited slightly more natural and smooth behaviour than the ADR one. We also observed that the ADR policies tended to cage the object, resulting in letters occluded by the fingers from all cameras. This made the pose estimation unreliable in those configurations (although the policy was able to absorb the errors and performed better overall). It is possible temporal smoothing of the pose estimator can help as long as it does not introduce any latency.

We also tried pose estimation without the letters on the faces and found that to be quite challenging — the network was confused by colour changes due to lighting in the real scene and the quality of the cameras used e.g. yellow turning into white in the image when the normal of a cube face is perpendicular to the camera. Having a robot cage as used in and with an constrained artificial light can prevent this from happening to an extent, but it also makes the system confined to the cage.

On the other hand, we were surprised positively about some aspects of our system:

Our pose estimator proved to be surprisingly robust even in real-world scenarios outside of Allegro Hand manipulation. This hints at the power of extreme randomisation with synthetic data in simulation for developing such general systems.

The ability to, at test time, adjust the speed of the policy by tuning the EMA value (see Section 2.3 - we trained with 0.15 but tested with 0.1) was very useful to avoid damaging the hardware while at the same time having agile policies.

The agility of our policies in the real world on the Allegro Hand, which we initially expected to be a limitation as it is relatively large and is motor- rather than tendon-driven, was a positive surprise to us.

Most of our experiments and trials conducted on the robot were done with a worn-out thumb on the Allegro Hand — one of the wires connecting a joint to the circuit board had a loose connection — and we were quite surprised that the policies produced high consecutive successes despite a malfunctioning actuator, suggesting the robustness of learning-based approaches. The slow turnaround time involved in repairing the hardware motivated us to do it ourselves regularly during the experiments, but it was only a temporary solution.

Since we do not regress to the pose via an end-to-end network, we found that our pose estimator was not tied to a particular camera configuration as keypoint detection worked reliably from different configurationsWhile extrinsics change with different camera configurations, the intrinsics remain the same. i.e. our earlier results as described in A.5 were obtained with a camera configuration that was different to the one we used afterwards with ADR.

Our best policies were trained with continuous actions unlike OpenAI et al. , where discrete actions were used.

Related work

In-hand manipulation is a longstanding challenge in robotics. Classical methods have focussed on directly leveraging the robot’s kinematic and dynamic model to analytically derive controllers for the object in hand. These approaches work well while an object maintains no-slip contacts with the hand, but struggle to achieve good results in dynamic tasks where complex sequences of making-and-breaking of contacts is needed.

Reinforcement learning has proven to be a powerful method for learning complex behaviours in robots with many degrees of freedom. Hwangbo et al. and Lee et al. showed how robust legged locomotion can be learned even over challenging terrain using a combination of deep RL, fast simulation, and domain randomisation. A crucial inspiration for our work is OpenAI et al. on learning in-hand reorientation of cubes via reinforcement learning in simulation and subsequent OpenAI et al. extending this task to solving Rubik’s cubes. A variety of recent works have leveraged new RL techniques and simulators in order to reproduce or extend in simulation the anthropomorphic in-hand manipulation capabilities shown in . However, these works have not shown sim-to-real transfer, demonstrating that this remains a significant challenge for learned in-hand manipulation. There have also been a variety of recent approaches attempting in-hand manipulation from the perspective of finger gaiting . However, these often fail to reproduce the agile dexterity present in human hands, as the limitations of such a sequential approach to control place corresponding limits on speed. Similarly, Allshire et al. and Shi et al. achieve real-world in hand manipulation using reinforcement learning, but on a platform that cannot mimic the dexterity of a human hand. Sievers et al. learned in-hand manipulation using tactile sensing, but the lack of vision meant that the kinds of re-posing behaviour and speed of the manipulation were limited. There has also been research investigating using reinforcement learning directly on hardware platforms as opposed to in simulation . However, this limits the complexity of learning that can be done due to the small number of trials available on real hardware platforms, as well as the wear-and-tear imposed.

Pose estimation for robotic manipulation is a widely studied area . However, relatively few works outside of OpenAI et al. have applied it to the problem of contact-rich, dexterous, in-hand manipulation, which introduces challenges that exclude many off-the-shelf pose estimators (due to the large degree of occlusions, motion blur, etc.).

Limitations

Despite our best efforts, the gap between simulations and the real world is still noticeable. Our non-ADR (manual DR) based policy achieves an average of 35 consecutive successes (see Figure 6(a)) in simulation, but only obtains an average of about 15 consecutive successes in the real world. While our ADR policies do perform better, they still fail to reach the average successes seen in simulations. Another limitation of our work is that our ADR policies do not consistently obtain high consecutive successes as seen in ADR (XXL) in OpenAI et al. . It is unclear whether we need a more reliable pose estimation e.g. marker-based systems, the policy needs more diversity in the data, or if it was due to the malfunctioning thumb on our Allegro Hand. Our best npd with ADR policies is around -0.2, and it is possible that further improvements can be made. This is something we would like to investigate in future work.

We also observed that there is still some sim-to-real gap in pose estimation. This is manifested when we played back the real states in sim (real-to-sim) with physics enabled, which sometimes resulted in interpenetrations. Therefore, we were not able to easily calibrate physics parameters of the cube, which could have improved our sim-to-real transfer.

Lastly, putting this work in context, it is worth remembering that some of the key reasons for successful transfer of this task involve:

Being able to simulate the task, and having a clear and well-defined reward function so that the policies can be trained in simulation.

Randomisation of interpretable parameters exposed by the simulator and providing a curriculum for training in simulation. For instance, we randomised friction, damping, stiffness, etc., which were crucial for the transfer. ADR provided a curriculum.

Ability to evaluate successful execution of the task so that we can track the improvements. In this work, we compared the current orientation with the desired, and if the difference was within a user-specified threshold (0.4 rad in our case), the goal-reaching was considered successful.

Many real-world tasks are hard to simulate and sometimes defining reward functions is not possible as it is in this task. Even if we could simulate and have a well-defined reward function, evaluating a successful execution of a task, e.g. cooking a meal, is not straightforward.

Acknowledgements

The authors would like to thank Nathan Ratliff, Lucas Manuelli, Erwin Coumans, and Ryan Hickman for helpful discussions. Maciej Bala assisted with multi-GPU training; Michael Lin helped with hardware and Zhutian Yang with video editing. Nick Walker provided valuable suggestions on the teaser video and proofreading. Thanks also to Ankit Goyal for the help with high frame-rate video capture and Jie Xu for proofreading.

Contributions

Viktor Makoviychuk implemented the first version of the Allegro Hand environment in Isaac Gym.

Ankur Handa, Arthur Allshire, and Viktor Makoviychuk developed the domain randomisations in Isaac Gym that assisted in sim2real transfer.

Arthur Allshire developed the vectorised Automatic Domain Randomisation (ADR) in Isaac Gym.

Ankur Handa, Arthur Allshire, Viktor Makoviychuk, and Aleksei Petrenko trained RL policies and added various features to the environment to improve sim2real transfer.

Denys Makoviichuk implemented the RL Library used in this project and helped implement specific extensions required for this work.

Yashraj Narang and Gavriel State advised on tuning simulations.

Dieter Fox and Gavriel State advised on experiments.

Ankur Handa developed the code to train pose estimation and data augmentation for the vision models.

Arthur Allshire wrote the vision data rendering system and developed the domain randomisations in Omniverse. Ritvik Singh and Jingzhou Liu generalised and extended the rendering pipeline.

Ankur Handa, Ritvik Singh, and Jingzhou Liu trained pose estimation models and did real-to-sim with vision to improve the pose estimator.

Alexander Zhurkevich helped speed up the inference.

Ankur Handa conducted the real-world experiments. Arthur Allshire helped with experiments in the early stages, and Viktor Makoviychuk helped with experiments in the later stages of the project.

Arthur Allshire wrote the code to perform policy & vision inference on the real robot. Ankur Handa maintained and extended it.

Arthur Allshire developed the live visualisation pipeline in Omniverse.

Karl Van Wyk managed the Allegro Hand infrastructure to run experiments on the real hand.

Balakumar Sundaralingam and Ankur Handa provided assistance repairing the Allegro Hand when it broke.

Karl Van Wyk developed an automatic calibration system for camera-camera and camera-robot calibration.

Arthur Allshire and Ankur Handa drafted the paper.

Yashraj Narang, Ritvik Singh, Jingzhou Liu, and Gavriel State helped to edit the paper.

Gavriel State, Dieter Fox, and Jean-Francois Lafleche provided resources and support for the project.

Ankur Handa edited the videos. Gavriel State, Arthur Allshire, Viktor Makoviychuk, Jingzhou Liu, Yashraj Narang and Ritvik Singh examined the videos and provided feedback.

Ankur Handa led the project. Ankur Handa, Arthur Allshire and Viktor Makoviychuk designed the roadmap of the project.

References

Appendix A Appendix

The costs are estimated from the AWS EC2 instance pricing as of OctoberOctober 19, 2022 at https://ec2pricing.net/. The estimated costs for OpenAI et al. may have been higher in 2018 and 2019 respectively.

For the OpenAI equivalent, a p3.16xlarge 8×\timesV100 machine at 24.48/hrandac6i.4xlargewith16CPUcoresat24.48/hr and a c6i.4xlarge with 16 CPU cores at0.68/hr add up to a total cost of (24.48+0.68(24.48 + 0.68\times384)384)\times50=50=14280. Similarly, for an equivalent of OpenAI with ADR, a 32-core c6i.8xlarge machine at 1.36/hr,theoverallcostcanbeestimatedtobe1.36/hr, the overall cost can be estimated to be(24.48×\times4 + 1.36×\times400)×\times24×\times14=$215,685.

While AWS does not offer NVIDIA A40s, it does provide NVIDIA A10Gs which are similar in performance. A g5.48xlarge instance with 8 NVIDIA A10G GPUs costs at 16.288/hr.ForourmanualDRexperimentstheoverallcomputeaddsupto16.288/hr. For our manual DR experiments the overall compute adds up to16.288×\times34=553.8andforADRitis553.8 and for ADR it is16.288×\times60=977.28intotal.Thecostforsuchexperimentsonp4d.24xlargeinstancesoffering8−GPUA100machinesat977.28 in total. The cost for such experiments on p4d.24xlarge instances offering 8-GPU A100 machines at32.77/hr would be 1114.18and1114.18 and1966.2 respectively.

A.2 Hardware Comparisons

A.3 PPO Hyperparameters

A.4 Isaac Gym Simulation Parameters

A.5 Progressive Improvements in Consecutive Successes in the Real World

A.6 Default KUKA configuration

A.7 Software Tools Used in the Work

We present a model card for DeXtreme in Table LABEL:tab:model-card, following Mitchell et al. .