Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning

Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, Nikita Rudin, Arthur Allshire, Ankur Handa, Gavriel State

Introduction

In recent years, reinforcement learning (RL) has become one of the most promising research areas in machine learning and has demonstrated great potential for solving sophisticated decision-making problems. Deep reinforcement learning (Deep RL) has achieved superhuman performance in very challenging tasks, ranging from classic strategy games such as Go and Chess , to real-time computer games like StarCraft and DOTA . It has also shown impressive results in robotic settings, including legged locomotion and dexterous manipulation .

Simulators play a key role in training robots improving both the safety and iteration speed in the learning process. Training a humanoid robot that walks up and down stairs in the real world can lead to damage to its machinery and the environment, including humans that are working on the robot. An alternative is to train inside simulators that offer an efficient and scalable platform via trial-and-error with no safety issues as observed in the real world. To date, most researchers have relied on a combination of CPUs and GPUs to run reinforcement learning system . Different parts of the computer tackle different steps of the physics simulation and rendering process. CPUs are used to simulate environment physics, calculate rewards, and run the environment, while GPUs are used to accelerate neural network models during training and inference as well as rendering if required.

However, switching back and forth between CPU cores optimized for sequential tasks and GPUs which offer large-scale parallelism is by nature inefficient, requiring data to be transferred between different parts of the system at multiple points during the training process. Therefore, scalability of deep reinforcement learning in robotics is faced with two critical bottlenecks: 1) enormous computational requirements and 2) limited simulation speed. These problems are especially challenging when learning long-horizon behaviours for robots with high degrees of freedom.

Popular physics engines like MuJoCo, PyBullet, DART, Drake, V-Rep etc. need large CPU clusters to solve challenging RL tasks naturally face these bottlenecks. For instance, in , almost 30,000 CPU cores (920 worker machines with 32 cores each) were used to train a robot to solve the Rubik’s Cube task using RL. In a similar task, used a cluster of 384 systems with 6144 CPU cores, plus 8 NVIDIA V100 GPUs, and required 30 hours of training for RL to converge.

One way to speed-up simulation and training is to make use of hardware accelerators. GPUs have enjoyed enormous success in computer graphics are also naturally suited for highly parallel simulations. This approach was taken by , and showed very promising results running simulation on GPU, proving that it is possible to greatly reduce both training time as well as computational resources required to solve very challenging tasks using RL. However, some bottlenecks were still not addressed in the work – simulation was on GPU but physics state was copied back to CPU. There, observations and rewards were calculated using optimized C++ code and later copied back to GPU where policy and value networks ran. Furthermore, only simplified physics-based scenarios were trained, rather than representative robotic environments, and no attempt was made to show sim2real.

To address these bottlenecks, we present Isaac Gym - an end-to-end high performance robotics simulation platform. It runs an end-to-end GPU accelerated training pipeline, which allows researchers to overcome the aforementioned limitations and achieves 2-3 orders of magnitude of training speed-up in continuous control tasks. Isaac Gym leverages NVIDIA PhysX to provide a GPU-accelerated simulation back-end, allowing it to gather experience data required for robotics RL at rates only achievable using a high degree of parallelism. It provides a PyTorch tensor-based API to access the results of physics simulation natively on the GPU. Observation tensors can be used as inputs to a policy network and the resulting action tensors can be directly fed back into the physics system. We note that others have recently begun attempting an approach similar to Isaac Gym with respect to running end-to-end training on hardware accelerators.

With the end-to-end approach, roll-outs of observation, reward, and action buffers can stay on the GPU for the entire learning process, eliminating the need to read data back from the CPU. This set-up permits tens of thousands of simultaneous environments on a single GPU, allowing researchers to easily run experiments locally on their desktops that previously required an entire data center and to solve previously out of reach tasks using just a small GPU server.

Isaac Gym provides a straightforward API for creating and populating a scene with robots and objects, supporting loading data from the common URDF and MJCF file formats. Each environment is duplicated as many times as needed, while preserving the ability for variations between copies (e.g. via Domain Randomization ). Environments are simulated simultaneously in parallel without interaction with other environments. Using a fully GPU-accelerated simulation and training pipeline can help lower the barrier for research, enabling solving of tasks with a single GPU that were previously only possible on massive CPU clusters. Isaac Gym also includes a basic Proximal Policy Optimization (PPO) implementation and a straightforward RL task system, but users may substitute alternative task systems or RL algorithms as desired. While the included examples use PyTorch, users should also be able to integrate with TensorFlow training libraries with further customization. An overview of the system is provided in Figure 2.

Background

There are many approaches to parallelizing physics simulations. We outline these approaches here and justify our design decisions in the context of GPU-accelerated simulation tailored towards learning algorithms. Isaac Gym was developed to maximize the throughput of physics-based machine learning algorithms with particular emphasis on simulations that require large numbers of environment instances executing in parallel.

When physics simulation runs on CPU, multiple threads can be used to distribute computation among the available cores. The most straightforward strategy is simulating one environment instance per thread. In this approach, scaling is limited by the number of physical cores in the system. On a 64-core hyper-threaded CPU, we could run up to 128 environments in parallel, but CPUs with a large number of cores are typically clocked lower to prevent overheating. Running tens or hundreds of threads comes with other potential pitfalls including synchronization, context-switching overhead, and memory bandwidth limitations. To scale further, we would need to use a multi-CPU setup or build a cluster, which introduces additional communication overhead.

Running a single environment instance per thread in its own dedicated physics scene can be inefficient. There is some overhead involved in setting up, executing, and gathering the results of each physics step. The simpler the environment, the more significant the overhead. To mitigate this, we can pack multiple environments into a single physics scene. For example, we could split 1024 environments into eight physics scenes with 128 environments each. Each scene can run in its own thread. Extra provisions are needed to ensure that environments in the same scene do not interact with each other physically, which can be done using contact filtering and other methods.

1.2 GPU Simulations

Running the physics simulation on GPU can result in significant speedups, especially for large scenes with thousands of individual actors. On the GPU, the physics engine can parallelize computations at the level of individual shapes, bodies, or joints. High-end GPUs require many thousands of objects to effectively utilize their streaming multiprocessor architecture. This makes them a good match for running simulations with thousands of environment instances. On GPU, we don’t need to worry about splitting the environments into multiple scenes. In fact, the opposite is generally true - we want to pack everything into a single scene to take advantage of the deep fine-grained parallelism and maximize the overall throughput.

Physics simulations on a GPU is not new. In previous work, we demonstrated good results with running GPU physics simulations for reinforcement learning . In this work, the GPU was used as a co-processor that accelerates the physics simulation, while the API for getting physics state and applying controls was CPU-based. There are, however, performance bottlenecks with this strategy. In a reinforcement learning pipeline, physics simulation is just one part of the system. After a physics step, we need to get the latest physics state to compute observations and rewards. If these computations are done on the CPU, we need to transfer the physics state from the GPU. While modern hardware architectures can achieve impressive data transfer speeds, large simulations can incur nontrivial overhead. Then, the raw physics state needs to be processed on the CPU to compute observations and rewards, which is subject to similar parallelization challenges as discussed above due to the limited number of CPU cores. Next, the observations and rewards need to be copied from system memory back to device memory for the reinforcement learning algorithm. After the learning step, a set of actions is generated by the policy network on the GPU. These actions need to be copied to the CPU so that they can be converted to physics simulation inputs. Those inputs end up being copied to the GPU again to run the next step of physics simulation on the device.

Isaac Gym eliminates those inefficiencies by keeping all of the computations on the GPU. Stepping physics, computing observations and rewards, and applying actions are performed on the GPU without ever copying large quantities of data between devices. Two new features were added to PhysX to facilitate this. First, PhysX GPU simulations can run without fetching the results to the CPU after every step. Second, a new direct GPU API was added to access the current state, submit state changes, and apply control inputs in GPU buffers. In Figure 3, we contrast the traditional RL experience collection pipeline with our high throughput fully GPU-based pipeline.

2 Simulation Setup

Isaac Gym provides a simple procedural API to create environments and populate them with actors. It supports loading assets from URDF and MJCF file formats. These assets can be instanced multiple times in simulation environments to create actors. In the underlying PhysX engine, single-body actors are created as rigid dynamics and multi-body actors are created as reduced coordinate articulations. During the setup phase, users can set initial actor poses, configure joint drives, and customize rigid body properties and physics materials. Most joint and rigid body properties can be changed during the simulation as well, which facilitates domain randomization without stopping and restarting the simulation. Below we provide definitions of some useful terms.

Actor: An entity composed of rigid bodies connected via joints. It can be created via direct loading of a URDF model or XML file composed of either meshes or primitive shapes.

Rigid Bodies: A primitive shape or a mesh model that comprises an actor is called a rigid body. The positions, rotations and velocities of a rigid body can be obtained via the API.

DOF States: Rigid bodies are connected by various joints. A joint can have 0 or more degrees of freedom. Fixed joints have no DOFs, revolute and prismatic joints have 1 DOF and spherical joints have 3 DOFs. The DOF states, which include joint position and velocity, can be obtained via the API.

The setup code runs on the CPU to allow flexibility in per-instance setup, but once the simulation starts Isaac Gym provides a tensor API that can be used to interact with the running simulation on either CPU or GPU. Users can specify the device to be used for the simulation and the tensor interface in the simulation parameters.

3 Tensor API

Isaac Gym provides a data abstraction layer over the physics engine. This allows us to support multiple physics engines with a shared front-end API. In this work, the physics engine is PhysX, although some limited tensor API functionality is available with the FleX physics engine as well.

Instead of calling physics engine functions directly, users can access all of the physics data in flat buffers. This data-oriented approach allows us to eliminate a lot of overhead caused by looping over tens of thousands of individual simulation actors in user code. Physics state is exposed to Python users as global tensors. For example, all rigid body states can be found in a single rigid body state tensor. Figure 4 shows a typical Isaac Gym scene composed of various copies of the same environment simulating different variations all running in parallel and the corresponding tensors associated with it. Control inputs can be applied using global tensors as well. For example, applying forces to all rigid bodies in the simulation can be done using a single function call that takes a tensor containing all of the forces. Users can create custom views or slices of the global tensors to suit their needs. When multiple environment instances are packed into the simulation, it is possible to create custom views of the data with the environment index as one of the dimensions. This makes it easy to vectorize observation and reward computations by running GPU kernels on multiple environments in parallel.

The core of Isaac Gym is implemented using C++ and CUDA. It is completely independent of any Python frameworks commonly used in machine learning. To make the data easily accessible to Python users, Isaac Gym provides utilities that can "wrap" the raw data buffers as tensor objects in common machine learning frameworks like PyTorch. The tensor-wrapping utilities make it possible to share the native CPU or GPU buffers with Python without any copying overhead.

A powerful feature of Isaac Gym is the ability to run the same code on either CPU or GPU by simply toggling a flag. Python users do not need to write custom CUDA or C++ kernels to compute observations, rewards, or actions. When physics state and control tensors are wrapped as PyTorch tensors, users can take advantage of TorchScript JIT to compile their Python functions to lower level scripts which orchestrate the training pipeline quickly.

3.2 Physics State Tensors

Physics state tensors are used to obtain state snapshots of a running simulation. Isaac Gym allows for interacting with the simulation using maximal and reduced coordinates. Physics state includes the kinematic state of rigid bodies and degrees of freedom (DOFs). Rigid body state consists of position, orientation (quaternion), linear velocity, and angular velocity. DOF state includes position and velocity. In the code snippet below we show how to access them through the API.

Revolute DOFs use radians and linear DOFs use meters for units. Additional state data includes contact forces, rigid body force sensors, and DOF force sensors. To support operational space control and inverse kinematics applications, Isaac Gym also provides Jacobian and generalized mass matrices which can be obtained for articulated actors.

The available state tensors are listed in Table 1. Most of the state tensors are read-only, except the root state tensor and the DOF state tensor. These two tensors play a special role, because they can be used to fully set the poses and velocities of actors. This can be used during environment resets, when new poses are generated or original poses need to be restored. The root state tensor captures the state of the root bodies of all actors. For single-body actors, the root state fully captures their poses and velocities in maximal coordinates. For articulated actors, the root state can be used to "teleport" them without changing the poses of the descendant articulation links. The DOF state tensor can be used to configure the descendant articulation links using reduced coordinates. Setting new DOF states does not affect the root state. For fixed-base articulated actors, such as mounted robotic arms, the DOF state tensor fully captures the articulation poses and velocities. Users can apply new root and DOF states for all actors at once or to a limited subset using an index buffer. This allows resetting a subset of environments without affecting the rest.

3.3 Physics Control Tensors

Physics simulation inputs include forces, torques, and PD controls such as position and velocity targets. Forces and torques can be applied to rigid bodies and DOFs. PD targets are applied to DOFs that have been configured to use position or velocity drives. Users can configure the drive parameters like stiffness and damping using a separate API. Table 2 lists the available control tensors. The control tensors are typically created in a higher-level framework like PyTorch, but can be efficiently shared with Isaac Gym using the tensor-wrapping utilities.

Physics Simulation

Robots are simulated using PhysX reduced coordinate articulations. Any individual rigid bodies may be simulated using either maximal coordinate rigid bodies or single-link reduced coordinate articulations. Articulations with a single link and rigid bodies are equivalent and interchangeable. We also support tendons to actuate degrees of freedom and they are simulated in PhysX using Fixed Tendon mechanics. The physics of tendons are described in detail in Section A.1. We tested the dynamics of tendons using the Shadow Hand simulation environment, described in Section 6.4.1.

We use the Temporal Gauss Seidel (TGS) solver to compute the future states of objects in our physics simulation. The TGS solver uses the observation that sub-stepping a simulation with a single gauss-seidel solver iteration yields significantly faster convergence than running larger steps with more solver iterations. It folds this process efficiently into the iteration process, calculating the velocity at the end of each iteration and accumulating these velocities (scaled by dt/Ndt/N, where NN is the number of iterations) into a per-body accumulated delta buffer. This delta buffer is projected onto the constraint Jacobians and added to the bias terms in the constraints. This approach adds only a few additional operations to a more traditional Gauss-Seidel solver, producing almost identical performance cost per-iteration. However, it achieves the same effect on convergence as having sub-stepped the simulation without the computational expense. With positional joint constraints, an additional rotational term is calculated for joint anchors to improve handling of non-linear motion to avoid linearization artifacts. This term is not necessary (and in fact undesirable) to add to contacts. Various parameters exposed to the user to tune the simulator are described in Table 3.

Environments

We implemented a diverse set of environments covering different application areas. Here we describe a subset of representative examples and key points related to the training. Benchmark results on the simulation performance and training results are presented in the subsequent sections.

All environments are trained using the Proximal Policy Optimization algorithm , using rl_games, a highly-optimized GPU end-to-end implementation from . This implementation vectorizes observations and actions on GPU allowing us to take advantage of the parallelization provided by the simulator. We list the environments used in our experiments below:

While Ant and Humanoid are relatively simple environments popularised by MuJoCo continuous control benchmarks, the strength of our simulator really shines when training on environments that are rich in complexity particularly robotic hands. Various meta-data related to simulation setup for these environments is in Table 4.

Unless stated otherwise, all experiments are done on a system with a single NVIDIA A100 GPU and a single 3.7GHz Intel i7-8700K CPU

All training runs for each environment are averaged over 5 seeds. The reward curves are plotted with μ±σ\mu\pm\sigma regions.

All the environments by default follow symmetric actor-critic approach with shared observations as well as shared network for policy and value functions. Sharing the network allows faster forward passes and improves training.

Moreover, for Shadow Hand and TriFinger, we also use an asymmetric actor critic approach with policy observations that are closest to real world settings while value function receives privileged state information from simulation as well as the observations received by the policy. This approach is naturally suited for sim-to-real transfers.

For all environments trained with feed forward networks we use a discount factor of γ=0.99\gamma=0.99 while LSTM networks use γ=0.998\gamma=0.998. We use a GAE discount factor, λ=0.95\lambda=0.95 and clipping ϵ=0.2\epsilon=0.2. Also, we use an adaptive learning rate and varying KL thresholds per environment.

Detailed hyper-parameters for each training task are shown in Table 17. Rewards and observations for each environment we used can be found in Appendix A.2.

Characterising Simulation Performance

We first characterise the simulation performance as a function of number of environments. As we vary this number, we aim to keep the overall experience an RL agent observes constant by decreasing the horizon length proportionally (i.e. number of steps in PPO) for a fair comparison. While we provide detailed training studies for many environments later, we characterise simulation performance only for Ant, Humanoid and Shadow Hand as they are sufficiently complex to test the limits of the simulation and also represent a gradual increase in the complexity. All three environments use feed forward networks for training.

We first experiment with the standard Ant environment where the agent is trained to run on a flat ground. We find that as the number of agents is increased, the training time, as expected, is reduced i.e. changing the number of environments from 256 to 8192 — an increase by 5 orders of magnitude — leads to a reduction in training time to reach 7000 reward by an order of magnitude from 1000 seconds (~16.6 minutes) to 100 seconds (~1.6 minutes). However, note that Ant reaches performant locomotion at 3000 reward in just 20 seconds on a single GPU.

Since Ant is one of the simplest environments to simulate, the number of parallel environment steps per second as depicted in the Figure 5(b) can go as high as 700K. We do not observe gains when increasing the number of environments from 8192 to 16384 due to reduced horizon length.

2 Humanoid

The Humanoid environment has more degrees of freedom and requires the agent to discover the gait that lets itself balance on two feet and walk on the ground. As observed in Figure 6 and Figure 7, the training times are increased by an order of magnitude compared to the Ant in Figure 5.

We also note in Figure 6 that as the number of agents is increased, in this case, from 256 to 4096, the training time needed to reach the highest reward of 7000 is reduced by an order of magnitude from 10410^{4} seconds (~2.7 hours) to 10310^{3} seconds (~17 minutes). However, performant locomotion starts happening at around a reward of 5000 at a training time of just 4 minutes. Going beyond 4096 environments for this set up resulted in no further gains and in fact led to both increase in training time and sub-optimal gaits. We attribute this to the complexity of the environment that makes it challenging to learn walking at such small horizon lengths.

We verified this by training on another set of environment and horizon length combinations where horizon length was increased by a factor of 2 compared to Figure 6. As shown in the Figure 7, the humanoid is able to walk even with 8192 and 16384 environments which have small horizon lengths of 32 and 16 respectively but sufficiently long to enable learning.

Also worth noting that due to the increased degrees of freedom the number of parallel environment steps per second is reduced from 700K for Ant to 200K for Humanoid as shown in Figures 6 and 7.

3 Shadow Hand

Lastly, we experiment with Shadow Hand to learn to rotate a cube resting on the palm to a target orientation using the fingers and the wrist. This task is challenging due to the number of DoFs involved and the contacts that are made and broken during the process of rotation. Our results with Shadow Hand environment follow similar trends. As the number of agents is increased, in this case, from 256 to 16384, the training time is reduced by an order of magnitude from 5×1045\times 10^{4} seconds (~14 hours) to 3×1033\times 10^{3} seconds (~1 hour). We find that the environment reaches performant dexterity of 10 consecutive successes at reward of 3000 in just 5 minutes.The experiments used Shadow Hand Standard variant as explained in Section 6.4.1. Further performance improvements continue to happen as more experience is collected. Additionally, we find that the horizon length of 8 for 16384 agents still allows learning re-posing the cube. The maximum effective frame-rate of 150K number of parallel environment steps per second was achieved with 16384 agents.

Characterising Environment Performance

We now provide details and performance metrics for individual environments mentioned in Section 4 trained using a PPO implementation that operates on vectorised states and actions.

The Ant model has four legs with two degrees of freedom per leg. On A100 with 4096 agents simulated in parallel we find that ant can learn to run and achieve a reward above 3000 in just 20 seconds, and fully converge in under 2 minutes. The average simulation performance achieved during training is 540K environment steps per second. The results are shown in Figure 9(a). For details of the reward function used, we refer to Appendix A.2.1 and for the observations used, we refer to Appendix A.2.1.

1.2 Humanoid

The Humanoid environment has 21 DOFs and on a A100 with 4096 agents simulated in parallel we can train it to run — a reward threshold of 5000 — in less than 4 minutes. This is 4x faster than our previous results in obtained using the same threshold. As shown in Figures 6 and 7, we achieve peak performance for this environment at 4096 agents. Figure 9(b) shows the evolution of reward as a function of time. For details of the reward function used, we refer to Appendix A.2.1 and for the observations used, we refer to Appendix A.2.1.

1.3 Ingenuity

We train a simplified model of NASA’s Ingenuity helicopter to navigate to a target that periodically teleports to different locations. The environment with trained with 4096 agents and achieves a reward of 5000 in just under 30 seconds. Forces are applied directly to the two rotors on the chassis, rather than simulating aerodynamics. We use a gravity value of -3.721 m/s2m/s^{2} to simulate martian gravity. In Figure 9(c) we show how the reward increases as a function of time.

1.4 ANYmal Robot Locomotion

ANYmal is a robot developed by ANYbotics for industrial maintenance. It is a four-legged dog-like robot, and has been used for experiments on navigation of rough and variable terrain. We train the robot to follow target X, Y, and yaw base velocities while minimizing joint torques. The target velocities are randomized at each reset and are provided as observations alongside the positional and angular velocities of the base, the measured gravity vector, most recent actions, and DOF positions and velocities. With 4096 agents simulating in parallel, we find that the robot is able to follow the targets in under 2 minutes as shown in Figure 9(d). The reward function is defined in A.2.2

In addition to the simple flat terrain environment, we have developed a rough terrain locomotion task for ANYmal and validated the approach by transferring trained policies to the real robot. The robot learns to walk on uneven surfaces, slopes, stairs and obstacles. In addition to the observations of the flat terrain environment it receives terrain height measurements around the robot’s base. For sim-to-real transfer we extend the reward function, add noise to the observations, randomize the friction coefficient of the ground, randomly push the robots during the episode and add an actuator network to the simulation. Following the approach used in , the actuator network is trained to model the complex dynamics of the series elastic actuators of the real robot.

We implement an automatic curriculum of increasing terrain difficulties. The robots start to learn on simple versions of the terrains, and when they are able to solve a certain level the difficulty is automatically increased. In order to avoid costly terrain generation during training, we create a single mesh with all terrain types and levels and change the robots’ reset location depending on their progress. With 4096 environments, we can train the full task on NVIDIA RTX A6000 and transfer to the real robot in under 20 minutes. We refer to for more details.

2 Humanoid Character Animation

We evaluate the performance of Isaac Gym on adversarial imitation learning tasks using an implementation of adversarial motion priors (AMP) . This technique enables physically simulated humanoid character to imitate complex behaviors from reference motion data. Instead of a manually engineered imitation objective, as is commonly used in prior systems , AMP learns an imitation objective using an adversarial discriminator trained to differentiate between motion from the dataset and motions produced by the policy.

Our character is modelled as a 34-DOF humanoid , and all motion clips are recorded from human actors using motion capture. Table 12 in Appendix A.2.2 details the observation features. The adversarial training process enables the character to closely imitate a diverse corpus of motions, ranging from common locomotion behaviors, such as walking and running, to more athletic behaviors, such as spin-kicks and dancing. Effective policies can be learned with approximately 39 million samples, requiring approximately 6 minutes with 4096 environments. The implementation provided by Peng et al., 2021 requires about 1 day (30 hours) on 16 CPU cores to simulate a similar number of samples in PyBullet. Therefore, Isaac Gym provides 300x or 2.48 orders of magnitude improvement in the training time.

3 Franka Cube Stacking

We use 16384 agents to train a Franka robot to stack a cube on top of an other. In this environment, we use a slightly different choice of action space, Operation Space Control (OSC), for learning. OSC is a task-space compliant controller that has been shown to enable faster policy learning compared to joint-space controllers and learn contact-rich tasks . Our OSC implementation is fully differentiable in Isaac Gym and we obtain convergence with this controller in under 25 minutes. Figure 12 shows the training results.

4 Robotic Hands

Large-scale simulation has the ability to solve not just individual instances but whole classes of problems in robotics, by leveraging the generality of the model-free reinforcement learning framework. Dexterous manipulations is one of the most challenging problems in robotics.

To show the performance of our simulator and the ability to realistically model contact we implemented 3 different hand training environments as shown in Figure13. Shadow Hand and Allegro Hand are trained to learn cube orientation while TriFinger learns to repose the cube in 6 degrees-of-freedom involving rotation and translation. We now focus on the specific training details for these environments.

Firstly, the Shadow Dexterous Hand. We follow the standard formulation where policy and value function both receive the same input as well as OpenAI observations with asymmetric formulation and domain randomisation from . Secondly, the TriFinger robot , which shows the ability to do 6-DoF manipulation by reposing the cube to a desired position and orientation, a task which has previously shown to be challenging for model-free reinforcement learning . We use asymmetric actor-critic and domain randomisation for TriFinger and demonstrate sim-to-real transfer on a real robot. Finally, we reuse system from the Shadow Hand to the Allegro hand with minimal changes to show the generality of our approach. These three environments are depicted in Figure 13 and the corresponding reward curves in Figure 14.

As mentioned, the task with Shadow Hand is to manipulate the cube to achieve a specific target orientation and is inspired by OpenAI et al. . We train with multiple variants on the Shadow Hand environment and describe them below:

In this setting, we use a standard formulation for training where the policy and the value function use feed forward networks and receive the same input observations. The default observations we used for the Shadow Hand Standard include joint position, velocities, forces, force-torque sensors reading from each fingertip, manipulated object position and orientation, linear and angular velocities, goal orientation, relative rotation between the current object and target rotations, actions applied on the previous step. For a detailed overview of observation and reward, see Appendix A.4. Also note that this variant does not use any randomisations.

We also reproduce results with OpenAI Shadow Hand experiments in Isaac Gym with observations used in dexterity work from OpenAI et al. . A key difference between this and the Shadow Hand Standard variant is that it uses asymmetric observations. The policy receives only the input observations that are possible to obtain in the real world settings while the value function receives the same observations in addition to the other privileged information available from the simulator. This variant should make it possible to transfer the policy to the real world, mimicking the setup in . The observations for the policy and value function are provided in Table 14. We experiment with both feed forward networks (SH OpenAI FF) and LSTMs (SH OpenAI LSTM). The LSTM networks are trained with a sequence length of 4.

It is worth noting that only networks trained with OpenAI observations use domain randomisation to closely match the results in OpenAI dexterity work .

For domain randomization we closely followed the approach proposed in and applied correlated and uncorrelated noise to observations, actions, as well as randomized cube size and all the key physics properties – masses, inertia tensors, friction, restitution, joint limits, stiffness and damping. Full details of these are available in Appendix A.4.1.

We outline a few important differences between our setup and the one used in the OpenAI work below:

While OpenAI used a success tolerance of 0.4 radpage 22, section C.1, paragraph Goals in , we use both 0.4 rad and a tighter tolerance of 0.1 rad. We focus on results with 0.4 rad in this section and provide results with 0.1 tolerance in Appendix A.4.2

We use a continuous as opposed to a discrete control space used in .

Our results are averaged with 5 seeds while OpenAI show results with only 1 seedpage 11, section 6.3, Ablation of Randomizations, Figure 8 in .

The randomizations used in our work do not include action delay and motor backlash.

We use an LSTM layer of 1024 hidden units after the input followed by an MLP layer of 512 hidden units. On the other hand OpenAI et al. used an MLP layer of size 1024 after the input followed by an LSTM layer of size 512 hidden units. We found our setting performs better with Isaac Gym.

We use a somewhat different reward function to OpenAI as shown in Appendix A.2.3.

Our experiments are only in simulation and unlike we do not attempt any sim-to-real transfer for the Shadow Hand experiment.

Figure 14(a), (b) and (c) show the reward curves for various settings we used for Shadow Hand. Shadow Hand Standard — trained with no randomization and uses symmetric actor critic setting with a feed forward network — is the fastest to reach a reward of 6000. This setting achieves 20 consecutive successes in under 35 minutes. Important to remember that this setting is not suitable for sim-to-real transfer as it includes some observations that may not be directly available in the real world.

We now focus on experiments with OpenAI observations and asymmetric feed-forward actor-critic. This setting is suited for sim-to-real transfer and the policy uses only the observations that are possible to obtain in the real world. As shown in Figure 15(b), we achieved more than 20 consecutive successes in less than 1 hour. In contrast, for the same performance it takes 30 hours on the OpenAI setup consisting of CPU based simulation and training setup running MuJoCo simulator on a cluster of 384 16-core CPUs with 6144 CPU cores in total and using 8 NVIDIA V100 GPUs for training. In Figure 15(a) we show that using LSTM networks, the performance increases and we can reach 37 consecutive successes in just under than 6 hours while OpenAI et al. achieve same performance in ~17 hours. Since OpenAI et al. show results only with 1 seed, comparing their result with our best seed we note that 37 consecutive successes with LSTM experiments can be achieved in just 2.5 hours. We provide the results for Shadow Hand OpenAI experiment with success tolerance of 0.1 in the Appendix A.4.

4.2 TriFinger

The TriFinger manipulation task, originating in , involves picking a cube lying on a flat surface and repositioning it to a desired 6-degrees-of-freedom pose. The manipulator has 3 fingers each with three degrees of freedom. In , it was shown that Isaac Gym training combined with Domain Randomization allows sim-to-real transfer. The environment is shown in Figure 13.

We use an asymmetric actor-critic formulation for this system as that allows to design a policy that uses input observations that are possible to obtain in the real world and therefore enable sim-to-real transfer. We show the reward and success rate in simulation in Figure 16. We also transfer results from simulation to the real world and note that our mean success rate in the real world is 55%. We refer to for more detailed analysis.

In particular, this example shows the ability of policies learned using Isaac Gym’s physics to generalize to the real world. Some of the behaviours leaned by the policy are shown in the Figure 17. It is worth noting that the robot is situated in a different location and therefore the sim-to-real transfer was done remotely.

4.3 Allegro Hand

We learn cube orientation with Allegro Hand and use the same reward as for the Shadow Hand as well similar observation scheme, with the only difference — smaller number of observations because of the different number of fingers in Allegro Hand — that it has 4 fingers instead of 5 and fewer degrees of freedom as a result, shown in Appendix A.2.3.

Figure 14(d) shows the reward curves for Allegro Hand and Figure 15(d) shows consecutive successes achieved. Interestingly, despite having fewer degrees of freedom this hand does not achieve as high consecutive successes as Shadow hand. This is because the wrist is fixed and fingers are slightly longer. We observed in Shadow hand experiment that having a movable wrist allows for better manipulation when reorienting the cube.

Summary

We show that Isaac Gym is a high performance and high-fidelity framework that allows blistering fast training on many challenging simulated robotic environments on a single NVIDIA A100 GPU that previously would have required large heterogeneous clusters of CPUs and GPUs using a conventional RL setup with CPU-only simulators. Moreover, the simulation backend is also suited for learning contact-rich manipulations as confirmed by our sim-to-real transfer demonstrations with ANYmal locomotion and TriFinger cube reposing.

Acknowledgements

We would like to thank the following for additional hard work helping us with this work.

Jonah Alben, Rika Antonova, Ayon Bakshi, Dennis Da, Shoubhik Debnath, Clemens Eppner, Dieter Fox, Animesh Garg, Renato Gasoto, Isabella Huang, Andrew Kondrich, Rev Lebaredian, Qiyang Li, Jacky Liang, Denys Makoviichuk, Brendon Matusch, Hammad Mazhar, Mayank Mittal, Adam Moravansky, Yashraj Narang, Oyindamola Omotuyi, Fabio Ramos, Andrew Reidmeyer, Philipp Reist, Tony Scudiero, Mike Skolones, Balakumar Sundaralingam, Liila Torabi, Cameron Upright, Zhaoming Xie, Winnie Xu, Yuke Zhu, and the rest of the NVIDIA PhysX, Omniverse, and robotics research teams. We also thank Jason Peng and Josiah Wong for the help in AMP and Franka Cube Stacking experiments.

Thanks are also due to open-source community projects like Matplotib, Python, NumPy, PyTorch, Tensorboard, Tensorboard Aggregator and SciencePlots which we used heavily in this work. We are thankful to Overleaf for hosting our latex project.

References

Appendix A Appendix

We simulate tendons as part of the Shadow Hand environment and describe the details of this simulation here.

Fixed tendons are an abstract mechanism that couple degrees of freedom (DOF) of an articulation. A fixed tendon is composed of a tree of tendon joints, where each joint is associated with exactly one axis of a link’s incoming articulation joint. In the following, when we refer to a tendon joint’s position, we mean the position of the axis of this associated articulation joint.

In addition, each tendon joint has a coefficient that determines the contribution of the (rotational or translational) joint position to the length of the tendon, which is evaluated recursively by traversing the tree: The length at a given tendon joint is the length at its parent tendon joint plus its joint position scaled by the coefficient.

Given the tendon length at each joint, the tendon applies a spring force (or torque) to the joint’s child link that is proportional to the deviation of the tendon length from a desired (tendon-wide) rest length. An equal and opposing force is applied to the parent link of the root tendon joint; conceptually, each tendon joint is a virtual joint drive between the root parent link and the tendon joint’s child link. In addition to the spring force, the tendon-joint applies a damping force that is proportional to and acting against the velocity of the virtual root-to-child link joint.

Analogous to the length dynamics, the tendon supports length limits that apply an additional force or torque that is proportional to the deviation from set limits.

A.1.2 Spatial Tendons

Spatial tendons create line-of-sight distance constraints between links of a single articulation. In particular, spatial tendons run through attachments that are positioned relative to an articulation link, and their length is defined as a weighted sum of the distance between the attachments in the tendon. It is possible to create multiple attachments per link, for example for tendon-routing purposes. In contrast to fixed tendons, spatial tendons are not constrained to follow the articulation topology.

Same as fixed tendons, spatial tendons may branch, in which case the tendon splits up into multiple conceptual sub-tendons, one for each root-to-leaf path in the tendon tree. Length and limit constraints are evaluated per sub-tendon, and have spring-damper dynamics that may both contract and extend the tendon (one may use appropriately set limits to achieve a one-sided, string-like constraint).

The sub-tendon constraint force acts on the leaf and root attachments, in the direction of its parent for the leaf, and in the direction of the child on the path to the sub-tendon leaf for the root. However, the force does not propagate further and act on any intermediate attachments between root and leaf.

A.2 Observations & Rewards

In this section we describe the reward and observations for each environment in detail.

Both the Ant and Humanoid environments use the same reward formulation, namely:

Reward described in Section A.2.1. Observations detailed in Table 6.

Reward described in Section A.2.1. Observations detailed in Table 6.

A.2.2 Locomotion environments

Observations detailed in Table 8. The reward function is as follows:

For the included flat-terrain environment, observations are detailed in Table 8 and the reward function is as follows:

Reward terms are defined in Table 10 and symbols in Table 10.

For rough terrain locomotion with sim-to-real, we extend the observations with 140 terrain heights around the robot’s base and use the more complex reward function:

AMP learns an imitation objective using an adversarial discriminator DD, trained to differentiate between motion from the dataset M\mathcal{M} and motions produced by the policy π\pi,

where pM(s,s′)p_{\mathcal{M}}(s,s^{\prime}) denotes the likelihood of observing a state transition from ss to s′s^{\prime} in the motion data, and pπ(s,s′)p_{\pi}(s,s^{\prime}) is likelihood of a state transition under the policy. The discriminator can then be used to specify rewards rtr_{t} for training a policy to imitate behaviors shown in the motion data

This objective, in effect encourages the policy to produce behaviors that fool the discriminator into classifying them as behaviors from the reference motion data.

Observations are detailed in Table 12. The reward function used is as follows:

We set wstack=16.0\mathsf{w_{stack}}=16.0, walign=2.0\mathsf{w_{align}}=2.0, wlift=1.5\mathsf{w_{lift}}=1.5, and wreach=0.1\mathsf{w_{reach}}=0.1

A.2.3 Robotic Hands

The reward function for Shadow Hand is as follows:

where wdist=−10\mathsf{w_{dist}}=-10 and wact=−2e−4\mathsf{w_{act}}=-2e-4.

There are two different variants of observations used. In the Shadow Hand Standard environment, the observations are as shown in Table 14. In The ShadowHand OpenAI environment, in order to compare to compare to , we use observations as shown in Table 14. Further details of the Shadow Hand environments are available in Appendix A.4. Below we provide the code snippet to compute the reward as used in our implementation.

Δit\Delta^{t}_{i} denotes the change across the timestep of the fingertip distance to the centroid of the object and was found to be helpful in . Formally, Δit=∣∣fti,t−tcurr,t∣∣2−∣∣fti,t−1−tcurr,t−1∣∣2\Delta^{t}_{i}=||\mathsf{ft_{i,t}}-\mathsf{t_{curr,t}}||_{2}-||\mathsf{ft_{i,t-1}}-\mathsf{t_{curr,t-1}}||_{2}, where tcurr,t\mathsf{t_{curr,t}} is position of the cube centroid and fti\mathsf{ft_{i}} denotes the position of the ii-th fingertip at time tt.

rot_dist\mathsf{rot\_dist} is the angluar difference between the current and target cube pose, rot_dist=2×arcsin⁡(min⁡(1.0,∣∣qdiff∣∣2)),qdiff=qcurrqtarget∗\mathsf{rot\_dist}=2\times\arcsin(\min(1.0,||\mathsf{q_{diff}}||_{2})),\mathsf{q_{diff}}=\mathsf{q_{curr}}\mathsf{q_{target}^{*}}. Following , a logistic kernel is used to convert tracking error in euclidean space into a bounded reward function, with K(x)=(eax+b+e−ax)−1\mathcal{K}(x)=\left(e^{ax}+b+e^{-ax}\right)^{-1}, where aa is a scaling factor; we use a=50a=50. See for a more thorough motivation and description of these reward terms.

The reward formulation is identical to that used in Shadow Hand - see Appendix A.2.3. The observations are also identical, save for the change in number of fingers.

A.3 Hyperparameters for Training PPO

A.4 Shadow Hand Details

As mentioned previously, we implemented two variants of the Shadow Hand environment. The Standard variant uses privileged policy observations and no Domain Randomization, in order to provide a quick training example to test Reinforcement Learning algorithms on. The OpenAI variant uses asymmetric observations, such that it would be possible to transfer the policy to the real world, mimicing the setup in .

Isaac Gym implements a high-level API that simplifies setting up physics domain randomization parameters and schedule in yaml configuration files and is very extensible. Here we detail the randomization parameters that we used.

Following unmodeled dynamics is represented by applying random forces on the object. The probability pp that a random force is applied is sampled at the beginning of the randomization episode from the loguniform distribution between 0.1%0.1\% and 10%10\%. Then, at every timestep, with probability pp we apply a random force from the 33-dimensional Gaussian distribution with the standard deviation equal to 1 m/s21~{}m/s^{2} times the mass of the object on each coordinate and decay the force with the coefficient of 0.990.99 per 5050ms.

Physical parameters like friction, link and object masses, cube size, joint and tendon properties, as well as correlated noise parameters are randomized every time an environment is reset, with a minimum interval of 720 steps. Table 18 lists all physics parameters that are randomized.

A.4.2 OpenAI Observations

We conduct experiments with Shadow Hand OpenAI observations with a tighter success tolerance of 0.1 radians and show the reward curves as well as the consecutive successes achieved with this training in Figure 18 and 19.

We achieve 20 consecutive successful cube rotations after training in just under 1 hour. This is similar to the performancepp 13, Section 6.5 titled Sample Complexity & Scale achieved by OpenAI et al. but with a cluster of 384 16-core CPUs and 8 V100 GPUs with training for 30 hours while we only need a single A100.

Using sequence networks like LSTMs improve the performance and we find that we are able to achieve 37 consecutive successful cube rotations after training in just under 6 hours. OpenAI et al. achieve similar performance in about 17 hours again on a cluster of 384 16-core CPUs and 8 V100 GPUs. We use a sequence length of 4 to train the LSTM. Various other parameters for this set up are in Table 17.

We also note that training with a tolerance of 0.1 rad and testing with a tolerance of 0.4 rad, we are able to even go up to 44 consecutive cube rotations.