Barkour: Benchmarking Animal-level Agility with Quadruped Robots

Ken Caluwaerts, Atil Iscen, J. Chase Kew, Wenhao Yu, Tingnan Zhang, Daniel Freeman, Kuang-Huei Lee, Lisa Lee, Stefano Saliceti, Vincent Zhuang, Nathan Batchelor, Steven Bohez, Federico Casarini, Jose Enrique Chen, Omar Cortes, Erwin Coumans, Adil Dostmohamed, Gabriel Dulac-Arnold, Alejandro Escontrela, Erik Frey, Roland Hafner, Deepali Jain, Bauyrjan Jyenis, Yuheng Kuang, Edward Lee, Linda Luu, Ofir Nachum, Ken Oslund, Jason Powell, Diego Reyes, Francesco Romano, Feresteh Sadeghi, Ron Sloat, Baruch Tabanpour, Daniel Zheng, Michael Neunert, Raia Hadsell, Nicolas Heess, Francesco Nori, Jeff Seto, Carolina Parada, Vikas Sindhwani, Vincent Vanhoucke, Jie Tan

I Introduction

There has been a proliferation of legged robot development inspired by animal mobility. Recent notable examples include the ETH ANYmal , the MIT Mini Cheetah , the KAIST RaiBo , Unitree A1/Go1, and the Boston Dynamics Spot robots. An important research question in this field is how to develop a controller that enables legged robots to exhibit animal-level agility while also being able to generalize across various obstacles and terrains. Through the exploration of both learning and traditional control-based methods, there has been significant progress in enabling robots to walk across a wide range of terrains . These robots are now capable of walking in a variety of indoor and outdoor environments, such as up and down stairs, through bushes, and over unpaved roads and rocky or even sandy beaches.

Despite advances in robot hardware and control, a major challenge in the field is the lack of standardized and intuitive methods for evaluating the effectiveness of locomotion controllers. Ad-hoc metrics are often used to present results, which complicates the comparing of results. To address this issue, it is essential to establish metrics that can accurately measure robot agility and to define a standard set of tasks that can serve as a common evaluation framework, similar to how the DeepMind Control Suite has been widely adopted in the field of reinforcement learning (RL).

A good benchmark set for agile legged locomotion should (1) be non-trivial or not easily exploitable, and (2) require a diverse set of primitive behaviors that quadrupeds showcase in real environments. Well-established benchmarks already exist to measure animal performance, for example, dog agility competitions . In these competitions, participants race their dogs through a pre-set obstacle course. A variety of obstacles including weave poles, jumps, tunnels, an A-frame, a seesaw, and a pause table test diverse locomotion skills. The performance is evaluated based on time, and there are penalties for errors such as completing obstacles in the wrong order, tackling an obstacle from the wrong direction, or touching the jump bars while leaping.

To solve the tasks in our Barkour benchmark suite, we introduce a simulation setup and two learning-based baselines as references. Our first approach involves training specialist policies in simulation that can overcome each individual obstacle. The specialist policies are then orchestrated by a high-level navigation controller that selects the appropriate specialist policy based on the location of the robot. In our second approach, we take inspiration from recent work on training generalist agents and develop a transformer-based generalist locomotion policy, named Locomotion-Transformer, which tackles all Barkour obstacles using a single policy network. To demonstrate the effectiveness of the learned agile skills, we deploy the simulation-trained policies in a zero-shot manner on a custom-built quadruped robot in a real-world Barkour setup.

Our contributions can be summarized as follows:

A benchmark (Barkour) for agile quadruped robot locomotion inspired by dog agility competitions.

Two learning-based approaches (specialist and generalist (Locomotion-Transformer) policies) that can complete the benchmark agilely which can serve as baselines to benchmark future algorithms.

A detailed analysis of zero-shot sim-to-real transfer using a custom-built quadruped robot.

II Related Work

Benchmarks are a driving force behind the development of artificial intelligence methods, such as ImageNet for computer vision and OpenAI Gym for Reinforcement Learning. In the field of legged robots, while many prior works focused on creating new algorithms and controllers, limited effort has been directed towards creating a systematic benchmark to assess the performance of these controllers, especially in the context of agility . Among these efforts, Eckert & Ijspeert proposed a suite of 1313 metrics to measure different aspects of the agility of legged robots including leaping, standing, balancing, and climbing. Their metrics are carefully designed and measure the robot performance in a comprehensive way. However, there are a few challenges in directly leveraging these metrics to develop a general agile locomotion controller: 1) a standardized environment was not provided as part of the benchmark, making it difficult for different groups to quickly iterate on a common set of tasks, and 2) having 1313 metrics for different agility skills means that researchers may opt to focus on a subset of the metrics instead of trying to push agility as a whole. Barkour is complementary to the metrics proposed in , with the key difference being the inclusion of a standardized and extensible obstacle course environment as well as a single metric to measure the overall performance of the controller in an intuitive way.

Combining reinforcement learning with dynamics randomization , system ID , locomotion primitives , and proper reward shaping , legged robots can exhibit stable locomotion blindly on general uneven terrains at moderate speeds. Further advancements in perception and vision have increased the robustness and adaptability of legged robots. The pioneering work from Miki et al. demonstrated how their robot, ANYmal, can walk on various types of uneven terrain using a noisy state estimation and heightfield . Rudin et al. demonstrated that by combining these insights with a highly parallelized simulator, one can obtain a high-performing locomotion controller within minutes . Recently, Agarwal et al. showed that it is possible to perform many visual locomotion tasks by consuming raw camera images in an LSTM-based network . Our approach builds on the findings of and leverages a GPU-based simulation environment similar to . However, our work distinguishes itself from previous studies by emphasizing both agility and generalization, and by showcasing several agile skills performed together.

III Barkour benchmark for measuring agility

Making robots move like animals or humans is a goal shared by many researchers in the robotics community. Animals and their behaviors have been an inspiration for many robots. In this work we aim to make this connection quantitative by introducing Barkour, an agility benchmark that assesses the overall agility skills of quadruped robots. The Barkour benchmark is inspired by dog agility competitions, which we believe provides a setting to quantitatively evaluate small to medium-sized quadruped robots. Dogs exhibit a wide range of agile locomotion skills and behaviors and dog agility events have well-defined rules. Dog agility competitions capture a diverse set of skills and have clear metrics based on timing and behaviors to compare performances of different dogs. A careful design of obstacles, metrics, and coverage of skill diversity is also crucial for creating a benchmark for robotics. Hence, we reuse many elements common to dog agility competitions: obstacle definitions, a time-based metric, and penalties for rules violations.

The agility score RagilityR_{\text{agility}} measures how fast a robot can successfully complete all obstacles in Barkour. The score is calculated based on real dog competitionsRegulations for Agility Trials and Agility Course Test (ACT) rules/scoring from the American Kennel Club (AKC). See standard course time for 8-inch Division Novice A and B Agility Standard Class. with some simplifications. A score of 1.01.0 indicates that the robot solved the entire course within allotted course time tallottedt_{\text{allotted}}. Starting from 1.01.0, the robot can receive two types of deductions: a 0.10.1 penalty for each failed or skipped obstacle and a 0.010.01 penalty for each full second the robot exceeds tallottedt_{\text{allotted}}. An episode is completed when the robot reaches the end table, otherwise it is terminated when RagilityR_{\text{agility}} reaches . The scoreTypically represented as $$ points in real dog competitions. calculation is given by:

In the context of this paper, we consider four types of obstacles with nominal obstacle sizes and allotted times given in Table I. Appendix -B provides detailed definitions of each type of obstacle, including the physical setup, acceptance criteria, etc.

Note that real dog competitions often include other types of penalties (e.g., smaller deductions if a dog retries an obstacle), whereas Barkour only deducts points for failed/skipped obstacles or excess time. This makes implementing the scoring mechanism in both sim and real much easier and significantly less error prone.

In this work, we chose a small area with a small number of obstacles and simpler metrics for ease of repeated and controlled hardware experimentation, and kept diversity of skills and measurability as high priorities. As is the case for real dog agility competitions, our proposed metric, environment, and simulation setup can be easily adapted to a different number of obstacles, a larger area, or other variables.

IV Two Baseline Solutions

One goal of this work is to demonstrate that by working towards our proposed Barkour benchmark, we can advance our techniques and insights in obtaining controllers with better agility. To achieve this, we present two learning-based methods to synthesize agile controllers for a real quadruped robot and establish baselines for the proposed Barkour benchmark. An overview of the two learning methods is shown in Fig. 3.

In the first baseline, we train three specialist policies for different tasks in a physics simulation with an on-policy Reinforcement Learning algorithm (Section IV-A). Each specialist policy is trained with a reward function and a terrain curriculum tailored to the corresponding task. We apply domain randomization for transferring the simulation-trained policies to the real world. We then design a high-level navigation controller (Section IV-C) that switches between different specialist policies and generates velocity commands for the specialist policies to overcome different obstacles.

In the second baseline, we train a single generalist policy that can handle all Barkour tasks. To achieve this, we collect a dataset with all the specialist policies and distill them into a unified Transformer-based locomotion policy, which we name Locomotion-Transformer. We then combine Locomotion-Transformer with a simplified navigation controller that only needs to provide the waypoints for the low-level policy to tackle the full obstacle course.

Our first baseline trains three specialist policies to cover all the core agility skills required in the Barkour benchmark: fast omni-directional walking on uneven terrain, climbing up and down a slope, and jumping over a board. The first policy is used to solve the start/end tables and weave pole tasks, the second solves the A-Frame task, and the third solves the Broad Jumping task. All the specialist policies are trained using PPO in LeggedGym , which uses the GPU-powered NVidia IsaacGym. Specialist model architecture details can be found in Appendix -H. We also investigated the TPU-powered Brax simulator for training, as presented in Appendix -L.

In our approach, perception is modeled as a heightfield in the vicinity of the robot following the method used by Miki et al . To construct the heightfield svs^{\text{v}}, we select a grid of locations on the terrain surrounding the robot and measure their relative heights with respect to the center of the robot’s torso. This grid translates and rotates with the robot, always maintaining alignment with its heading. The shape of the heightfield is customized to suit the requirements of each task, which are elaborated in the corresponding sections below.

We choose the action space to be the target motor angles measured from a nominal standing pose.

IV-A2 Reward Function

Similar to Rudin et al. , our reward function consists of multiple terms that fall into two categories: Task and Regularization. Task rewards encourage the robot to perform the desired skills, such as running forward or turning, while regularization rewards shape the behavior of the robot, such as low energy consumption and high stability. Section -H lists the reward terms that we use in this work.

IV-A3 Omni-directional Walking Policy

The Omni-directional Walking Policy (OWP) is trained to follow a velocity command in all directions on uneven terrains. It demonstrates the ability of the controller to quickly adjust the speed profile and the robot’s orientation, which are critical skills for the weave pole and pause table tasks.

OWP takes all three categories of observations as input and is trained with randomly sampled velocity commands. The training environment for OWP follows the general uneven terrain curriculum designed by Rudin et al , which consists of mild slopes, stairs, and random steps. Please refer to Appendix -E for details of the observation space, the velocity sampling, and the reward structure.

IV-A4 Slope Climbing Policy

IV-A5 Jumping Policy

IV-A6 Domain Randomization

IV-B Locomotion-Transformer: A Generalist Locomotion Policy

Our second baseline aims to train a single generalist locomotion policy that can tackle all obstacles, thereby removing the necessity for ad-hoc switching between individual specialist policies and promoting generalization capabilities to different obstacle and terrain configurations. Our generalist policy, Locomotion-Transformer, is trained by distilling the specialist policies via behavioral cloning with a Transformer sequence model , similar to . Although prior work has demonstrated that generalist policies can also be learned via reinforcement learning by simultaneously training on multiple environments, the diverse curricula and reward structures in our problem setting makes it challenging.

Collecting the right interaction data that covers the state distribution on which the policy is going to operate is critical to successful policy distillation . To learn the generalist, we collect an offline dataset by rolling out individual specialist policies in simulated environments that are same as the training settings of the specialist. We collect 17,63617,636 simulated episodes, equivalent to 57.5857.58 hours of robot time, across four different types of environments: random steps, stairs, gaps, and slopes. More details on the data collection setup used in this work and a data card (Table VI) can be found in Appendix -J.

IV-B2 Transformer Model

As shown in Figure 3, Locomotion-Transformer is a causal Transformer that takes a history of velocity commands vˉ\bar{v}, proprioceptive states s0:tps^{\text{p}}_{0:t}, actions a0:t−1a_{0:t-1}, and the most recent heightfields stvs^{\text{v}}_{t} over a fixed context window as inputs in the following order:

svs^{\text{v}} contains all types of heightfields that are originally designed for each individual environment (see details in Section IV-A). The model predicts the next action ata_{t} at the last position, and is trained on an L2 regression loss. We tokenize the most recent elevation image with a two-layer convolutional encoder network for each type of heightfield. The proprioceptive states (along with the velocity commands) and actions are each tokenized with one projection layer. We use a context window of 0.30.3s, the same as in the specialist policy, which amounts to a size of W=15W=15. For terrain perception, we combine both heightfields from OWP and JP to cover the terrain near and in front of the robot. We use ReLU activation to encode each observation input. The model architecture hyperparameters can be found in Appendix -K.

IV-C High Level Navigation Controller

The specialist policies allow the quadruped to tackle different individual obstacles. However, to complete the Barkour benchmark, the robot must also make decisions on how to navigate the field and transition between behaviors while following specific rules of the obstacle course. To achieve this, we design a high-level state machine-based navigation controller to command our locomotion policies and guide the robot through the entire course. The navigation controller has access to the full state of the environment and the ground truth robot pose, and outputs appropriate velocity commands for the robot.

Specifically, we place a sequence of waypoints around the obstacles that serve as sub-goals to guide the robot through each obstacle in the Barkour benchmark. Each waypoint specifies a desired position and heading orientation for the robot to reach. Every timestep, the high-level navigation controller computes linear and angular velocity commands to take the robot to the desired poses. An error tolerance is also encoded in each waypoint to determine when to switch to the next waypoint. When using the specialist policies, each waypoint also includes a behavior type to dictate which specialist policy to use. On the other hand, the generalist Locomotion-Transformer policy only needs the velocity commands. More details about this calculation can be found in Appendix -I.

V Robot Hardware

Exploring the influence of the Barkour benchmark on enhancing the agility of quadruped robots necessitates considerable controller development and thorough real hardware experimentation. This poses significant challenges on the reliability and repeatability of the robot hardware, especially given the highly agile movements we strive for. Moreover, quadruped animals exhibit diverse body configurations compared to typical quadruped robots, which can significantly affect their capacity for agile motion. Consequently, we believe that hardware optimization and customization are critical in bridging the agility gap between legged robots and their animal counterparts.

VI Experiments and Discussions

We evaluate the specialist and generalist policies within our hierarchical control framework on Barkour tasks. We aim to answer the following questions:

Can we solve Barkour benchmark using the frameworks proposed in Section IV and how does it compare to animal agility?

How important are the design choices we made during specialist and generalist policy training?

The benchmark is effective for several reasons. Firstly, it requires diverse motion capabilities to complete the course, which effectively exposes potential limitations on agile skill discovery. Secondly, the benchmark effectively tests the maneuverability and control precision of the locomotion controller at high speeds due to the complex routes required to finish the course as illustrated in Fig. 6, and scoring being tied to the completion time of the course. In the event that the robot misses the correct gate in the weave poles section, touches the jump board, or if the robot is too slow, the overall score is penalized. Lastly, we found that the benchmark is well-suited for real animal counterparts, as demonstrated in our experiments involving two small dogs. The comparison between the performance of the real dogs and robots highlights opportunities for improvement in both hardware (such as the need for flexible spines) and algorithmic approaches. It is worth noting that the real dogs never failed to complete any of the obstacles and were significantly faster than our provided baselines.

VI-B Barkour Evaluation using Combined Specialist Policies

The decomposition of these runs (Table IV into the different obstacles shows that the specialist policies complete the weave poles at around the nominal (target) speed, while the broad jump is faster and the A-frame is slower. The robot completes the weave poles and A-frame with a 100% success rate, while the broad jump, requiring particularly more agile behavior, succeeds 38% of the time. We also provide the detailed analysis of each specialist policy below:

VI-B2 Slope Climbing Policy

VI-B3 Jumping Policy

VI-C Barkour Evaluation using Locomotion-Transformer

We now evaluate the performance of the generalist policy, which is distilled using the data generated by the specialist policies. Unlike the combined specialist policies that require the high-level navigation policy to select which specialist policy to use, our Locomotion-Transformer model absorbs all the low-level skills of the specialists and thus can be deployed on the robot without explicit hints on which obstacle it is handling. Therefore, we use a simplified navigation policy that only outputs the velocity command to the policy during evaluation.

We evaluate Locomotion-Transformer on the real Barkour setup with 1919 trials, which can be seen in Figure 8. In general, we observe slightly lower performance for the Locomotion-Transformer policy compared to the combined specialist policies (also seen in Table III). A major reason for this is that the Locomotion-Transformer policy does not have information about the obstacle type being tackled and has to infer relevant information from terrain and proprioceptive observations, which increases the difficulties of the task.

On the other hand, we want to highlight two advantages of the generalist Locomotion-Transformer policy. First, by using a single generalist policy, we achieve smoother transitions between different obstacles. As seen in the supplementary video, when transitioning between different specialist policies, the robot may exhibit jerky motions when the previous specialist policy reaches a state that is unfamiliar to the following policy. Meanwhile, the Locomotion-Transformer achieves smooth transitions among different behaviors including different gaits.

We have shown that the generalist Locomotion-Transformer policy is able to execute a wide range of high-level commands provided by a navigation controller and achieves a high Barkour score. However, the training is based on simulation data from specific expert policies, which raises two follow-up questions:

Is the Transformer-based framework also capable of learning high-level behaviors?

Is it possible to train a Locomotion-Transformer policy from a hardware dataset?

We answer these questions affirmatively in Appendix -D by training and deploying a course-specific Locomotion-Transformer policy from hardware data.

VI-D Is Training Specialist Policies Necessary?

We have demonstrated that it is possible to train individual specialist policies for each task and distill them into one generalist policy to solve the Barkour benchmark. However, one question remains: Is it possible to train a single agent using RL that combines the capabilities of all specialist policies subsection IV-A?

To answer this question, we train a single multi-task policy using our IsaacGym-based training pipeline. We construct a training terrain curriculum by mixing the terrains from all three specialist policy types. During RL training, 50%50\% of the data are from the slope environment for training Slope Climbing Policy, 30%30\% are from the general uneven terrain for training Omni-directional Walking Policy, and 20%20\% are from the gap environment for the Jumping Policy. We follow the same curriculum for each terrain type as described in Section IV-A. We use the reward function from training the Omni-directional Walking Policy to obtain a policy that takes a velocity command as input.

We deploy the trained policy on the real robot and find that the resulting policy can robustly walk down the start table and finish the weaving pole task. However, it cannot climb up the A-Frame even though we have already adjusted the training data distribution to bias towards the slope task. This demonstrates that the steep slope necessitates the use of a specialist training. In addition, the policy cannot successfully perform the broad jump. This experiment also demonstrates the difficulties and diversity of the skills required to solve the proposed Barkour benchmark.

VI-E Locomotion-Transformer Ablations

We investigate several design choices behind Locomotion-Transformer, including model architecture, model size, training dataset size, and context length. To enable rapid model iteration, we report PyBullet simulation Barkour scores. For each experiment, we train three models with different random seeds and report the mean and standard deviation of the respective mean evaluation scores.

To understand if the transformer architecture is necessary, we trained an MLP (a larger version of the specialist policy architecture) on the same dataset. As shown in Fig. 11(a), the distilled MLP policy performs substantially worse than the Transformer policy.

VI-E2 Context length

In Fig. 11(b), we show the Locomotion-Transformer performance with different context lengths. Using a longer context length is more effective, suggesting that the Transformer is able to leverage past state information.

VI-E3 Dataset size

In Fig. 11(c), we train Locomotion-Transformer models with 1%, 10%, and 100% of the training data. While 1% (∼\sim176 episodes) is insufficient for training a competitive policy, 10% (∼\sim1760 episodes) and 100% result in similar performance. We note that using a larger or more diverse dataset may lead to improved sim-to-real performance, which we leave as investigation for future work.

VI-E4 Model size

Fig. 11(d) shows that increasing the size of the Locomotion-Transformer model can lead to improved performance, which mirrors similar results in other domains such as language. However, deploying these models on real robots places strict bounds on inference latency and therefore model size.

VII Conclusion

We present Barkour, a benchmark to evaluate the agility of quadruped robots. Inspired by dog agility competitions, Barkour is a testbed with an intuitive scoring mechanism that requires combination of various agile skills. Furthermore, it can be easily adapted to robots of various sizes or extended by adding or rearranging obstacles while retaining the same metrics.

To set a strong baseline and make progress towards the Barkour benchmark, we explore a learning-based sim-to-real approach and proposed two baseline solutions. In both solutions, we first train a set of specialist policies that excel at tackling individual tasks by leveraging recent developments in fast simulation techniques with on-policy RL algorithms. While in the first solution, we manually design a state machine to switch between different specialist policies. In the second solution, we distill these skills into a generalist Transformer-based policy, named Locomotion-Transformer, that can automatically and smoothly transition between different obstacles. Our results show that Locomotion-Transformer can exhibit multiple agile skills required by the proposed benchmark and can automatically switch between them depending on the sensed environment and the commands received from a high-level navigation controller. As a validation of our approach, we also demonstrate that the policies can generalize to other obstacle configurations.

VIII Limitations and Future Work

Our proposed approach sets a strong baseline on the benchmark. However, as the scores reflect, Barkour is not fully solved and there is still notable room to push towards dog-level agility by improving speed and robustness. We believe that bridging this gap necessitates a collective endeavor from the research community, and the suggested Barkour benchmark can help effectively track this progress.

One limitation of the current proposed baseline methods is that we use privileged information such as the CAD model of the environment and the position of the robot (via a Motion-Capture system) in the world frame. An important future work direction is to explore Barkour using only on-board sensors for both low-level locomotion skills and high-level navigation controller.

An equally exciting direction for future research on Barkour is to evaluate the impact of modifications to robot hardware, different form-factors, and sensors on performance or training speed. Finally, we are also looking into evaluating Barkour in an interactive setting, closer to real-world dog agility competitions, with a human leading a robot through the course.

Acknowledgments

We would like to thank Marissa Giustina, Gus Kouretas, nubby Lee, James Lubin, Sherry Moore, Thinh Nguyen, Krista Reymann, Satoshi Kataoka, Trish Blazina and the rest of the robotics team at Google DeepMind for their feedback and contributions.

References

-A Author Contributions

Methods (conception, benchmark definition, architecture development, implementation, policy training, ablations…) Ken Caluwaerts, Atil Iscen, J. Chase Kew, Kuang-Huei Lee, Lisa Lee, Wenhao Yu, Tingnan Zhang, Vincent Zhuang.

Software infrastructure Ken Caluwaerts, Erwin Coumans, Daniel Freeman, Atil Iscen, J. Chase Kew, Yuheng Kuang, Lisa Lee, Ofir Nachum, Ken Oslund, Francesco Romano, Wenhao Yu, Tingnan Zhang, Daniel Zheng.

Hardware development Ken Caluwaerts, Jason Powell, Stefano Saliceti, Jeff Seto, Ron Sloat, Daniel Zheng.

Leadership (managed or advised on the project) Ken Caluwaerts, Adil Dostmohamed, Raia Hadsell, Nicolas Heess, Atil Iscen, Bauyrjan Jyenis, Michael Neunert, Francesco Nori, Carolina Parada, Stefano Saliceti, Jie Tan, Vikas Sindhwani, Vincent Vanhoucke, Daniel Zheng.

Paper (writing, figures, visualizations) Ken Caluwaerts, Daniel Freeman, Nicolas Heess, Atil Iscen, J. Chase Kew, Kuang-Huei Lee, Lisa Lee, Stefano Saliceti, Jie Tan, Wenhao Yu, Tingnan Zhang, Vincent Zhuang.

Operations (data collection, hardware maintenance) Omar Cortes, Linda Luu, Jason Powell, Diego Reyes, Ron Sloat.

-B Barkour Definitions

-B2 Weave Poles

-B3 A-frame

To successfully complete the A-frame obstacle, the robot must pass (center of torso) across the line segment defined by the part of the A-frame touching floor on the first side, then move across the top of the A-frame and finally across the line segment at the bottom of the opposite end.

The A-frame is challenging because the rather small feet of the robot do not provide enough friction for semi-static behaviors. The lack of friction also penalizes jittery motions that lose contact with the ground. Moreover, the sensitivity to friction and restitution of the surface and the feet deformations also make the problem harder to model perfectly in simulation. The A-frame requires the robot to approach with some speed and maintain good contact to generate sufficient friction while keeping forward momentum. The robot must also keep its balance and transition to a soft landing during the downhill segment.

-B4 Broad Jump

-B5 Scoring

The score is based on succeeded obstacles and completion time. Timing starts when the robot’s center leaves the start table and stops when it reaches the end table.

-C Barkour Experiment Setup

To be able to restart an experiment without assistance, we incorporate a default policy for walking back to the start table and an automated recovery policy that triggers when the robot has tipped over. Fig. 12 shows two typical examples of the robot successfully recovering from unexpected conditions. In the first example, the robot accidentally rolls over after completing the A-frame. This triggers a brief switch to the recovery policy, after which the Locomotion-Transformer policy repositions the robot towards the broad jump obstacle. In the second example, the robot suboptimally jumps into the air at the broad jump (likely due the cable pulling it at an inopportune time). After landing sideways with the knees hitting the floor, the generalist policy quickly recovers and the robot continues towards the end table.

-D Distilling Hardware Data into a Course-Specific Locomotion-Transformer

As indicated in the main text, training a Locomotion-Transformer policy based on expert policies in simulation raises a couple questions around the flexibility of the method:

Is the Transformer-based framework also capable of learning high-level behaviors?

Is it possible to train a Locomotion-Transformer policy from a small hardware dataset?

To explore these topics, we collect a small dataset (64 episodes, 112K samples total) of Barkour runs that successfully completed the whole course from Fig. 2 on hardware using the navigation controller and specialist policies or generalist Locomotion-Transformer policies. Next, we use the distillation pipeline but remove the navigation controller and directly feed the robot’s position and orientation to the Locomotion-Transformer (Fig. 13). In other words, instead of a generalist policy that accepts a broad range of commands from a navigation controller, we train a course-specific policy. More precisely, the course-specific policy directly maps the robot’s position, orientation, proprioceptive sensors (IMU, joints), and heightfield to motor commands.

Finally, we qualitatively verified the robustness of course-specific Locomotion-Transformer policy by placing the robot near but not on the nominal trajectory for the course. As shown in Fig. 15, the policy is able walk towards the nominal trajectory and complete the course. In some cases (Fig. 15 right plot, farther from nominal trajectory), the robot stands still until it receives a small push, after which it completes the course.

-E OWP Observation Space and Training

-F SCP Terrain Curriculum and Training

-G JP Training Curriculum

-H Details on Reward function for Specialist Policy Training

The specialist policy and value models consist of separate fully-connected MLPs each with four layers of sizes 1024, 512, 256, and 128. Each hidden layer is followed by an ELU activation and all inputs are flattened and fed directly into the network.

-I Details on Navigation Controller

The navigation controller sends linear and angular velocity commands to the policy based on the next waypoint’s relative position and target yaw angle:

where α\alpha and β\beta are coefficients for frontal linear velocity and sideways linear velocity, ∥d∥\lVert\mathbf{d}\rVert is the distance between the center of the robot’s torso and the next waypoint, ϕ\phi is the angle between the vectors from the robot to the next waypoint and the robot’s heading, γ\gamma is the coefficient for angular velocity and Δ(ψ)\Delta(\psi) is the difference between robot’s heading angle and the waypoint’s orientation. Lastly, vxv_{x}, vyv_{y}, and ω\omega are clipped based on the policies’ capabilities (depending on how fast policies are trained to walk and turn).

-J Details on Locomotion-Transformer Data Collection

-K Details on Locomotion-Transformer Architecture

As described in Section IV-B, Locomotion-Transformer is a causal Transformer , following the standard GPT implementation . We have two layers and the context length is 15. The transformer input token and the internal representations are 256-dd. For tokenizing elevation maps, we use one convolutional encoder to map each elevation map into a 64-dd embedding vector, concatenate all embedding vectors, and project the concatenated vector into a 256-dd token. All the convolutional encoders have two convolution layers with kernel size 3, channel size 8, and stride size 1, followed by a 64−d-d projection layer. Proprioceptive states and velocity commands are tokenized together with a 256-dd projection following a concatenation. Actions are also tokenized with a 256-dd projection. All tokenizer hidden layers are followed by ReLU activation.

-L Accelerating Training with Brax

Because of the difference in implementation details between the various physics engines used, we designed a simple protocol to align the actuation parameters between the Brax robot model and the other (sim and real) models. We placed the robot on a raised platform and recorded a trace of its joint angles when actuating to a target position from a reference neutral pose. We then differentiably optimized the joint gain and damping parameters until the joint trajectories had converged.

-L2 Training Environments

We followed the environment setup outlined in Sec. IV, sharing the observation space, actuation model, reward functions, and domain randomization scheme. However, we did not use a curriculum, and instead uniformly sampled over environment variability.

We provide a reference training curve of the omnidirectional walking policy in Fig. 17 and validated the resulting policies on hardware.

For the A-Frame environment, we trained policies for traversing the A-Frame as well as sloped terrain variants for different slope angle/friction combinations. The behavior of these policies tended to depend sensitively on the distribution of slopes and frictions being randomized over. These sloped-traversal policies typically performed well in sim, but had trouble transferring to the real, rugged A-frame obstacle when deployed to the robot.

We explored additional variants of locomotion rewards, including the “Fetch” reward from Brax’s default set of environments , as well as various penalization and reward schemes for encouraging more graceful motions. Notably, Fetch worked out of the box to produce a gait capable of locomotion (without domain randomization), but the motion quality and strain on the robot were too jittery for practical use.

We also explored alternative environment definitions, including different observation space features, and different bread-crumbing mechanisms for the target, but ultimately settled on the scheme described in the main text.

-L3 Validation

We evaluated cross-engine correspondence fidelity by qualitatively verifying IsaacGym policies running in Brax, Brax-learned policies running in PyBullet, and Brax policies running on the robot (see, e.g., Fig. 18). Reaching qualitatively reasonable transfer quality required domain randomization as well as motor-level gain and damping coefficient optimization. We defer a more thorough analysis of sim-to-real transfer between Brax and the robot to future work. At a high level, more aggressive domain randomization and larger observation noise typically resulted in better transfer at the expense of more conservative gaits.

-L4 Conclusions and Future Work

Brax proved to be a fast iteration platform, supporting near-interactive reward shaping experiments and seamless scaling to large numbers of accelerators. In particular, we note Brax’s efficiency to facilitate high-throughput hyperparameter sweeps for effective reward scaling coefficient settings (Fig. 19). The primary obstacles to its uptake were largely related to development velocity, feature completeness, and motion quality: existing, functioning pipelines needed to be duplicated within Brax, Brax did not always have the necessary features out-of-the-box that already existed in, e.g., IsaacGym, and Brax policies simply weren’t as smooth, owing to less overall researcher-time spent refining them. Nonetheless, Brax remains a compelling platform for sim-to-real robotics research. The explorations described above were based on Brax v1 as the next generation of Brax (v2) was not yet available at the time of writing.