Learning to Jump from Pixels

Gabriel B. Margolis, Tao Chen, Kartik Paigwar, Xiang Fu, Donghyun Kim, Sangbae Kim, Pulkit Agrawal

Introduction

One of the grand challenges in robotics is to construct legged systems that can successfully navigate novel and complex landscapes. Recent work has made impressive strides toward the blind traversal of a wide diversity of natural and man-made terrains . Blind walkers primarily rely on proprioception and robust control schemes to achieve sturdy locomotion in challenging conditions including snow, thick vegetation, and slippery mud. The downside of blindness is the inability to execute motions that anticipate the land surface in front of the robot. This is especially prohibitive on terrains with significant elevation discontinuities. For instance, crossing a wide gap requires the robot to jump, which cannot be initiated without knowing where and how wide the gap is. Without vision, even the most robust system would either step in the gap and fall or otherwise treat the gap as an obstacle and stop. This inability to plan results in conservative behavior that is unable to achieve the energy efficiency or the speed afforded by advanced hardware.

State-of-the-art vision-based legged locomotion systems can traverse discontinuous terrain by walking across gaps and climbing over stairs. However, often simplifying assumptions are made in the control scheme such as fixed body trajectory , statically stable gait , or restricted contact pattern . These assumptions result in conservative and non-agile locomotion. For instance, such systems can walk across small gaps, but cannot jump across big ones.

Planning agile behaviors, such as jumps, on discontinuous terrain offers a different and complementary challenge to traversing continuously uneven terrain. Executing a jump requires planning the location of the jump, the force required to lift the body, and dealing with severe under-actuation during the flight phase. Past work has demonstrated standing jumps in simulation , on a real robot , and running jumps in simulation . The most relevant to our work is the demonstration of MIT Cheetah 2 running and jumping over a single obstacle . However, this system was heavily hand-engineered: it assumes straight-line motion, uses a specialized control scheme developed for four manually segmented phases of the jump, and employs a specialized vision system for detecting specific obstacles. Further, the robot was constrained to a fixed gait. Consequently, this system is specific to jumping over one obstacle type, and substantial engineering effort would be required to extend agile locomotion to diverse terrains in the wild.

Traversing discontinuous terrains in more general settings requires a system architecture that can automatically produce a diverse set of agile behaviors from visual observations. To study this problem, we constructed a gap-world environment containing flat regions and randomly placed variable-width gaps. While these environments are much simpler than “in-the-wild”, traversing them successfully requires solving many of the core challenges in vision-guided agile locomotion.

Our proposed method, Depth-based Impulse Control (DIC), employs a hierarchical scheme where a high-level controller processes visual inputs to produce a trajectory of the robot’s body and a “blind” low-level controller ensures that the predicted trajectory is tracked. This separation eases the task for both the controllers: the high-level is shielded from intricacies of joint-level actuation and the low-level is not required to reason about visual observations, allowing us to easily leverage advances in blind locomotion. Instead of using low-level controllers that track robot’s center of mass, a scheme typically known as whole-body control (WBC) , we make use of a whole-body impulse controller (WBIC) that reasons about impulses and is therefore appropriate for dynamic locomotion such as jumps. Model-free deep reinforcement learning is used to train the high-level controller that predicts the commands for WBIC from depth images captured from an on-board camera in real-time. We first train our agents in simulation and then transfer them to the real world using the MIT Mini Cheetah robotic platform (Figure 1).

Our overall contribution is a system architecture that enables the robot to: (a) cross a sequence of wide gaps in real-time using depth observations from a body-mounted camera in the real world; (b) requires no dynamics randomization for sim-to-real transfer; (c) does not assume fixed gait and results in emergence of different gaits as a function of robot velocity and task complexity; (d) achieves the theoretical limit of jump width with fixed gaits and even wider jumps with variable gaits and (e) outperforms prior work by making better use of the full range of agile motion afforded by the hardware.

Method

Our approach, DIC, is guided by the intuition that a wide range of agile behaviors can be generated by using an adaptive gait schedule and commanding the body velocity of the quadruped. A high forward velocity results in running, whereas different ratios of vertical and forward velocity can control the height and the span of a jump. The adaptive gait schedule allows the robot to change when its foot contacts the ground and thus further expands the range of feasible contact locations and applied forces. As shown in Figure 2, we solve the problem of mapping depth observations to velocity and gait-schedule commands by training a high-level trajectory generator (Section 2.1) with model-free deep reinforcement learning (Section 2.3).

To ensure that the robot tracks these commands, one possibility is to simultaneously train a low-level controller using RL that converts the high-level velocity and gait commands into joint torques. Such a scheme has two drawbacks: (i) sim-to-real transfer issues and (ii) large data requirement for training. Another possibility is to leverage an analytical model of the robot and solve for joint torques using trajectory optimization – a scheme commonly known as whole-body control (WBC) . One issue, however, is that a typical WBC tracks the robot’s center-of-mass (CoM) , which is infeasible during the flight phase of agile motion due to under-actuation of the robot’s body. To overcome this issue, we leverage a prior control scheme built on the intuition that changes in body velocity can be realized by modifying the forces applied by the robot’s feet on the ground. This frees the controller from the requirement of faithfully tracking the CoM and instead tracks the contact timing and the ground forces applied by the feet. This approach, called whole-body impulse control (WBIC) , enables tracking of highly dynamic trajectories set by the high-level controller (see Section 2.2). Our proposed method, Depth-based Impulse Control, integrates WBIC with a vision-aware neural network (Figure 2).

Whole-body State The robot’s whole-body state at time tt is fully defined as

Rollout Procedure The iterative execution routine for our high-level policy and an analytical model-based low-level controller is given by Algorithm 1. The high-level policy πθ\pi_{\theta} (Section 2.1) selects action at\textbf{a}_{t}, which the whole-body trajectory generator (WTG; Section 2.2) converts to target whole-body trajectory Xt:t+HdesX^{des}_{t:t+H}. The low-level controller tracks the whole-body trajectory over horizon HH by regulating contact forces. In our experiments, H=10H=10 and the MPC and high-level policy timesteps are 0.0360.036s.

Let the high-level policy be at=πθ(st,ot,at−1)\textbf{a}_{t}=\pi_{\theta}(\textbf{s}_{t},\textbf{o}_{t},\textbf{a}_{t-1}) where at\textbf{a}_{t} is the action and st,ot\textbf{s}_{t},\textbf{o}_{t} denote the robot’s internal state and the terrain observation respectively. The action at previous time-step is fed as input to encourage the predicted actions to change smoothly. π\pi is represented using a neural network.

With fixed gait, the robot’s desired contact state is a cyclic function of time and does not depend on the high-level controller. For instance, the contact schedules for trot and pronk gaits correspond to:

where dd is the gait cycle duration. In our experiments with fixed gaits, we set d=10d=10. In this scenario, π\pi only sets the robot’s velocity.

For variable gait, the high-level action space is expanded to predict one of the two possible contact states of the feet (atc∈\textbf{a}_{t}^{c}\in). In our setup, variable pronk corresponds to choosing one of these states at every time step:

We can further relax the assumption about the gait and let the policy choose the contact state for each foot independently (atc∈4\textbf{a}_{t}^{c}\in^{4}) at every time step. We call this unconstrained gait, where:

determines the contact state of each foot. Flexibility in the contact state allows for emergence of terrain dependent agile gaits.

2 Low-Level Controller

The Whole-body Trajectory Generator (WTG) converts action at\textbf{a}_{t} into an extension of the desired whole-body trajectory at time t+Ht+H, denoted as

where the action is converted to a velocity command as pb˙(at)=[atx˙,aty˙,atz˙,α˙=0,β˙=0,atγ˙]\dot{\textbf{p}_{\text{b}}}(\textbf{a}_{t})=[\textbf{a}_{t}^{\dot{x}},\textbf{a}_{t}^{\dot{y}},\textbf{a}_{t}^{\dot{z}},\dot{\alpha}=0,\dot{\beta}=0,\textbf{a}_{t}^{\dot{\gamma}}], from which pb(at)\textbf{p}_{\text{b}}(\textbf{a}_{t}) and pb¨(at)\ddot{\textbf{p}_{\text{b}}}(\textbf{a}_{t}) are fixed for consistency with the previous target Xt+H−1desX^{des}_{t+H-1} assuming linear interpolation between timesteps. The generator computes foot position targets pfraibert,pf˙raibert,pf¨raibert\textbf{p}_{\text{f}}^{\text{raibert}},\dot{\textbf{p}_{\text{f}}}^{\text{raibert}},\ddot{\textbf{p}_{\text{f}}}^{\text{raibert}} such that the contact locations satisfy the Raibert Heuristic (Section B.1) and swing trajectories are represented as three-point Bezier curves.

Whole-body Trajectory Tracking operates at high frequency with no direct access to terrain information. It consists of a hierarchy of three controllers described in and summarized below:

A Model Predictive Controller (MPC) solves a convex program fdes=MPC(Xt:t+Hdes,Xt)\textbf{f}^{\hskip 1.42271ptdes}=\text{MPC}(X^{des}_{t:t+H},X_{t}) to convert the desired whole-body trajectory Xt:t+HdesX^{des}_{t:t+H} and current whole-body state XtX_{t} into target ground reaction forces fdes\textbf{f}^{\hskip 1.42271ptdes} for each foot at each timestep. MPC operates at 40 Hz.

A Whole-Body Impulse Controller (WBIC) applies differential inverse kinematics qdes,q˙des,τdes=WBIC(Xtdes,Xt,fdes)\textbf{q}_{des},\dot{\textbf{q}}_{des},\tau_{des}=\text{WBIC}(X^{des}_{t},X_{t},\textbf{f}^{\hskip 1.42271ptdes}) to find the target position qdes\textbf{q}_{des}, velocity q˙des\dot{\textbf{q}}_{des}, and feedforward torque commands τdes\tau_{des} for all joints to optimally track the current step of the the whole-body trajectory XtdesX^{des}_{t} and desired ground reaction forces fdes\textbf{f}^{\hskip 1.42271ptdes}. WBIC operates at 500 Hz.

A Proportional-Derivative Plus Feedforward Torque Controller takes as input a target position qdes\textbf{q}_{des}, target velocity q˙des\dot{\textbf{q}}_{des}, and feedforward torque command τdes\tau_{des} as well as the current position and velocity for each joint. It computes an output torque for each motor at 40 kHz.

3 Neural Network Training

Network Architecture The high-level policy πθ(at∣st,ot,at−1)\pi_{\theta}(\textbf{a}_{t}|\textbf{s}_{t},\textbf{o}_{t},\textbf{a}_{t-1}) is modeled using a deep recurrent neural network that includes a convolutional neural network (CNN) for processing the raw terrain observation ot\textbf{o}_{t}. The output features of CNN are concatenated with proprioceptive inputs st\textbf{s}_{t}, previous action at−1\textbf{a}_{t-1}, and a cyclic timing parameter and passed through a sequence of fully connected layers to output a probability distribution over at\textbf{a}_{t}. Figure 3 illustrates the architecture of the policy network.

Initialization and Termination For each training episode, the robot is initialized in a standing pose on flat ground. The locations of gaps and their widths are randomized. An episode terminates if any of three terminal conditions are met: (1) the body height is less than 20 centimeters; (2) body roll or pitch exceeds 0.70.7 radians; or (3) a foot is placed in a gap. The maximum episode length is 500 steps, equivalent to 25 seconds of simulated locomotion.

Reward Function The reward rtr_{t} at time tt is defined as:

The first term rewards forward progress pt,xb−pt−1,xbp^{b}_{t,x}-p^{b}_{t-1,x}, where pt,xbp^{b}_{t,x} is the projection of the body frame position at time tt onto the xx-axis in the world frame. The second term applies a soft safety constraint by penalizing when the body velocity vtbv^{b}_{t} exceeds VthreshV_{thresh}. The third, fourth, and fifth terms incentivize stability by penalizing the roll, pitch, and yaw of the body, denoted as αtb,βtb\alpha^{b}_{t},\beta^{b}_{t}, γtb\gamma^{b}_{t} . The sixth term rewards smooth motion by minimizing q˙\dot{q}. In training with variable and unconstrained gaits, we found this term critical to promote exploration of lower-frequency gaits. The parameters in the reward term are set to: c1=1.0,c2=0.5,c3=0.02,c4=0.05,c5=0.15,c6=0.03,Vthresh=1.0c_{1}=1.0,c_{2}=0.5,c_{3}=0.02,c_{4}=0.05,c_{5}=0.15,c_{6}=0.03,V_{thresh}=1.0m/s.

Policy Optimization The parameters of the neural network (θ\theta) are optimized using the PPO algorithm, Adam optimizer with learning rate 0.00030.0003 and batch size 256256. During training, 3232 environments are simulated in parallel. We find that policies converge within 60006000 training episodes, equivalent to 6060 hours of simulated locomotion or 12 hours of computation.

Asymmetric-Information Behavioral Cloning Learning directly from depth images presents two challenges: (1) Partial observations: a front-facing depth camera can only provide information about the terrain in front of the robot, not the terrain underneath its feet, making the contact-relevant terrain partially observed. (2) Sensory variance: the depth image obtained from a body-mounted camera is dependent on the robot pose. This introduces variance in perception across trials, even when the robot is traversing the same terrain.

Variance makes learning more challenging, and partial observations necessitate the use of a recurrent network architecture. These factors make learning directly from depth images less sample-efficient than learning from heightmaps. In addition, rendering depth images is more computationally expensive than cropping heightmaps, which makes learning from depth images less wall-clock efficient.

Experimental Setup

Hardware: We use the MIT Mini Cheetah , a 9kg electrically-actuated quadruped that stands 28cm tall with a body length of 38cm. A front-mounted Intel RealSense D435 camera provides real-time stereo depth data and an onboard computer run the trajectory-tracking controller described in Section 2.2. Data from the depth camera is processed by an offboard computer that communicates the output of the high-level policy to the robot via an Ethernet cable.

Simulator: We train high-level policy using the PyBullet simulator. To obtain data from the mounted depth camera, we use a CAD model of our robot and sensor’s known intrinsic parameters.

Gap World Environment: To evaluate the ability of our system to dynamically traverse discontinuous terrains, we define a test environment consisting of variable-width gaps and flat regions. The difficulty of traversing gap worlds depends on the proximity of gaps as well as gap width, with closer and wider gaps presenting a greater challenge to the controller. Our training dataset consists of randomly generated gaps with uniform random width between Wmin=4W_{\text{min}}=4 and Wmax∈W_{\text{max}}\in centimeters, separated by flat segments of randomized width 0.50.5 to 2.02.0 meters. Our test dataset contains novel terrains drawn from the same distribution.

Baselines: We compare our method to a model-free baseline, Policies Modulating Trajectory Generators, and a model-based baseline, Local Foothold Adaptation. For details of these baselines, refer to Appendix C, D.

Results

Fixed Gait We train Depth-based Impulse Control to cross gaps using trotting and pronking gaits. For both trotting and pronking, our visually-guided approach succeeds at above 90% of gap crossing attempts up to the theoretical limits derived in the supplementary material (Section B). Figure 4(a) reports the performance of our method relative to this theoretical limit. Ideal performance is derived from maximum stride length given velocity, foot placement, and contact schedule constraints. Note that while the theoretical limits are derived assuming zero yaw, the learned trotting controller learns to move with nonzero yaw, thus extending the foot placements further apart and beating the ideal. Our method also outperforms blind locomotion (Figure 4(a)) and a Local Foothold Adaptation baseline (Figure C.2), particularly on large gaps.

Unconstrained Gait We relax all constraints on contact schedule and train a controller with a vision-adaptive contact schedule to cross wide gaps. Figure 4(b) reports the performance of unconstrained gait gap crossing in simulation. Unconstrained gait policies outperform those with fixed gait, crossing gaps that are much wider. When trained with extremely wide (40- to 70-cm gaps), DIC learns to select a variable-bounding contact schedule which achieves superior performance to trotting and pronking for very large gaps (Figure 6). When we restrict the maximum gap size to 40cm or less, a variable-timing pronking gait emerges in the unconstrained gait controller. Figure 5 illustrates the variable contact timings and velocity modulation of the variable pronking controller in simulation. Similar to concurrent work which has demonstrated the emergence of variable gaits for energy minimization on flat ground; we observe emergent gait adaptation for safe traversal of discontinuous terrain.

Ease of Training Our method successfully navigates gaps of different width with different gaits using the same reward function and trajectory generator structure. In contrast, we found that the PMTG baseline was highly sensitive to the tuning of the reward and trajectory generator for each gait and environment. We first tuned the trajectory generator, residual magnitudes, and reward function of PMTG for sim-to-real forward locomotion on flat ground; details and video of the baseline can be found at the project website\refwebsite{}^{\ref{website}}. We found the parameters that succeeded at sim-to-real on flat ground were prohibitively conservative and failed to learn any gap-crossing behavior when the maximum gap width WmaxW_{\text{max}} was 10cm for trotting or 20cm for pronking. To overcome this issue, we applied specialized reward design and expanded the range of the trajectory generator parameters. While the re-tuned agent was able to cross gaps longer than the aforementioned range, the resulting behaviors overrode the TG with irregular gaits indicative of simulator exploitation.

2 Real World Performance

Deployment We deploy DIC in fully real-time fashion on the MIT Mini Cheetah robot , directly making use of depth images and an onboard state estimator. In this setting, we record successful gap crossings up to 16cm. We refer the reader to the project website for video evaluation\refwebsite{}^{\ref{website}}.

To study the impact of sensor noise on transfer, we also deploy DIC using ground-truth state information via motion capture and terrain heightmap. With these adjustments, we are able to consistently cross gaps up to 26cm on the real robot. Figure 7 plots motion capture data from three such deployments each for adaptive trotting (left) and adaptive pronking (right). The relevant cross-section of the terrain surface is drawn in dark green. Although the foot placements of the robot differ across runs due to noise in the system dynamics, DIC adapts to avoid stepping in a gap in each case.

From these experiments, we identify two main challenges which prevent our method from transferring for wider gaps: (i) drift in state estimation caused by sensor noise and imprecise knowledge of contact timing; (ii) violation of the assumption made by the low-level controller that the robot’s feet do not slip while in contact with the floor, especially during aggressive motion. We refer the reader to the project website for video of example failure cases\refwebsite{}^{\ref{website}}.

3 Vision and Behavioral Cloning

Behavioral Cloning (BC) Table 1 illustrates that behavioral cloning from heightmaps to depth images offers an advantage over learning directly from depth images in most cases after 10M training steps and 1M cloning steps. We note that cropping heightmaps is faster than rendering depth images, resulting in an additional wall-clock time benefit to BC. These results also demonstrate that the combination of behavioral cloning with a variable gait schedule is beneficial, with the cloned Variable Pronk achieving highest performance for wide gaps of any fixed or variable gait policy.

Recurrent Architecture We find that student policies with recurrent architecture consistently yield higher final performance than without, particularly for environments with larger gaps which require more dynamic motion (Table 1). This suggests that the hidden state is helpful in forming a useful representation of unobserved terrain regions given the observation history.

Related Work

Model-free RL for locomotion is shown to benefit from acting over low-level control loops rather than raw commands . Robust walking methods including RMA as well as recent work on ANYmal and Cassie learn conservative, vision-free policies to predict joint position targets for a PD controller and achieve sim-to-real transfer using a combination of reward shaping, system identification, domain randomization, and asymmetric-information behavioral cloning. Previous work in simulation has applied model-free reinforcement learning to traversal of discontinuous terrains in simulation. notably applied model-free RL to the problem of crossing stepping stones with physically simulated characters, but this method did not use realistic perception or take measures to promote sim-to-real transfer.

Model-based control for locomotion has achieved highly dynamic blind walking , running , and jumping over obstacles using known quadruped whole-body and centroidal dynamics. Other works have applied model-based control to terrain-aware navigation of a mapped environment, typically with complete information about the terrain . In general, control strategies based on known models are high-performing and robust where the state is known and the model is sufficiently accurate. In contrast, model-free controllers excel at incorporating unstructured or partially observed state information when large data is available.

Interfacing Model-based and Model-Free Methods. A previous line of work has leveraged model-free perception for foothold selection. locally adapted foot placements to safe footholds predicted by a CNN. RLOC similarly uses a learning-based online footstep planner in combination with a learning-modulated whole-body controller to perform terrain-aware locomotion. Unlike our method, uses a complete terrain heightmap as observation, plans by targeting foot placements, and is limited to relatively conservative fixed walking and slow trotting gaits. On the other hand, concurrent work applies RL to modulate a model-based controller’s target command without perception. demonstrated that using a model-free policy to choose contact schedules for a reduced-order model leads to the emergence of efficient gait transitions during blind flat-ground locomotion. demonstrates the integration of a model-free high-level controller with a centroidal dynamics model. This framework deployed with a fixed trotting gait is demonstrated to achieve flat-ground and conservative terrain-aware locomotion. Unlike our work, does not demonstrate gaits with flight phases or plan from realistic terrain observations.

Conclusion and Discussion

We have presented a vision-based hierarchical control framework capable of traversing discontinuous terrain with gaps. The combination of model-free high-level trajectory prediction and model-based low-level trajectory tracking enables us to simultaneously achieve high performance and robustness.

While our system advances the state-of-the-art, there are many avenues for improvement. First, while we are able to train policies that can jump gaps as long as 66 centimeters in simulation, we can only transfer to gaps up to 26 centimeters in the real world. We identify a few obstacles to transfer in Section 4.2 and further note the limitation that the contact force optimizer does not account for robot’s kinematic configuration, sometimes resulting in infeasible or overly conservative target impulses.

Another challenge in deploying our system in the wild is that the small onboard computer on the Mini Cheetah robot does not have space to run the neural network alongside the low-level controller. We are in the process of upgrading the onboard computer. Finally, while we only present results in the gap-world environment, it should be possible to use this method on combinations of rough continuous terrains and additional classes of discontinuous terrains such as stairs. We leave such experimentation to future work.

The authors acknowledge support from the DARPA Machine Common Sense Program. This work is supported in part by the MIT Biomimetic Robotics Laboratory and NAVER LABS. We are grateful to Elijah Stanger-Jones for his support in working with the robot hardware and electronics. The authors also acknowledge MIT SuperCloud and the Lincoln Laboratory Supercomputing Center for providing HPC services.

References

Appendix A Details of Low-Level Controller

Complete information about the Model-Predictive Controller and Whole-Body Impulse Controller used in this work may be found in . A simplified description of the control loop is given in Algorithm 2. The convex Model-Predictive Controller solves for target ground reaction forces over the planning horizon given the target whole-body trajectory and current state. The whole-body impulse controller then tracks these ground reaction forces at high frequency.

Appendix B Derivation of Theoretical Gap-Crossing Limits

Given the velocity of the base and the duration of the next contact, the Raibert Heuristic selects foot placements such that each leg’s lever angle of incidence on the ground is equal to its angle of departure. This foot placement follows the formula

where Δtdli\Delta t_{d}^{l_{i}} is the duration of the next placement of foot ii, vv is the estimated robot body velocity, vcmdv^{\rm cmd} is the commanded robot velocity, and kk is a tunable gain term.

B.2 Theoretical Limit on Fixed-Gait Gap Crossing

A quadruped stepping at a fixed cycle frequency ff and moving at velocity v=vcmdv=v^{cmd} will locomote a distance of vf\frac{v}{f} each gait cycle, and so the nominal foot placement for any given foot under the Raibert heuristic will advance by a distance of vf\frac{v}{f}. In the pronking gait, wherein all legs contact the ground simultaneously, this places the upper limit on gap crossing at vf\frac{v}{f}. For a trotting gait, wherein pairs of diagonal legs meet the ground in alternating timing, this limit is reduced by half to v2f\frac{v}{2f}. The Mini Cheetah nominally trots at a frequency of f=3f=3Hz, theoretically limiting its maximal gap crossing at a velocity of v=1v=1m/s to 3333cm in the pronking case, or 1717cm in the trotting case. Table B.1 lists derived limits for a few example gait frequencies, body velocities, and gaits. In practice, the popular ANYmal C by ANYbiotics, over twice as long and 5 times as massive as the Mini Cheetah, is rated by its manufacturers to cross gaps of up to 2525cm.

B.3 Upper Bound on Gap Crossing Probability for Fixed Gait

Adhering to the Raibert heuristic with constant velocity and gait yields a distance dd between footsteps of d=v2fd=\frac{v}{2f} for trotting and d=vfd=\frac{v}{f} for pronking, as the analysis above shows. If we assume that a gap of width hh is randomly positioned in the robot’s path, the probability that any single foot steps into the gap is hd\frac{h}{d}. The probability that a single foot avoids the gap is then 1−hd1-\frac{h}{d}. The probability that none of the four feet step into the gap is less than or equal to the probability that any one foot avoids the gap. Thus, the probability of avoiding a randomly placed gap of width hh for a blind controller at constant velocity is upper bounded at 1−2fhv1-\frac{2fh}{v} for trotting and 1−fhv1-\frac{fh}{v} for pronking. Intuitively, the upper bound gap crossing probability decays linearly from one for a gap of width to zero for a gap of width dd.

B.4 Upper Bound on Gap Crossing Probability with Local Foothold Adaptation

The baseline controller of performs locomotion with fixed gait at fixed velocity. However, if a foothold is detected to be unsafe, a local grid search is performed for the nearest safe foothold within some maximum displacement. Assume the maximum displacement Δ\Delta. Then, the probability of failure to cross a gap is the probability that a foot steps in the gap at least distance Δ\Delta from the edge. This results in single foot failure probability of h−2Δd\frac{h-2\Delta}{d} for gap of width hh and distance between footsteps dd. The probability of avoiding a randomly placed gap of width hh with maximum foot adaptation Δ\Delta is therefore upper bounded at 1−2f(h−2Δ)v1-\frac{2f(h-2\Delta)}{v} for trotting and 1−2f(h−2Δ)v1-\frac{2f(h-2\Delta)}{v} for pronking. Intuitively, the upper bound gap crossing probability decays linearly from one for a gap of width 2Δ2\Delta to zero for a gap of width d+2Δd+2\Delta.

Appendix C Local Foothold Adaptation Baseline

Local Foothold Adaptation Baseline commands a constant velocity and contact pattern to a whole-body impulse controller, and adjusts foot placement locations locally by applying a safety heuristic to a terrain heightmap and searching for the nearest safe location to the nominal foothold. In our evaluation, we assume that this method has privileged access to the true terrain heightmap.

Figure C.2 presents a comparison between the performance achieved by our method (Jumping with Pixels) and the theoretical performance limits derived for rule-based foot placement adaptation (FPA) as described in B.4. Our method outperforms FPA across a range of maximum foot displacement values fmax∈[2cm,4cm,6cm]f_{max}\in[2\text{cm},4\text{cm},6\text{cm}]. Performance improvement is greater for wide gaps. Prior work which implemented FPA on the same robot we use applied a maximum foothold adaptation of 4cm. We additionally note that foot placement adaptation is complementary to the velocity and contact schedule adaptation of our approach. We expect that future work combining these techniques will combine the performance improvement of each.

Appendix D PMTG Baseline

Policies Modulating Trajectory Generators (PMTG) Baseline augments the action space of model-free RL using a parametric trajectory generator (TG) capable of producing cyclic leg motions. Given a timing parameter (tt) that cycles between and 11 and trajectory parameters (a) – stride frequency, length etc., TG outputs joint position targets qdes=TG(t,a)\textbf{q}_{des}=\text{TG}(t,\textbf{a}). The policy also directly predicts residuals (Δqdes\Delta\textbf{q}_{des}). The output command is therefore qdes+Δqdes\textbf{q}_{des}+\Delta\textbf{q}_{des}.

The reward function for training the baseline controller is defined as rt=c1∗rdx+c2∗rvthres+c3∗rr+c4∗rp+c5∗ry+c6∗rGC−c7∗pGAr_{t}=c_{1}*r_{dx}+c_{2}*r_{v_{thres}}+c_{3}*r_{r}+c_{4}*r_{p}+c_{5}*r_{y}+c_{6}*r_{GC}-c_{7}*p_{GA}, where c1,2,3...7c_{1,2,3...7} are the coefficients of each reward terms respectively. The individual terms are defined as follows:

Forward Distance Reward (rdxr_{dx}) : This term maximizes the forward distance moved by the robot in x-direction. rdxr_{dx} = pt,xb−pt−1,xbp^{b}_{t,x}-p^{b}_{t-1,x}

Velocity Threshold Reward (rvthresr_{v_{thres}}): This term penalizes if the robot body velocity ∣∣vtb∣∣2||v^{b}_{t}||_{2} exceeds threshold velocity VthreshV_{thresh}.

Orientation Reward (rr,rp,ryr_{r},r_{p},r_{y}) : This term incentivizes stability of the robot, maintaining zero roll, pitch and yaw.

Gap Crossing Reward (rGCr_{GC}): Unlike , which intensively shape the reward near the gap, we considered a binary reward(0/500) which incentivizes if the robot body successfully reaches the other side of the gap.

Gap Avoidance Penalty (pGAp_{GA}): This term penalizes and terminates the environment when number of gaps crossed is 0 after 180 time-steps.

We found that Gap Crossing Reward (rGCr_{GC}) and Gap Avoidance Penalty (pGAp_{GA}) were critical for baseline policy as opposed to our proposed controller. Figure D.3 shows that the PMTG baseline without these reward terms learns to trot in place and avoids crossing any gaps of width 10cm, while the PMTG baseline with reward shaping is able to demonstrate a gap crossing behaviour. We find that learning with our method of whole-body trajectory modulation achieves higher performance in this environment and converges more rapidly than the baseline.

Appendix E PMTG Sim-to-Real Results

To verify our implementation of PMTG used for baseline comparison, we trained and deployed a simple forward walking policy on flat ground. The results of deployment are illustrated in Figure E.6.

We trained trotting and pronking policies in simulation to cross gaps of different width. We found that parameters such as TG frequency ff and residual Δpx\Delta p_{x} are mainly responsible for exploration over large gaps. We performed an experiment with policies trained over different range of these parameters as follow:

Trot/Pronk Relaxed : Policies trained with Δpx∈[0cm,10cm]\Delta p_{x}\in[0cm,10cm] and f∈[3Hz,5Hz]f\in[3Hz,5Hz]

Trot/Pronk Conservative : Policies trained with Δpx∈[0m,7cm]\Delta p_{x}\in[0m,7cm] and f∈[3Hz,4Hz]f\in[3Hz,4Hz]

Trot/Pronk Conservative (Train on Flat) : Policies trained with Δpx∈[0cm,4cm]\Delta p_{x}\in[0cm,4cm] and f∈[2.5Hz,3.5Hz]f\in[2.5Hz,3.5Hz]. These are the parameter ranges which was used to train a policy deployed on the flat-ground as shown in Fig. E.6

Figure D.4 shows the performance of policies trained with each parameter range. We note that only the relaxed policy is able to learn successful gap crossing behavior for large gaps, particularly for the pronking gait. Figure D.5 illustrates simulated body and foot trajectories of the converged gap-crossing policies trained with each set of TG parameter ranges. We observe that although policies trained with relaxed ranges achieve improved gap-crossing performance, their trajectories include foot dragging, high footswings, and unrealistic velocity changes. The policies with relaxed TG use larger residual commands than those trained with conservative TG, which enables them to discover these unrealistic patterns of simulator exploitation.

Appendix F Depth Image Preprocessing

The depth images produces by the depth sensor used in this work contain imperfections that are not modeled in simulation. To mitigate this, we apply standard preprocessing techniques before passing each depth image to the controller:

Downsampling: For computational efficiency, we downsample each image from its original resolution of 480×360480\times 360 to 160×120160\times 120 using nearest-neighbor interpolation.

Invalid Band Crop: Since our depth sensor is steroscopic, there is an ”invalid band” of variable size on one side of the image. In this band, some objects in the scene are only detected by one camera, preventing their distance from being accurately estimated. In the case of our system, we found that cropping the left side of the image by 20 pixels (after downsampling) avoided the invalid band while retaining sufficient terrain information to perform the task.

Depth Filter: We clip the points in the depth image to the range [0.10.1m, 1.01.0m].

Hole-Filling Filter: We apply a hole-filling filter from the pyrealsense2 package to generally reduce noise artifacts in the image.

Figure F.7 provides example depth images before and after preprocessing.