IEEE Robotics and Automation Letters · 2020 · an animated walkthrough

Perceptive Locomotion in Rough Terrain – Online Foothold Optimization

The paper in 106 seconds · narrated · sound on
Transcript

Meet ANYmal, a four-legged robot from ETH Zurich. It climbs a platform a third as tall as its legs, crosses pallets and loose bricks, and trots up real stairs. But how does it know where to put each foot?

On smooth floors, wheels win. In clutter, legs can go further, but only if every foot lands in a good spot. A blind robot feels the ground, but can't see it. In the paper's stair-gap test, its toe caught the beams, every single time. Think of jogging across stepping stones. You glance ahead, and pick each stone as you go.

ANYmal does the same. Its depth camera builds an elevation map: a grid of ground heights, in 3.5 cm cells. For each step, it starts from where a blind robot would land. Then it checks every cell within 35 cm, and throws out the unsafe ones. The rest are scored: comfortable for the leg, flat, and away from edges. The best cell becomes the foothold, the spot where the foot lands. All four footholds take 1.8 milliseconds on average, about one tick of the controller. And when a leg has to climb high, it gets extra time for the swing.

In simulation, it crossed the stair gaps in all twelve runs. Blind walking failed every time. On the real robot: a 22 cm platform, a slope and bricks that move underfoot, and real stairs. No pre-scanned terrain, and no motion capture. Just the robot's own eyes, fast enough to keep trotting.

Footage: Robotic Systems Lab, ETH Zurich (the paper's official video). Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video. Animations are schematic.

The story in plain words

Researchers at ETH Zurich taught the four-legged robot ANYmal to look at the ground just ahead and pick a safe spot for every foot, fast enough to keep trotting over steps, gaps and loose bricks.

  1. Why this matters

    Legs should beat wheels where the ground is cluttered: stairs, rubble, building sites. Yet in 2020 most legged robots worked on mildly rough ground, where wheels are faster and use less energy. Legs only pay off when every foot is placed with care.

  2. What makes it hard

    When a robot trots, two feet are always in the air, and each must land a fraction of a second later. The robot has to choose those spots from its own camera, on board, while moving. It is like jogging across stepping stones: there is no time to stop and plan.

  3. What people did before

    Early systems found good footholds, but needed the terrain scanned in advance and cameras in the room tracking the robot. Others planned slowly, taking minutes or walking with only one foot lifted at a time. A neural network could correct steps quickly, but it had to be trained first.

  4. What this paper does

    For each foot, the robot takes the spot a blind robot would step on and looks at every 3.5 cm patch of its height map within 35 cm. It throws out unsafe patches, scores the rest (flat? away from edges? comfortable for the leg?) and picks the best. All four feet take 1.8 thousandths of a second.

  5. What they showed

    In simulation it crossed a “stair of gaps” in 12 of 12 runs, where blind walking always failed. The real robot walked a 2-minute obstacle course with a 22 cm platform, a third of its leg length, and climbed real stairs.

  6. Why it's a step forward

    A simple, exhaustive search turned out fast enough to run inside the control loop of a trotting robot, with a map it builds itself. It only plans one step ahead, though, and it doesn't check whether the knees will hit the stairs, which limited it on steep stairs.

Words used below
Elevation map
a grid of small cells, each storing the ground height there
Foothold
the spot where a foot lands
Nominal foothold
where the foot would land if the robot were blind
Trot
a gait where diagonal legs move together, two feet in the air
Swing
the part of a step when the foot is in the air
1 / 7
blind guess height map chosen foothold
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

Where wheels stop, legs can go on

In short: legged robots only beat wheeled ones on messy ground, and only if they can see the ground and choose where every foot goes, while walking.

The paper opens with an honest observation. In 2020, legged robots were mostly used on mildly rough ground, and there their wheeled cousins are faster and more energy-efficient. The case for legs is cluttered terrain, where a robot can adapt where it puts its feet and how high it lifts them. To do that it needs to perceive the terrain and plan footholds from what it sees.

What came before, and where it stopped

  • LittleDog, 2008–2010. Early terrain-aware controllers scored footholds from height maps, one with a hand-made cost, one with a ranking learned from expert examples. Both relied on terrain scanned in advance and on motion-capture cameras in the room, and the first used a slow, statically stable gait.
  • ANYmal, 2018 (the same lab, including three of these authors). This work picked footholds from a map of safe and unsafe spots, but also only with statically stable gaits: one foot moves at a time and the body stays balanced over the other three.
  • Optimizing body and footholds together (HyQ robot). This worked, but took several minutes, so the terrain had to be processed offline. It was shown with a crawling gait.
  • A neural network for foothold corrections (HyQ, 2019). It was fast and continuous, but had to be trained first.
  • Full footstep planners lay out every step to a goal. They take more computation than a fast trot allows.

This paper aims at the gap: pick each foothold locally (one step ahead), from a map the robot builds on the fly, fast enough for a dynamic trot, and without training a network.

Paper Fig. 1. ANYmal on the lab obstacle course. A depth camera at the front of the body maps the ground ahead, and each foot's landing spot is chosen from that map.
Scene 2

Walking blind into a stair of gaps

In short: on a course of narrow beams with gaps between them, a robot that can't see fails every time. The one that picks its footholds from the map gets through every time.

To show why seeing matters, the authors built a test in simulation, inspired by an earlier HyQ experiment: five beams, each 0.12 m deep, spaced 0.24 m apart, each 0.05 m higher than the one before, with the last one widened into a platform. The robot trotted across it going up and going down, at 0.3 and 0.5 m/s, from different starting positions: 12 runs in total, all successful.

The comparison is a blind trot: the same controller, minus everything that uses the map (no foothold search, no timing changes). Going up, it always failed to get past the second beam. Going down, its legs always got stuck between two beams.

The traces also show something the animation only hints at. At 0.3 m/s the robot's natural step is shorter than the beam spacing, so it steps in place on some beams to keep its average speed. At 0.5 m/s the step is longer, so feet sometimes swing over a beam entirely. The authors note that for every course there is a speed at which each foot touches each beam exactly once. Picking that speed automatically is left for future work.

Paper Fig. 5. Climbing up the beams. Top: body height, which rises smoothly. Middle: body speed for the slow (blue) and fast (red) runs. Bottom: the paths of all four feet at 0.3 m/s. Each arc lifts over the next beam (grey) and lands on top of one.
Paper Fig. 6. The same test going down, with foot paths shown at 0.5 m/s. Here the steps are longer than the beam spacing, so some beams are swung over.
Scene 3

A height map built on the move

In short: the robot turns its depth camera's points into a grid of ground heights around itself, many times a second, and keeps that grid lined up with where it thinks it is.

ANYmal carries an Intel RealSense D435 depth camera at the front of its body. Each scan has about 61,624 3-D points and arrives 6 times a second. The points are dropped into a grid map of 1.6 m × 1.6 m around the robot with 3.5 cm × 3.5 cm cells, each holding one height. This is the elevation map. Extra filters sharpen edges, clean up what isn't visible, compute the ground's tilt and fill in patches hidden behind obstacles.

Two practical problems got specific fixes:

  • Drift. The robot's estimate of its own position slowly drifts, mostly in height. An earlier method fused sensor data to correct this, but was too slow for a fast-moving robot. Here a simple drift compensation filter nudges the map height to agree with the latest scan.
  • Its own feet. Walking down, the front feet enter the camera's view and would show up as spikes in the map, so points on the feet are thrown away.
Animated (illustrative). Without correction, a small height error builds up and the stored map floats off the real ground. With the drift filter, every new scan pulls it back. The drift rate shown is made up for clarity.

Speed. On the CPU the map updated only about 2.5 times a second, too slow for a fast-walking robot. So the authors rewrote it for the GPU (using CuPy), making updates 14× faster. On an on-board Jetson AGX Xavier, the whole pipeline takes 28.9 ms on average, from receiving a point cloud to publishing the map. Only 8.4 ms of that is spent on the GPU itself, so the map keeps up with the camera's 6 scans per second.

Paper Fig. 2. The whole walking system. The operator's joystick command and the elevation map (green) feed the motion generator (red): a timetable of steps, the blind-guess footholds, the batch search that refines them, a planner for the body's path, and swing-leg paths. A whole-body controller and joint-level impedance control then drive the robot.
Scene 4

Check every cell, keep the best

In short: for each foot, the robot looks at every map cell near where it would normally step, throws out the unsafe ones, and takes the cheapest of the rest. This is the paper's main contribution.

Step 1: the blind guess. From the commanded speed and the step timetable, the robot predicts where each leg's hip will be when the foot lifts off and when it touches down. The nominal foothold sits halfway between those two points, pushed down by a default leg vector. It is the spot that best matches the robot's preferred leg pose, with no terrain involved.

Animated (illustrative). The nominal foothold is the midpoint of the hip's position at lift-off and its predicted position at touch-down, plus a default leg vector (paper eq. 12).
pnom = ½ · (pthigh, lift-off + pthigh, touch-down) + pthigh→foot  // hip positions predicted from the commanded velocity and the step timing

Step 2: the batch search. Around that guess, the robot visits every cell inside a 35 cm circle, a few hundred cells. Each cell is first checked against hard rules:

  • Max step height. The highest point on the line from the foot's last stance position to the cell must not be too far above the cell or the stance foot. This rejects steps that would push the joints to large torques or near straight-leg poses.
  • Leg collision. Cells closer than a minimum distance to another foot are rejected, because feet tend to collide when side-stepping.

Leg over-extension is deliberately not a hard rule. The hip prediction is too rough, so such a rule would be too conservative. Over-extension is handled later by impedance control. The feasible cells get a cost (next section), and the cheapest wins. Unlike some planners, the search area doesn't need to be split into flat patches first. And once the search area is set, the nominal foothold is ignored, so a guess made “blind” doesn't bias the result.

Why it is fast enough. Every cost term is positive, so a cell can be abandoned as soon as its running total passes the best value found so far. Finding all four footholds takes 1.8 ms on average, at most about twice that. That is close to the 2.5 ms control step. The authors note that the CNN approach from HyQ is up to 12.5× faster, and that run time grows quadratically with grid resolution. If the nominal foothold falls outside the map (the robot outran its map), the nominal foothold is used as is.

Paper Fig. 3. The batch search for the left-front leg. Cell colors are heights. The white outline is the search area around the nominal foothold (green), and the blue marker is the chosen foothold.
Scene 5

What makes a foothold good

In short: each cell's cost adds up several wishes, most important first. The leg should be comfortable, the ground flat and edge-free, the target shouldn't jump around, the stance should stay wide, and the choice should stay close to the last one.

The paper lists six objectives in descending order of weight:

  1. Default leg configuration. Land where the hips will be at lift-off and touch-down, plus the default leg vector. This keeps the walk at the commanded speed. A velocity-feedback term shifts the target when the body is pushed, so the robot can catch itself.
  2. Foothold score. An extra map layer from 0 (flat, safe) to 1 (edges, slopes, rough). It mixes how much the slope varies nearby (edges), the average slope, and how far the height differs from its local average (roughness).
  3. Push over. Footholds keep updating even after the swinging foot passes its highest point, which makes the robot react better. But a target that suddenly jumps across an edge causes violent leg motions. So for swinging legs, the cost includes the worst foothold score on the line between the previous choice and the new cell.
  4. Support area. Prefer cells far from the other feet, for a wider, more stable stance.
  5. Previous foothold. Stay close to the last choice. This stops the target flipping from one side of a neighboring foot to the other.
  6. Leg over-extension. A light pull toward cells under the hips, so the leg doesn't stretch too far.
sf = λ1·σ(s(h)) + λ2·s̄(h)² + λ3·|h − h̄|  // edges + slopes + roughness, weights sum to 1, normalized to [0, 1]
Animated (illustrative). Why the push-over term exists. If the target hops across an edge in mid-swing, the foot has to change course abruptly. Making it costly to cross high-score cells keeps the target steady.

A tilt knob for stairs. The default leg vector points between straight down and the local ground normal. A weight κ between 0 and 1 sets the mix. With κ below 1, the robot leans back when climbing, which helps its knees clear the step edges.

Animated (illustrative). κ blends “straight down” with “perpendicular to the stairs”. The stairs are drawn at 36°, like the real staircase.

The weights λ, κ and the objective weights are not given as numbers in the paper. The cost maps in Scene 5 only show their order and shape, not their real values.

Scene 6

Give a high step more time

In short: a leg that must climb over something tall gets a longer swing, so its motors don't have to spin faster than they can. The other legs wait their turn.

The timetable of lift-offs and touch-downs (the contact schedule) normally repeats on a fixed beat. But a high-stepping leg must travel much farther than a normal one in the same time, so its joint speeds are much higher. Humans do the same thing: our strides get longer in time as the step height grows. The robot takes the swing-over height (how far the highest obstacle on the foot's path rises above both the old and the new foothold) and stretches the swing:

t̃swing = tswing + ko · hmax  // k_o: obstacle gain; h_max: height above the old stance foot + above the new foothold

To keep the order of lift-offs, the other legs' stance phases stretch too. The paper's Fig. 4 shows the case where the left-front leg takes a big step. Its touch-down is delayed, and the next pair's lift-offs are delayed by the same amount, so three feet stay on the ground rather than one.

Paper Fig. 4. Top: one stride of a regular trot (colored squares are stance, dots are swing). Bottom: the same stride when the left-front leg (LF) takes a large step. Its swing is longer, and the other legs' timing shifts to keep the order.

Two more controller details that matter on rough ground

  • Impedance control. Besides the torques from the whole-body controller, each joint is pulled toward a target angle, like a spring and damper. The targets come from inverse kinematics that avoids the straight-leg singularity. This handles over-extended legs and stiffens the joints when a foot slips or gets blocked.
  • Body tilt. The body is aligned with a plane fitted through the current and the planned footholds, so it tilts before reaching a slope, not after.
Scene 7

What the tests showed

In short: the method worked reliably wherever the knees didn't hit the terrain. Steep stairs showed its limits, and stiffer joint control helped there.

All experiments ran on ANYmal, a fully torque-controlled quadruped. Control and state estimation ran at 400 Hz on an on-board Intel i7-7600U. Mapping ran on a separate on-board Jetson AGX Xavier.

Stairs in simulation (144 runs)

Two straight staircases with 12 steps, 0.29 m deep, either 0.10 m (moderate) or 0.17 m (steep) high. The robot was commanded to walk at 0.3 m/s, starting from different positions, and a run counted as a success if it reached the other floor without falling. Each combination was repeated 18 times.

DirectionImpedance controlModerate stairsSteep stairs
Upon18 / 1810 / 18
Upoff18 / 183 / 18
Downon18 / 1814 / 18
Downoff18 / 1812 / 18

What this means: moderate stairs never caused trouble. On steep stairs, going up mostly failed because the knees hit the steps, and going down mostly failed from over-extended swinging legs. Stiffer joint control helped most on the way up, from 3 to 10 successes out of 18.

On the real robot

Obstacle course. An operator drove the robot at 0.3 m/s and steered it into the hardest parts of a lab course for 2 minutes of continuous walking: bricks crossed 3 times, a slope 2 times, a platform 4 times. The platform is 0.22 m high, 33 % of the robot's maximum leg length. The slope and bricks weren't fixed and moved when stepped on.

Paper Fig. 7. Snapshots from the 2-minute obstacle run: up onto the 0.22 m platform, over the loose slope and the bricks.

Real stairs. The same dimensions as the steep simulated stairs: 0.17 m high, 0.29 m deep, about 36°, walked at 0.25 m/s. Across three attempts, the robot failed on average after 18 steps, mostly from knee collisions, as in simulation.

Paper Fig. 8. ANYmal trotting up one floor of a real staircase, with steps 17 cm high and 29 cm deep.
Wrap-up

Limits and what came next

In short: the method plans one step ahead on a simplified balance model and doesn't guarantee the legs can reach or clear the terrain. The authors name exactly these as the next problems.

  • Flat-ground balance. The body planner uses a balance criterion (the zero-moment point) that is only valid on flat ground, and it assumes the body's tilt and height stay constant over the planning window. The authors call these “very limiting”.
  • No reachability guarantee. Footholds aren't checked for whether the leg can actually reach them without hitting the terrain. Together with knee collisions, this makes stair climbing hard. Reliable stair climbing was shown when the stairs were gentle enough not to cause knee collisions or over-extension. Everyday stairs showed both the potential and this weakness.
  • One step at a time. Footholds are adapted locally, not planned as a sequence toward a goal.

The planned next steps: replace the flat-ground balance criterion with one that holds on rough terrain, and handle footholds the legs can't reach.