Science Robotics · 2022 · an animated walkthrough

Learning robust perceptive locomotion for quadrupedal robots in the wild

The paper in 105 seconds · narrated · sound on
Transcript

Meet ANYmal. Hiking an Alpine trail, crossing deep snow, climbing stairs, and exploring caves, with zero falls. Its secret: knowing when to trust its eyes.

Legged robots could go where it’s too far or too dangerous for people, stepping over ledges and rubble that stop wheels. People and dogs move briskly on rough ground because they look ahead. Giving robots that skill is a grand challenge.

So the most robust robots walked blind, feeling the ground only through their legs. They find a step by bumping into it: safe, but slow. Seeing helps, but outdoors the height map, the robot’s picture of the ground, often lies. Grass looks like a wall, snow hides the ground, and slipping feet shift the map.

This paper trains in simulation. First, a teacher that sees the true terrain learns to walk by trial and error, called reinforcement learning. Then a student copies it, using only the robot’s own sensors, with the map deliberately corrupted.

The student’s memory learns when to trust its eyes. On a soft foam block, the map says solid, the foot sinks, and a gate shuts the map out: the legs take over.

On the real ANYmal, with no extra training, it walked twice as fast as the blind controller, climbed 30 cm steps, and beat the signposted hiking time to an Alpine summit. It also carried team CERBERUS’s robots through the DARPA Subterranean Challenge, which they won, without a single fall.

Footage: Robotic Systems Lab, ETH Zurich, and Takahiro Miki (SubT testing-day clip), from the project page videos. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

A four-legged robot learned, in simulation, when to trust its eyes and when to trust its legs, so it can walk fast over real mountains, snow and stairs without falling, even when its view of the ground is wrong.

  1. Why this matters

    Legged robots can go where it is too far or too dangerous for people: hazardous sites such as search-and-rescue scenes, caves and tunnels, one day other planets. Legs can step over ledges, slopes and gaps that stop a wheeled vehicle of the same size.

  2. What makes it hard

    People and dogs cross rough ground briskly because they look ahead. A robot’s “eyes” (depth cameras, laser scanners) are often wrong outdoors: tall grass looks like a wall, snow and shiny floors confuse them, and slipping shifts the whole picture. It is like hiking a dark trail with a flashlight that sometimes lies.

  3. What people did before

    Planning methods pick footholds on a terrain map and assume the map is right. The most robust learned controllers, including the lab’s own earlier work, walked blind: they felt the ground through their legs. That is very safe, but slow, and high steps stop them. Outdoors, controllers avoided falls by switching vision off or adding hand-made reflex rules.

  4. What this paper does

    The robot gets a memory that compares what the map predicts with what its feet actually feel, and a learned “valve” that lets the map in when it agrees and shuts it out when it doesn’t. It is trained in simulation on deliberately corrupted maps, with no hand-written rules.

  5. What they showed

    On the real ANYmal robot: twice the walking speed of the blind controller (1.2 m/s, a brisk walk), steps of 30.5 cm, and a 2.2 km mountain hike whose summit it reached in 31 min, faster than the 35 min on the trail sign. Zero falls in every deployment.

  6. Why it’s a step forward

    For the first time on rough terrain, one controller has both the speed of seeing and the robustness of feeling. It walked four robots through the DARPA underground challenge that their team won. Still open: it guesses when it can’t see, for example at a hidden cliff edge.

Words used below
Proprioception
body sense: joint angles, joint speeds, tilt and acceleration.
Height map
a grid of ground heights the robot builds from its depth sensors.
Height samples
208 heights read from the map in rings around the four feet.
Reinforcement learning
learning by trial and error, guided by a score (reward).
Belief encoder
a memory network that fuses body sense and map into a best guess of the terrain.
Gate (α)
a learned valve between 0 and 1: how much of the map to let through.
1 / 8
proprioception (body sense) exteroception (height map) belief (fused)
Scene 1 of 7

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

Why it matters: going where people can’t

In short: legged robots could do jobs that are too far or too dangerous for people, but only if they can cross rough ground quickly and without falling.

The paper opens with the need. Legged robots can carry out missions in places that are too far away or too dangerous for humans, such as hazardous sites and, one day, the surfaces of other planets. Legs can cross steep slopes, steps and gaps that would stop a wheeled or tracked vehicle of similar size. Several commercial legged robots were already being deployed in the real world.

But no legged robot matched an animal on real-world terrain. People and dogs walk and run briskly over rough ground because they look ahead and plan where to put each foot, and they combine what they see with what they feel, adjusting to slippery or soft ground even when they can’t see well. The paper calls giving robots this ability “a grand challenge in robotics”.

The abstract states the stakes: robots that can operate on their own in remote and hazardous places would open up exploration of areas people rarely reach. And seeing matters for speed and energy: a robot that perceives the ground before touching it can adapt its gait ahead of time instead of reacting after each misstep.

The difficulty in one sentence

Seeing the ground lets a robot move fast, but outdoors its sensors are often wrong, and a robot that trusts wrong information falls. The rest of the page is about how the paper resolves this.

Scene 2

Walking blind vs. walking with eyes

In short: Robots that only feel the ground are safe but slow; robots that see the ground can plan each step, and this paper wants both.

People and dogs walk briskly over rough ground because they look ahead and plan where to put their feet. Legged robots have struggled to do the same in uncontrolled outdoor terrain.

Before this paper, the most robust learned controller for ANYmal (Lee et al., 2020) walked blind. It used only proprioception: joint positions and velocities, the IMU (inertial measurement unit, which senses tilt and acceleration), and a short history of its own actions. Body sense is always available and never lies, so this controller could handle mud, snow, rubble and vegetation. The price is speed. The robot learns about a step only when a foot hits it, so it walks cautiously, and tall steps stop it entirely.

The classical alternative builds a terrain map, plans footholds on it, and optimizes a trajectory. These methods assume the map is essentially correct. Some store per-cell uncertainty, but they use it only through hand-written rules like “avoid risky cells”. That does not cover the many ways a map can fail in the field.

What the paper sets out to do

  • Use exteroception (terrain perception) to anticipate the ground and move fast.
  • Keep the robustness of proprioception when perception is wrong or missing.
  • Learn how to combine the two end-to-end, with no hand-written rules about when to trust which.
Scene 3

Why the map can’t be trusted

In short: The robot’s picture of the ground is often wrong outdoors, in several different ways, so blindly trusting it makes a robot fall.

The robot never sees raw images. Depth sensors produce point clouds, which are registered using the robot’s estimated pose and fused into a robot-centric 2.5D elevation map: one height value per grid cell. The policy then reads height samples from this map in five rings around each foot. The rings have radii of 8, 16, 26, 36 and 48 cm with 6, 8, 10, 12 and 16 points, which is 52 samples per foot and 208 in total.

Animated, 3-D. The 208 height samples the policy reads, on a staircase. Watch what a 15 cm map drift does: samples near each step edge now report the wrong height (red). The drift size is illustrative; ring radii and counts are the paper’s.

Using the map as an abstraction layer makes the controller independent of the sensor: the same policy ran with LiDAR and with stereo cameras. But the map inherits every weakness of the sensors and the pose estimator:

Paper Fig. 3. (A) Height samples (red) drawn from the elevation map on a staircase. (B–G) Real cases where the map is incomplete or misleading. Each case is explained below.
PanelSituationWhat the map showsWhy
BReflective metal floorA deep trench that isn’t thereLight-based depth sensing produces large outliers on reflective surfaces
CDeep snowAn almost empty mapSnow is bright and textureless, so stereo matching fails
DOverhanging branchesTall wallsA 2.5D height field can’t represent anything overhanging
EVegetation, soft snowSolid obstacles at the surfaceGeometry alone can’t tell soft from rigid
FSlippery, loose groundA map shifted relative to realitySlipping feet break the pose estimator’s stable-foothold assumption, so the map drifts
GWalking uphill, past edgesHoles behind structuresThe sensors are on the robot, so anything behind an edge is occluded

What this means: map errors are not rare corner cases. They are part of normal outdoor operation, so a controller that trusts the map blindly will fall.

Scene 4

Step 1 · a teacher that sees everything

In short: In simulation, a “teacher” that is allowed to see everything learns to walk well by trial and error.

Why train in two stages

On the real robot, many things that matter are hidden: the true terrain, friction, contact forces, and whether the map is lying. That makes the task a partially observable decision problem (a POMDP: the robot must act without seeing the full state of the world). The policy has to build a belief from its history of observations, which in practice means a recurrent network. Training such a network with reinforcement learning (RL: trial and error guided by a reward score) from scratch is slow. So the paper uses privileged learning. First a teacher is trained with RL while it sees the full true state, which makes the problem much easier. Then the teacher is distilled into a student that sees only what the robot sees.

Paper Fig. 6. The whole pipeline: (1) teacher trained with PPO (Proximal Policy Optimization, a standard reinforcement-learning algorithm) on privileged information, (2) student distilled from the teacher using noisy height scans, (3) deployment with an elevation map built from onboard LiDAR or depth cameras.

Simulation setup

  • Simulator: RaiSim, with ANYmal C and a learned actuator network to narrow the sim-to-real gap.
  • Scale: 1,000 robots in parallel; the policy runs at 50 Hz; 250 steps per robot per PPO iteration.
  • Task: follow a random velocity command (vx, vy, ω) in the body frame.
  • Domain randomization: body and leg masses, initial pose and velocity, external pushes on the body, and occasional low foot friction to cause slipping.
  • Episode ends on body–ground collision, large body tilt, or exceeding a joint torque limit.

What each policy observes

GroupContentsDim.TeacherStudent
Proprioceptioncommand (3), body orientation (3), body velocity (6), joint positions (12) and velocities (12), joint position history (36), joint velocity history (24), joint target history (24), CPG phase info (13; CPG = central pattern generator, the built-in stepping rhythm)133✓✓
Exteroceptionheight samples, 52 per foot208cleannoisy
Privilegedcontact states (4), contact forces (12), contact normals (12), friction (4), thigh & shank contact (8), external force & torque (6), airtime (4)50✓—

Action space: a rhythm plus corrections

Each leg has a phase variable φl that drives a nominal stepping motion: the foot rises 0.2 m along a cubic curve, comes back down, and stays on the ground for the stance half of the cycle. The policy doesn’t output torques. Each step it outputs a phase offset Δφl per leg and a residual joint target Δqi for all 12 joints. Joint PD controllers then track the target:

qitarget = IK( p(φl + Δφl + Δφ0) ) + Δqi   // Δφ₀ = base frequency, IK = analytic inverse kinematics

Network

The height samples of each foot go through a shared MLP (multilayer perceptron, a plain feed-forward neural network) encoder ge (hidden layers 80, 60) to a 24-d feature, giving a 96-d exteroceptive latent. The privileged state goes through gp (64, 32) to a 24-d latent. These are concatenated with proprioception and fed to the main MLP (256, 160, 128, LeakyReLU).

Reward

The main terms reward tracking the commanded linear and yaw velocity, and penalize sideways drift and unwanted body motion. The rest are shaping terms: foot clearance (lift the swing foot above its surroundings, but not absurdly high), shank and knee collisions, joint velocity and acceleration, a soft limit that stops the knee flipping, smoothness of joint targets, torque, and foot slip.

Terrain and curricula

  • rough (Perlin)
  • rough discrete
  • large steps ≤ 0.4 m
  • boxes ≤ 0.25 m
  • grid steps
  • step stairs
  • stairs: standard · open · ledged · random

Stairs are built from boxes rather than the height map. Height-map risers are slightly slanted, and the policy learned to exploit them in simulation, which then failed on real, vertical stairs. An adaptive terrain curriculum uses a particle filter over the terrain parameters to keep each terrain “challenging but achievable” as the policy improves. A second curriculum scales domain randomization and several penalty terms by a factor ck+1 = ck0.98 that rises toward 1.

Scene 5

Step 2 · a student trained on bad data

In short: A “student” learns to copy the teacher using only what the real robot can sense, with its view of the ground deliberately spoiled during training.

The student runs in the same simulated world but only receives proprioception and noisy height samples n(oe). It is trained with supervised learning on two losses:

L = Lbc + 0.5 · Lre
L_bc : squared distance between the student’s and the teacher’s action in the same state
L_re : squared distance between the decoded and the true (clean heights, privileged state)

The training data comes from rolling out the student, not the teacher, in the style of DAgger (dataset aggregation, a standard imitation-learning recipe). The student therefore visits the states its own mistakes lead to, and the teacher labels what to do there. Its MLP head has the same shape as the teacher’s and starts from the teacher’s weights. The recurrent part is trained with truncated backpropagation through time over 10 steps, using 300 robots × 400 steps per iteration.

The noise model

Every height sample is perturbed both in where it is read and in what height it reports:

xp = rpcos θp + εpx + εfx + wx
yp = rpsin θp + εpy + εfy + wy
hp = h(xp, yp) + εpz + εfz + wz + εoutlier
TermScopeResampledMimics
εpeach scan pointevery stepsensor noise
εfall points of one footevery steplocal map misalignment, jitter
wall points of one footonce per episodepersistent offsets: drift, soft ground
εoutlierrandom subset of pointsevery stepspikes, e.g. reflections
Animated. One foot, the four noise terms switched on one at a time. Per-point noise scatters samples; the per-foot shift moves all 52 together; the per-episode offset stays put; outliers spike. Magnitudes are illustrative.

On top of this, each training terrain is split into cells with their own height offset, which mimics transitions such as the edge of a patch of vegetation or deep snow. The noise parameters are set by one of three regimes, drawn at the start and again in the middle of every episode:

RegimeShareDominant noiseStands in for
Nominal60%small everywherea normal, good map
Large offset30%large per-foot offsetspose drift, deformable ground
Large noise10%large per-point noise and outliersocclusion, sensor failure: no usable map

The student starts on flat ground with no noise. The terrain curriculum switches on after 10 epochs, and the noise ramps up linearly from epoch 20 to 100.

Paper Fig. 7. (A) The noise terms around one foot, where each zi sets one standard deviation. (B) The same step seen with zero, nominal, large-offset and large noise. (C) The belief encoder and (D) the belief decoder, covered in the next section.
Scene 6

The belief encoder decides what to trust

In short: The student’s memory compares what the map predicts with what the feet feel, and a learned gate decides how much of the map to use at each moment.

This is the core of the paper. The student is a recurrent belief encoder followed by the same MLP head as the teacher. The encoder’s job is to produce a belief state bt that matches the teacher’s latent features (le, lpriv). In other words, it has to recover what the teacher would know, from noisy and partial inputs.

b′t, ht+1 = GRU( otp, lte, ht )   // l^e = g_e(noisy heights)
α = σ( ga(b′t) )   // 96-d attention, one weight per exteroceptive feature
bt = gb(b′t) + lte ⊙ α   // 120-d belief (96 terrain + 24 privileged)
  • The GRU (gated recurrent unit, a common kind of memory network; 2 layers × 50 units) is the memory. It carries what the feet have felt and how the map has behaved over the last moments. That is what lets the robot notice “the map said solid, but my foot sank”.
  • The gate α is a learned, per-feature valve on a skip connection. When the map is consistent with the body’s experience, clean exteroceptive features pass straight into the belief. When it isn’t, the gate closes and the belief is formed from the recurrent state, which is dominated by proprioception.
  • The decoder uses the same gating idea to reconstruct the 208 clean heights and the privileged state from the belief. It provides Lre during training, and afterwards it lets you look inside the policy on the real robot.
Animated. The soft-foam experiment. The map (dashed) says “solid block”. When a front foot sinks in, the gate α drops and the belief (purple) follows the legs instead of the map. Gate values are illustrative, not measured.

Looking inside the belief on the real robot

Paper Fig. 5. Red dots are the height samples fed to the policy; blue dots are the terrain decoded from its belief. (A) Soft foam block, (B) transparent acrylic block, (C/D) stairs with sensors uncovered and covered, (E) slippery platform with the decoded friction coefficient.
  • Soft foam (A): the map says “solid block”, so the robot first prepares to step on it. On contact the foot sinks, and the belief is revised downward. It stays revised after the foot lifts off, because the GRU remembers.
  • Transparent block (B): the sensors don’t see the acrylic, so the robot walks as if on flat ground until a foot hits the step. Then the belief jumps upward and the gait changes to climb.
  • Sensors covered (C/D): with sensors working, the robot climbs the stairs without touching a riser. With the sensors physically covered, it receives random values, bumps the first riser, updates its belief, and still climbs up and down.
  • Slippery platform (E): the decoded friction coefficient drops sharply once the robot is on the platform, and it steps faster to keep balance. The sliding feet also corrupt the pose estimate and the map. The policy relies on proprioception until the map settles again.
Real run. The soft-obstacle test from panel A: to the sensors the yellow foam block looks like a solid step, but it squashes as soon as ANYmal puts weight on it. The robot adjusts and walks on without falling. Footage: Robotic Systems Lab, ETH Zurich (Movie S4).

Does the gate matter? (ablation)

Paper Fig. S2. Four students compared: GRU or MLP encoder, each with or without the gate. (A) Training curves. (B) Success at climbing a step under small and large exteroceptive noise, 100 trials per height.
TerrainSmall noise · gateno gateLarge noise · gateno gate
rough0.690.750.881.00
large step0.720.861.141.23
grid steps1.441.672.222.21
standard stair0.850.961.391.44
boxes0.931.121.611.68

Mean squared difference between student and teacher actions (lower is better), selected rows of Table S4.

What this means: a recurrent encoder beats a feed-forward one, and the gate helps most when the map is good, because it lets clean terrain information through. When the map is bad, gated and ungated are nearly equal. The gate closes and the policy behaves like a proprioceptive one, so it gains foresight without losing robustness.

Scene 7

From sensors to joints

In short: On the real robot, the trained student is simply switched on: sensors build a map, the policy reads it 50 times a second and moves the legs.

The trained student is deployed on ANYmal C zero-shot: no fine-tuning on the real robot. Two sensor setups were used interchangeably without retraining: two Robosense Bpearl dome LiDARs, or four Intel RealSense D435 depth cameras.

  1. Mapping, 20 Hz. A GPU elevation-mapping pipeline registers point clouds using the estimated pose and updates the map in a Kalman-filter fashion. It adds drift compensation and ray casting to clear stale obstacles. Fast mapping mattered: at 1.2 m/s, a slow map would always lag the robot.
  2. Sampling. Each control step, 208 heights are read from the latest map around the current foot positions. Where the map has no data, a random value is filled in. The policy learned to handle exactly this in the “large noise” regime.
  3. Policy, 50 Hz. The belief encoder and MLP turn proprioception, height samples and the velocity command into phase offsets and residual joint targets.
  4. Joints. CPG phase plus inverse kinematics plus residuals gives 12 joint position targets, tracked by the actuators’ PD loops.

The velocity command can come from a joystick or a navigation planner. The obstacle-course test used a pure-pursuit path follower, and team CERBERUS used autonomous planners in the DARPA SubT Challenge. The planner doesn’t need to know the ground type or switch modes; the controller handles that.

Scene 8

Results in the wild

In short: Out in the world the controller was twice as fast as the blind one, climbed higher steps, kept human hiking pace on a mountain, and never fell.

Deployments across seasons

The controller was used in alpine, forest, underground and urban settings, in rain, snow, fog and darkness, with zero falls across all deployments. It handled stairs in any direction (sideways, diagonally, turning on the stairway) without a dedicated stair mode.

Paper Fig. 1. Natural terrain (slippery, steep, vegetated), search-and-rescue scenarios (stairs, payloads, fog), caves (reflective surfaces, loose ground, low light, water), and winter snow, often several at once.

A hike on Mount Etzel

A 2.2 km loop in Switzerland with 120 m of climb, slopes up to 38%, wet rock, high steps and tree roots. A hiking planner rates it “difficult”. ANYmal needed no human help, and stopped only to reattach a detached shoe and swap batteries.

Paper Fig. 2. The Etzel route and scenes along it, including the steep rocky ascent and the slippery, root-covered forest descent.
SegmentANYmalReference for humans
To the summit31 min35 min (official signpost)
Whole route78 min76 min (hiking planner)

What this means: on a trail made for people, the robot kept human hiking pace for over an hour with no falls.

Controlled comparison with the blind baseline

Paper Fig. 4. (A, B) Step climbing, 10 trials per height with a 5 s limit. (C, D) Obstacle course of a ramp, a 20 cm platform, stairs and blocks, tracked with pure pursuit. (E–H) Maximum forward, lateral and yaw speed on flat ground and over a 20 cm step.
TestThis paperBlind baseline
Highest step climbed reliably30.5 cmsuccess drops at 20 cm
Obstacle course time33 s, unassisted75 s, with human help
Forward and lateral speed, flat ground1.2 m/s0.6 m/s
Turning rate3 rad/s0.6 rad/s

What this means: seeing the terrain doubles walking speed, makes turning five times faster, and raises the step limit by about 10 cm. The baseline’s front legs got stuck at the riser, while this controller lifted its legs early and leaned forward so the hind legs could swing over. Above about 32 cm it hesitated: it had learned that such steps are at the edge of what the robot can physically do.

Real run. The 30.5 cm step, the highest one this controller climbed reliably. It sees the step coming, lifts its front legs high before touching it and leans forward so the hind legs can follow. Footage: Robotic Systems Lab, ETH Zurich (Movie S2).
Paper Fig. S1. The same comparison in simulation: success rate at crossing 4 m at 0.7 m/s, over a 41 × 41 grid of terrain parameters for grid steps (A) and stairs (B). The perceptive controller succeeds over a much larger part of the grid.

DARPA Subterranean Challenge

This was the default locomotion controller of team CERBERUS, which won the SubT Final. It drove four ANYmals through more than 1,700 m of tunnel, urban and cave courses, amid dust, fog, water and smoke, without a single fall.

Real run. A team CERBERUS ANYmal in a cave passage on a DARPA SubT Challenge testing day, shown sped up (about 20×). Narrow walls, loose sand and rocks, and poor light. Footage: Takahiro Miki, Robotic Systems Lab, ETH Zurich (linked from the paper’s project page).
Discussion

Limitations and what’s next

In short: The robot still guesses when it can’t see (for example behind a cliff edge), and it only walks, it doesn’t climb or scramble.

  • Uncertainty is only implicit. When a cliff edge or stepping stone is occluded, the policy tends to assume the surface continues, and could step off. Estimating uncertainty explicitly might let it slow down or probe the ground with a foot first. A learned occlusion model could also help.
  • The elevation map is a bottleneck. It throws away texture and material cues present in raw sensor data, and the pose estimator that builds it is not trained with the rest of the system. Learning directly from raw sensors could make the controller faster and more robust.
  • It walks, and only walks. Maneuvers far from normal gait, such as pulling a leg out of a narrow hole or climbing onto a high ledge, are out of scope.