IEEE ICRA · 2024 · Best Service Robotics Paper · an animated walkthrough

Learning to walk in confined spaces using 3D representation

The paper in 110 seconds · narrated · sound on
Transcript

Under a tilted concrete slab, over loose rubble, between tables and benches, into gaps barely taller than itself. ANYmal picks its own body height. How?

After an earthquake, robots could check places too dangerous for people. Legged robots cross rubble well, but walking tall, they bump into fallen slabs. They must crouch at the right moment. To place their feet, they use a height map: the ground seen from above, like a floor plan. On a floor plan, a slab looks like a solid wall. Earlier crawling robots needed flat floors, or moved slowly and trusted a perfect map.

This paper gives the robot a 3-D picture instead: small cubes called voxels, each empty or full. The gap stays open. Control is split like a driver and a car. The driver network reads the voxels and picks how low to crouch; the car, a proven walking controller, does the walking. They learn in simulation, mostly by trial and error, in randomly built worlds full of steps and low ceilings. A teacher with perfect vision learns first; a student copies it from noisy cubes, like real sensors give.

Why not always crouch? Crouched, it can’t climb a 25 cm box, about a step and a half. Standing, it hits the ceiling. Adapting does both. In simulated tests it reached the goal every time in eight of nine layouts. Always standing or always crouching each fell short in four.

On real ANYmal robots it crawled through a mock collapsed building, choosing its own height while a person only steered. Next: moving obstacles.

Footage: Robotic Systems Lab, ETH Zurich. Animations are schematic. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

Engineers at ETH Zurich taught a four-legged robot to see the world in 3-D and decide for itself when to duck, so it can crawl under fallen slabs and into gaps where people cannot safely go.

  1. Why this matters

    After an earthquake, or inside an old tunnel or factory, robots could check places too dangerous for people. Often the only way in is a low gap under a fallen slab, over loose rubble. A legged robot can lower its body to fit; a wheeled robot of the same size can’t.

  2. What makes it hard

    To place their feet, today’s learning-based walking robots look at a map of ground heights: one number per patch of floor. It’s like a floor plan seen from above, where a table is drawn as a solid block. Under a slab the map shows a wall, and the gap disappears.

  3. What people did before

    Learned controllers already walk fast and reliably over steps, slopes and rubble, even when their sensors are wrong, but only where nothing hangs overhead. Robots that did crawl under things did it on flat floors under a single obstacle, or planned every step slowly with a map they had to trust completely.

  4. What this paper does

    Think of a driver and a car. A “driver” network looks at a 3-D picture of the space around the robot, made of small cubes that are empty or filled, and tells the “car”, a proven walking controller, how low to crouch and how to tilt. Both are trained in simulation, mostly by trial and error, in randomly built virtual worlds full of steps and low ceilings.

  5. What they showed

    In simulated tests the robot reached a goal 6 m away every time in 8 of 9 layouts, and 6 times in 10 under the lowest 0.5 m ceiling, about knee height for an adult. Always walking tall, or always crouching, each fell short in 4 of the 9: crouched, it can’t climb a 25 cm box, about one and a half stair steps. Real ANYmal robots crawled through a mock collapsed building.

  6. Why it’s a step forward

    It keeps the fast, robust walking of earlier controllers and adds awareness of what is overhead, in a map format any 3-D sensor can fill. Honest limits: the test scenes didn’t move, the robot avoids touching things rather than pushing past them, and a person still chooses the direction.

Words used below
voxel
a small cube of space, marked empty or filled: a 3-D pixel
height map
the ground as a grid of numbers, one height per square
policy
the trained neural network that turns what the robot senses into commands
reinforcement learning
learning by trial and error in simulation, with rewards for good behaviour
teacher / student
a network with perfect simulated senses trains one that uses realistic, noisy senses
1 / 8
low-level walking policy 3-D perception (voxels) high-level policy
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

Why getting under things matters

In short: the places robots are most useful, like collapsed buildings, are full of low gaps, and getting in means crouching at the right moment.

Robots are attractive for jobs that are dangerous for people: exploring caves and tunnels, inspecting industrial sites, searching damaged buildings. Legged robots are good at this because they can walk on uneven, unstable ground where wheels slip or get stuck, and this lab’s robots have already hiked mountains and forests and explored underground tunnels.

One ability of legged robots is still barely used: they can change their posture. By bending its legs, a quadruped (a four-legged robot) can lower its body or tilt it, and so fit through gaps that a wheeled robot of the same size cannot. The catch is that crouching has a cost. A crouched robot has less leg left to climb with, so it should crouch only when a ceiling actually comes down, and stand up again for the next step.

The paper’s goal is to add that ability to the robust, learned walking that already works on rough ground, so that one controller can cross rubble and duck under a slab without a human steering its body.

Paper Fig. 1. The goal in pictures: ANYmal crawling through a mock collapsed building, under tilted concrete slabs and timber, over loose gravel and unstable steps. It lowers and tilts its body to fit.
Scene 2

What’s hard: a map that can’t see a ceiling

In short: today’s walking robots see the ground as a map of heights, like a floor plan seen from above, and such a map cannot show that there is space under a slab.

Learned perceptive controllers, including this group’s earlier work, sample the terrain from an elevation map (also called a height map): a grid where each cell stores one height. The paper calls it a 2.5-D representation, and it cannot hold an overhanging structure. Under a slab, either the cell records the slab’s top and the map shows a wall, or it records the ground and the robot is blind to the ceiling. Earlier systems that did crawl under things either walked on flat ground under a single obstacle, or used planning with a two-layer map (floor and ceiling) and a slow static gait that trusted the map to be right.

Animated. The same slab in both maps. The height map turns it into a wall; the voxel grid keeps the free space under it.

The paper’s answer is a 3-D occupancy voxel grid (a voxel is a 3-D pixel): every small cube of space around the robot is either occupied or free. Voxels have two practical advantages beyond representing ceilings. Data from different sensors (dome LiDARs, which are laser scanners that measure distance all around; and depth cameras) can be merged into the same format, so one policy works across sensor setups without retraining, unlike a policy trained on depth images. And voxels can be accumulated over time with standard mapping tools, which gives the robot a memory of what it saw.

Scene 3

Two layers: steer and walk

In short: the robot gets two “brains”: a new one that looks at the 3-D surroundings and decides body height and tilt, and the existing walking controller that carries it out.

The controller is a hierarchy of two neural-network policies, trained in four stages (the numbers in Fig. 2):

  1. Low-level teacher: learns to walk over rough terrain while following a 6-D command, using privileged simulator information (things a real robot cannot measure, like exact contact forces).
  2. Low-level student: distilled from the teacher (trained to imitate it) so it works from noisy on-board observations, as in the authors’ earlier work.
  3. High-level teacher: learns, in procedurally generated confined spaces, to produce those 6-D commands, using exact spherical scans.
  4. High-level student: distilled from it, using noisy occupancy voxels. This is what runs on the robot.

The 6-D command is forward velocity vx, sideways velocity vy, turning rate ωz, and body roll ξ, pitch θ and height h. At deployment the operator still gives only the 3 velocities; the high-level policy decides the body pose on its own. Here a policy means a trained neural network that maps what the robot senses to commands. The split keeps what already worked: the walking layer inherits the robustness of the earlier rough-terrain controller, and the new layer only has to learn spatial awareness.

Animated. The high-level policy outputs a residual on top of the operator’s velocity, so doing nothing already means “follow the operator”. Roll, pitch and height are output directly.
Paper Fig. 2. The four training stages. Top right (1): the low-level teacher tracks a velocity command plus extra roll, pitch and height commands on rough terrain. Top left (3): the high-level teacher sees a spherical ray scan (green dots) in confined training worlds. Bottom (4): the deployed high-level student turns noisy occupancy voxels into a latent with a 3-D convolution, fuses it with proprioception in a recurrent belief encoder, and feeds the low-level student (2).
Scene 4

The walking layer takes posture commands

In short: the walking controller learned to walk over rough ground at whatever body height and tilt it is told, including crouched on stairs.

The low-level policy extends the perceptive locomotion pipeline of Miki et al. (Science Robotics, 2022) with three extra inputs: desired roll, pitch and body height. It keeps that controller’s action space (a phase offset per leg for a periodic gait generator, plus residual joint targets) and its rewards for velocity tracking, torque, joint velocity and acceleration, and foot slip. It is trained by reinforcement learning, trial and error in simulation, using the standard PPO algorithm (Proximal Policy Optimization).

At the start of each training episode a command is drawn at random: roll and pitch from a normal distribution with mean 0 and standard deviation 0.25 rad (about 14°), body height uniformly between 0.1 and 0.6 m above ground. Height above ground is taken as the mean of five ground samples, so the value doesn’t jump at step edges. Two rewards are added for tracking:

rorientation = exp(−α · erp)  // e_rp: roll or pitch error
rbase height = exp(−α · eh)  // e_h: |desired − actual| body height above ground

What the policy observes

InputContents
Commandvx, vy, ωz, roll, pitch, height
Own body (proprioception)body velocity, orientation, a short history of joint position errors and velocities, action history, leg phases
Surroundings (exteroception)height samples around each foot at five radii
Privileged (teacher only)contact states, forces and normals, friction, thigh/shank contacts, external forces and torques, swing phase duration

What this means: the walker still perceives only the ground around its feet. It needs no idea of the ceiling, because someone else, the high-level policy, tells it how low to go.

The student version is distilled by behavior cloning (copying the teacher’s actions) with heavy randomization of the exteroceptive measurements (what it senses of its surroundings), following the earlier paper. On hardware (ANYmal, operator commanding), it tracked combined height, pitch and velocity commands (Fig. 6) and walked crouched over stairs, steps, loose ground and a wooden pallet ramp (Fig. 7).

Real run. The walking layer on its own, with a person sending the posture commands: ANYmal crouches low, tilts its body forward and back, and stands up again, all while stepping. This is the ability the high-level policy later uses. Footage: Robotic Systems Lab, ETH Zurich.
Paper Fig. 6. A real-robot run: commanded (red) and measured (blue) forward speed, body height and pitch. Height follows steps between about 0.2 and 0.6 m; pitch follows ±0.38 rad commands, reaching roughly ±0.3 rad, while the robot also walks.
Paper Fig. 7. (a) High and low body height combined with pitch neutral, up and down. (b) Crouched walking on stairs, a step, rubble and a wooden pallet under a concrete slab.
Scene 5

A generator for confined worlds

In short: to learn, the robot needs many different practice worlds; a generator builds them automatically from terrain pieces and adds low ceilings.

A policy only learns what its training worlds show it. To get a wide variety of confined spaces, the authors wrote a procedural terrain generator and open-sourced it as leggedrobotics/terrain-generator. It builds a terrain mesh by tiling mesh parts, using Wave Function Collapse (WFC), an algorithm common in game level design. Whether two tiles may sit side by side is decided from the terrain heights along their shared edge, so steps meet steps and ramps meet the level they lead to.

Animated. WFC fills a grid one cell at a time. Placing a tile removes incompatible options from its neighbours, and the next cell collapsed is one with few options left.

The paper defines 1570 tile types: steps, stairs, ramps and rough ground at different heights. After the mesh is built, boxes of different sizes are added above it to make overhangs and tight spaces (Fig. 3).

Paper Fig. 3. Top: tiles are connected one by one into a larger terrain. Bottom: overhanging boxes are then added on top of the mesh.

For the simulators: the low-level policy was trained in the RaiSim physics simulator; the high-level policies in Isaac Gym, a GPU-based simulator, which also simulated the sensors (spherical scans, depth cameras and voxel maps). Teacher training used 1000 parallel environments (low- and high-level), student training 300.

Scene 6

Teacher with exact scans, student with noisy voxels

In short: first a “teacher” learns with perfect vision in simulation, then a “student” learns to copy it using the kind of noisy 3-D data a real robot’s sensors give.

High-level teacher

The high-level teacher is trained with reinforcement learning (PPO) to follow a target velocity by commanding the (frozen) low-level policy through confined, rough terrain. It observes the velocity command, body velocity, joint positions and velocities, orientation and previous actions, plus the same foot height samples as the low-level policy and a spherical scan: sparse rays cast in all directions from the robot. The authors chose sparse rays because depth images or voxel grids would slow reinforcement learning down.

Rewards: velocity tracking and a base distance reward that pays for keeping the body away from whatever the rays hit. Penalties: a large one for hitting an obstacle with the body, plus penalties on joint speed, acceleration and torque, and on vertical velocity and orientation, to avoid unnecessary motion.

rbase = exp(−α · (dmax − min(d, dmax)))  // d: ray-cast distance, d_max: threshold, α: coefficient
Animated. The reward is 1 while everything is at least dmax away and falls as the closest hit gets nearer. The paper does not give α or dmax; the values here are for illustration.

High-level student

The deployed student sees the world as occupancy voxels: a 32 × 32 × 32 grid with 0.08 m cells, so a cube of about 2.56 m around the robot. In simulation these voxels are corrupted with noise so the policy tolerates real sensor data. The network:

  • a 3-layer voxel encoder of 3-D convolutions (filters that slide over the voxel grid looking for shapes), with standard normalization and activation layers, compressing the grid to a short list of numbers (a latent);
  • a memory unit, a GRU (gated recurrent unit, 2 layers, hidden size 128), that fuses the compressed voxels with proprioception (the robot’s sense of its own joints and motion) over time, plus a decoder that reconstructs privileged state;
  • a small fully connected network (an MLP, multi-layer perceptron) that outputs the command for the low-level policy.

The loss is the mean squared difference between student and teacher actions plus how well it reconstructs the privileged state.

Animated. The kind of corruption a student sees in training: missing, spurious and misplaced voxels. The paper says the voxels are noisy but does not list the noise model; the types and rates here are illustrative.
Simulation, from the authors’ video. The student policy walking through a generated training world. The two insets show what it could see: the clean voxel grid, and the noisy version it actually trains on. Footage: Robotic Systems Lab, ETH Zurich.
Scene 7

Adapting height is the whole point

In short: always crouching or always standing each fails somewhere; only a robot that changes its height as it goes handles a ceiling and a step together.

Is a high-level policy really needed? Maybe the robot could just crouch all the time. The simulation test answers this. The robot starts 6 m from a goal and must reach it within 30 s, on three kinds of evaluation terrain (Fig. 5): an overhanging box, a box obstacle on the ground, and both together. The policies were trained on generated terrain, never on these tests.

Two baselines skip the high-level policy: High always walks at normal height, Low always crouches. High fails once the overhang comes down to 0.7 m. Low passes every overhang but fails the 0.25 m box, and it fails every combined case. Only the adaptive policies handle the combination, because they must crouch and keep enough leg to climb.

Animated. The combined case: the body goes down for the overhang while the legs still step up onto the box, then comes back up. Motion is illustrative.
Paper Fig. 5. Time-lapse of the student policy on the three evaluation terrains. It drops its body under the overhang, keeps normal height over the ground box, and does both in the combined case.
Scene 8

Results in simulation and in the field

In short: in simulation the deployed controller nearly matched its perfect-vision teacher, and on real robots it crawled through a mock collapsed building.

Success rates from Fig. 4, for every setting tested:

TerrainSettingHighLowStudentTeacher
Overhanging (overhang height)
0.9 m1.001.001.001.00
0.7 m0.001.001.001.00
0.5 m0.001.000.601.00
Obstacle (box height)
0.25 m1.000.001.001.00
0.15 m1.001.001.001.00
0.05 m1.001.001.001.00
Overhanging + obstacle (0.25 m box; overhang height)
1.0 m1.000.001.001.00
0.7 m0.670.001.001.00
0.5 m0.000.001.001.00

What this means: each fixed posture fails somewhere; the adaptive student succeeds everywhere except the lowest overhang, where it gets through 60 % of the time, and stays close to its teacher. Averaged over the nine settings (our calculation): High 63 %, Low 56 %, Student 96 %, Teacher 100 %.

Paper Fig. 4. The same numbers as heat maps: blue is success, red failure. The Student and Teacher columns are almost entirely blue.

On real robots

The policies ran on ANYmal C and ANYmal D, with either dome LiDARs or depth cameras. LiDAR points were voxelized directly; depth-camera data went through voxel mapping. In the field test the operator gave only x, y and yaw velocities, and the robot chose its own body height, roll and pitch. The course was a simulated collapsed building plus tables and benches, with ceilings at varying angles, loose gravel and unstable steps. The paper reports the robot handled all of it robustly (Fig. 1 and 8); it gives no success counts for the field tests.

Real run. The full system in the mock collapsed building: ANYmal keeps its body low under the tilted concrete slab while it climbs over loose rubble, with the operator giving only the walking direction. The original video plays this part sped up. Footage: Robotic Systems Lab, ETH Zurich.
Paper Fig. 8. Two field runs with the full system. Top: the voxel map and the body trajectory (red), which dips as the robot goes under a structure. Bottom: photos at the numbered moments.
Beyond the paper

Limitations and what’s next

In short: it works for still scenes, without touching obstacles, and a person still chooses where to go.

  • Static scenes only. The authors note limits in highly dynamic environments.
  • No body contact. The robot avoids touching the ceiling; the paper names using body contacts in tight spaces as future work.
  • Operator in the loop. The operator still steers (x, y, yaw); the high-level policy decides posture, not where to go.
  • Short evaluation. The simulation test has three terrain types with three parameters each; the field results are qualitative.

About this page: animations are schematic. Motions, ray patterns, noise and reward constants are illustrative and marked so; numbers presented as results come from the paper (Fig. 4, Sec. III–IV).