Doctoral thesis · ETH Zurich · 2025 · an animated walkthrough

Bridging Perception and Control for Legged Locomotion and Navigation in the Wild

The thesis in 111 seconds · narrated · sound on
Transcript

This robot hikes mountain trails, walks through snow, crouches under collapsed concrete, and explores dark caves. To get here, it had to learn when to trust its own eyes.

Legged robots could go where people shouldn’t. But underground, the radio link fades, so the robot must see the ground and choose its own way. That’s hard, because its sensors lie. Tall grass looks like a wall. Snow looks solid, until the legs sink in.

The thesis answers in four steps: a fast terrain map, sight plus touch, a 3-D view for ceilings, and common sense.

First, the map. The graphics chip turns about 43,000 laser points into a grid of ground heights in under 7 milliseconds.

Then, a walking brain trained in simulation. Its belief encoder, a small memory, stops trusting the map when the feet disagree. It hiked a 2.2 km mountain loop in 78 minutes, close to human hiking time, without a single fall.

Next, ceilings. A second policy reads a small 3-D grid of cubes, called voxels, and tells the walker when to crouch. On the real robot, it chose its own posture through a mock collapsed building.

Last, common sense. A vision-language model, an AI that describes images, rates what is risky. Here it reads a road-work sign and takes the grass.

With all four pieces, the robot hiked 10 kilometres down from Gornergrat above Zermatt, needing a human only twice. The lesson: give a learning robot simple, trustworthy inputs, and what it learns in simulation holds up in the wild.

Footage: Learning robust perceptive locomotion… and Learning to walk in confined spaces…, Robotic Systems Lab, ETH Zurich. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

Over four projects, this thesis taught a four-legged robot to use its eyes without being fooled by them, and then to use common sense, until it could hike ten kilometres down a Swiss mountain largely on its own.

  1. Why this matters

    Robots with legs can go where wheels can’t and where people shouldn’t: collapsed buildings, caves, steep trails, maybe other planets. In those places the radio link is weak or slow, so nobody can steer the robot step by step. It has to walk and choose its way by itself.

  2. What makes it hard

    The robot’s “eyes” (laser scanners and depth cameras) measure shape, and in the wild they are often wrong. Tall grass looks like a wall. Snow looks solid until the legs sink in. A shiny floor looks like a hole. And shape alone can’t tell a lake from a lawn, or read a sign that says “keep off”.

  3. What people did before

    Classic controllers planned every footstep on a map and assumed the map was right. That works in a lab. Newer controllers trained by trial and error in simulation were very robust, but blind: they felt the ground with their feet, like walking in the dark, so they were slow and tripped on high steps.

  4. What this thesis does

    It builds one system in four steps: a map that turns sensor points into terrain fast (Ch 2); a walking controller that mixes sight with touch and learns when not to trust sight (Ch 3); a 3-D view so the robot knows about ceilings and crouches (Ch 4); and a vision-language model that adds common sense about what is risky (Ch 5).

  5. What it showed

    The robot hiked a 2.2 km mountain loop in 78 minutes, about the time a hiking planner gives people, without a fall. The same controller helped the team that won the DARPA Subterranean Challenge. The robot crouched through a mock collapsed building, read road-work signs, and hiked 10.1 km above Zermatt with two human interventions.

  6. Why it’s a step forward

    The walking controller became the lab’s default for its ANYmal robots. The main lesson: give the learning robot simple, cleaned-up inputs (heights, cubes, risk values) so that what it learns in simulation holds up outside. The open questions are richer raw senses, learning on the real robot, and AI models that understand physics better.

Words used below
Touch (proprioception)
the robot’s sense of its own joints and body motion
Sight (exteroception)
measuring the world outside: lasers, depth and color cameras
Elevation map
a grid around the robot with one ground height per cell
Reinforcement learning
training a controller by trial and error, here in simulation
Voxels
small 3-D cubes marking where something is
Vision-language model
an AI trained on internet images and text that can describe a scene
1 / 8
touch sight: shape, maps, voxels meaning: common sense
Scene 1

Read the full section with the thesis figures ↓

The thesis, chapter by chapter

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the thesis; the text is a plain-language walkthrough. Chapters 2 to 4 are published papers, each with its own, more detailed explainer linked in its section; Chapter 5 is covered here in the most detail.

Scene 1

Why legged robots need to see and decide

In short: legged robots could help in disaster zones, caves, mountains and on other planets, but only if they can walk and find their way without a person steering every step.

The thesis opens with the promise of legged robots: legs handle uneven ground, forests and snow that stop wheels and tracks. The places where that matters most are also where remote control breaks down. Radio signals are unreliable in tunnels and caves, and on a distant planet every command arrives late. So the robot needs autonomy: it has to move robustly and pick its own path.

The introduction points to evolution: the ability to see is thought to have been a turning point for how animals move, because eyesight gives foresight. Robots need the same, and they need it to keep working when the view is poor. That is the thread through every chapter: sight must be combined with touch, and, at the end, with meaning.

Thesis Fig. 1.1. Where the thesis ends up: the robot hiking autonomously below the Matterhorn, using shape, meaning and common sense together.
Scene 2

What goes wrong in the wild

In short: the robot’s sensors are often wrong outside the lab, and some dangers can’t be seen from shape at all; earlier controllers either trusted the map too much or didn’t use it.

The thesis names three challenges. Complex terrain: slopes, steps, gaps and tight spaces, with ground that slips (gravel, ice) or gives way (mud, snow). Perception you can’t fully trust: lidar suffers in fog, dust and rain; stereo cameras struggle with textureless, shiny or badly lit surfaces; time-of-flight cameras miss dark surfaces and get noisy in sunlight; and errors in the robot’s own position estimate smear the map. Meaning: geometry says nothing about whether a flat surface is a gravel path or a lake, or whether a sign says “no entry”.

Before this work, two kinds of controllers existed. Model-based controllers planned footholds and motions on a terrain map. They could be precise, but they assumed an accurate map, so they were mostly shown in controlled settings. Learned controllers, trained by reinforcement learning in simulation, had recently been moved onto real robots with great robustness, but they used touch only, so they had to bump into a step to find it.

Thesis Fig. 3.1. The conditions the Chapter 3 controller met over several seasons: steep and slippery slopes, vegetation, fog, caves with water and poor light, snow on stairs.
Scene 3

The arc: one system, four chapters

In short: each chapter fixes the main limitation of the one before it, and all of them run together on the same robot, ANYmal.

  1. Ch 2, a fast terrain map. Walking needs a map built in real time from many sensor points. Moving the work to the graphics chip made that possible.
  2. Ch 3, sight + touch. With a map in hand, the robot still needs a controller that doesn’t fall when the map lies. A learned controller mixes the map with touch and decides how far to trust it.
  3. Ch 4, ceilings. A map with one height per cell can’t show a table top or a low beam. A small 3-D grid and a second “posture” policy let the robot crouch.
  4. Ch 5, common sense. Up to here the robot was steered by a person with a joystick, and it only knew shapes. A vision-language model adds meaning (risk, preferred paths, signs) so the robot can navigate on its own.

The later chapters reuse the earlier ones: Chapter 4’s walker is the Chapter 3 controller with three extra commands, and Chapter 5 paints its risk values into the Chapter 2 map and walks with a Chapter 3-style controller.

ChapterPaperDetailed explainer
2 · Elevation mapping on the GPUIROS 2022 (extension: MEM, IROS 2023)GPU elevation mapping · MEM
3 · Perceptive locomotionScience Robotics 2022Robust perceptive locomotion
4 · Confined spaces with voxelsICRA 2024 (best service robotics paper award)Walking in confined spaces
5 · World knowledge for navigation (ROAM)in preparation at the time of the thesiscovered on this page (scene 7)

What this means: this page tells the overall story; for the details of any one step, follow the link for that chapter.

The thesis also lists follow-up work that used the Chapter 3 controller: CERBERUS in the DARPA Subterranean Challenge, a team of legged robots exploring planetary analog sites, and navigation and locomotion for wheeled-legged robots, which extended its training method.

Scene 4 · Chapter 2

A fast terrain map on the GPU

In short: the robot turns tens of thousands of sensor points into a grid of ground heights many times per second, on the graphics chip, and cleans up the errors that field work revealed.

An elevation map is a grid around the robot that stores one ground height per cell, often called 2.5-D because it can’t store two heights in one place. The lab’s earlier mapping software ran on the main processor and could not keep up with the flood of points needed for fast, agile walking. Chapter 2 rebuilt it to run on the GPU, where thousands of points are processed at once.

Field deployments shaped its features. Drift compensation: when the robot’s estimate of its own height drifts, fresh points no longer line up with the map, so the map is shifted to match them. Visibility cleanup: rays from the sensor that pass through a cell prove it is empty, which removes ghosts such as an overhang that was recorded as a wall. It also adds a learned traversability filter, an upper-bound layer and the smoothing, inpainting and plane-finding filters that the walking controllers use.

Animated. Drift compensation: the new scan sits 10 cm above the map, the average gap is measured, and the map moves to meet it. Values illustrative.
Step (Jetson Xavier, 43,017 lidar points)Time (ms)
Move points into the map frame, count height errors1.194
Drift compensation0.742
Height update + ray casting0.648
Overlap clearance0.003
Learned traversability filter4.102
Surface normals0.168
Total6.857

What this means: a complete update takes about 7 thousandths of a second on the robot’s own small computer; the traversability filter is the slowest part.

Sensor settingPointsMap updates / sSensor rate
Depth camera, thinned6,27649.460
Depth camera, raw407,04016.160
Dome lidar43,07419.9920

What this means: with the lidar, the map keeps up with every single scan; even 400,000 raw camera points per frame still give 16 updates per second, which the thesis calls fast enough for its use.

Thesis Fig. 2.8. Processing time as the number of points grows. The earlier CPU mapper slows down steeply; the GPU version stays low even with extra work such as ray casting.
Thesis Fig. 2.1. One map, several layers: raw heights, how walkable each cell is, surface directions, filled-in gaps, and flat planes found for footstep planning.

The map ran on four ANYmal robots of team CERBERUS in the DARPA Subterranean Challenge and became the perception backbone of the later chapters. It is open source (elevation_mapping_cupy). An extension, MEM, adds color and semantic layers from cameras, which Chapter 5 relies on. Details: GPU elevation mapping explainer · MEM explainer.

Scene 5 · Chapter 3

Walking by sight and touch

In short: a learned walking controller that uses the map when it is right and falls back on touch when it isn’t, trained in simulation and used on real robots without changes.

The key part is a belief encoder: a small memory network that takes touch (joint states, body motion, recent history) and sight (heights sampled from the elevation map around each foot) and forms a running estimate of the terrain. A learned gate inside it decides how much of the map to let through.

It is trained in two stages, like a teacher and a student. A teacher policy learns to walk in simulation with perfect knowledge of the terrain. A student then learns to copy the teacher while seeing only what a real robot would see: map heights that are deliberately corrupted with noise, offsets and gaps. To copy the teacher well, the student has to learn when the map can be believed. The resulting controller was put on the robot with no fine-tuning.

Animated. Why sight helps: as the step grows, the touch-only controller starts failing at 20 cm; the controller that also sees handles steps up to 30.5 cm, and above 32 cm it holds back.
Test (thesis Sec. 3.2.4)Sight + touchTouch only
Highest step climbed reliably30.5 cmsuccess drops at 20 cm
Top speed, flat ground (forward and sideways)1.2 m/s0.6 m/s
Turning speed3 rad/s0.6 rad/s
Obstacle course (platforms, stairs, blocks)no help neededstuck at all three obstacles

What this means: seeing the ground made the robot twice as fast, five times quicker at turning, and able to climb steps about 10 cm higher.

Thesis Fig. 3.5. Looking inside the controller. Red dots are the map heights it received; blue dots are what it believes. On a soft block the belief drops once the foot sinks; a see-through obstacle is added after contact; with covered sensors it still climbs the stairs by touch.
Thesis Fig. 3.2. The Etzel hike: 2.2 km, 120 m of climbing, slopes up to 38 %. The robot reached the summit in 31 minutes (the trail sign says 35) and finished the loop in 78 minutes, close to the 76 minutes a hiking planner suggests.

In the thesis’s words, this was the first time a quadruped completed such a hike without a fall; the only stops were to reattach a shoe and swap batteries. A person with a joystick gave the direction and speed. Details: perceptive locomotion explainer.

Scene 6 · Chapter 4

Crouching through tight spaces

In short: a second policy looks at a small 3-D grid around the robot and tells the walker how low to crouch and how to tilt, so the robot can get under low ceilings while still walking over rubble.

Legs let a robot lower its body, something wheeled robots of the same size can’t do. Earlier crouching demos used simple flat floors with one obstacle, or slow, careful gaits that assumed a perfect map. Chapter 4 splits the job in two. The low-level walker is the Chapter 3 controller extended with three more commands (body height, roll and pitch), trained on rough terrain. A high-level posture policy sees the space around the robot and sets those commands.

The posture policy first learns as a teacher with clean “spherical” distance scans, then a student learns to copy it from noisy voxels: a 32 × 32 × 32 grid of 8 cm cubes, which can be filled from lidar or depth cameras alike. Training worlds were generated automatically by fitting together 1,570 kinds of terrain tiles (steps, stairs, ramps) and dropping boxes on top as ceilings.

Animated. As the ceiling slopes down, the posture command lowers the body and tilts it; the walker below carries it out. Motion illustrative.
Simulation test (Fig. 4.4), success rateAlways tallAlways crouchedCh 4 student
Ceiling only, at height
0.9 m1.001.001.00
0.7 m0.001.001.00
0.5 m0.001.000.60
Box on the ground only, box height
0.25 m1.000.001.00
0.15 m and 0.05 m1.001.001.00
0.25 m box + ceiling at height
1.0 m1.000.001.00
0.7 m0.670.001.00
0.5 m0.000.001.00

What this means: each fixed strategy fails one kind of obstacle; only changing posture on the fly handles a step and a low ceiling together (the student matched its teacher everywhere except the lowest ceiling alone).

Real run. The robot lowers its body to pass under a concrete slab while stepping over a pallet and loose gravel; the operator only sets direction and speed. Footage: Robotic Systems Lab, ETH Zurich.
Thesis Fig. 4.2. Four training stages: the walker as teacher then student, and the posture policy as teacher (clean scans) then student (noisy voxels through a 3-D encoder and memory).
Thesis Fig. 4.1. The real test: a mock collapsed building with loose gravel underfoot and tilted slabs overhead, taken with the robot choosing its own posture.

On the real ANYmal C and D, the operator only gave direction and speed; the robot chose its own height and tilt in a mock collapsed building with loose gravel, unstable steps and tilted slabs. Details: confined spaces explainer.

Scene 7 · Chapter 5

Common sense from a vision-language model

In short: an AI model that understands images and text judges what is risky, a fast network turns its judgement into a map layer, and trained policies walk and navigate on that map, so the robot can read signs and avoid water on its own.

Up to Chapter 4 a person with a joystick chose where to go. For full autonomy, the team first added GPS-based navigation along a planned route, and found that shape alone was not enough: the robot also has to avoid snow, lakes, private pastures and places that signs close off. People handle this with common sense. The system in Chapter 5, called ROAM, borrows it from a vision-language model (VLM; the experiments used GPT-4V) trained on internet-scale images and text.

A VLM is slow (seconds per answer) and weak at exact positions, so it isn’t used to steer. Instead, about every 10 s it gets the camera images, a description of the robot, and its own recent notes, and returns a short list of labels such as gravel_path, grass, water, each with a risk from 0 to 1, a steppable flag and a preferred-area flag. An open-vocabulary segmentation network then finds those labels in every camera pixel 6 times per second, and the values are painted into the elevation map as extra layers. A sign detector watches for signs; when it finds one, the robot pauses and asks the VLM straight away.

Animated. How a pixel gets its risk: each label’s value is weighted by how likely the pixel shows that label. The weights (path 0.1, grass 0.5, lake 1.0, anything else 1.0) are the thesis’s own example.
risk(pixel) = Σlabels vlabel · plabel  // v from the VLM, p from a softmax over the segmentation scores

Two policies trained by reinforcement learning use the map. The walking policy (a Chapter 3-style teacher–student controller) also receives a steppable scan and is penalized in simulation for stepping onto non-steppable patches. The navigation policy (10 commands per second) gets the risk layer and treats areas next to the preferred path as if they were raised, so staying on the path is the easy choice. A fallback returns the robot to an important earlier waypoint, rated by the VLM, when it gets stuck.

Animated. Slow reasoning, fast reflexes: the VLM speaks every ~10 s, but the map, navigation and walking layers run many times per second, so a slow answer never delays a footstep. (Map rate from thesis Fig. 5.5, other rates from the text.)
Thesis Fig. 5.5. The ROAM pipeline, from camera images through the VLM and segmentation into map layers for walking (steppable) and for navigation (risk, preferred area).

Compared with other navigation systems (1.2 km course in Zurich)

ChallengeROAMGeometry only (Field Local Planner)Goal images (ViNT)
Global path wrongly runs through bushes / grasswent around, stayed on the pathgot stuck in bushes; walked on grassneeds images recorded beforehand
Mannequin lying on the pathlabelled “human”, walked aroundstepped on its legneeded several interventions, stepped on it
Road-work signs at start and end of a closed sectionread both, switched to the grass and backwalked into the work zonehit the barrier, drifted into tall grass

What this means: only the system with common sense handled all three without a person stepping in; the others lacked either meaning or true autonomy.

Thesis Fig. 5.2. The three comparison challenges, including the model’s answers before and after reading the road-work signs; a memory of its own recent notes kept the closure in mind after the sign was out of view.

Smaller tests showed the same common sense: with a “WARNING: Exploding Cans” sign, the robot placed its feet between the cans; with “DO NOT STEP HERE” on a wooden pallet it used only the grass, and with warning signs on the grass it used only the pallet. With meaning switched off, it cut across grass and railway tracks and walked into snow; with meaning on, it kept to the road.

Scene 8

What it adds up to

In short: from a joystick-steered hike without falls to a 10 km mountain descent chosen largely by the robot itself.

Field testUsesDistanceNotes
Etzel hikeCh 3 controller, joystick2.2 km120 m up, 78 min (planner: 76), no falls
DARPA SubT, four ANYmalsCh 3 controller> 1.7 kmtunnel, urban and cave courses, no falls; team CERBERUS won
Seealpsee hikeROAM (Ch 5)2.38 kmlake shore, cows, pastures; a few short interventions
Zermatt hikeROAM (Ch 5)10.1 kmGornergrat to Furi, 1,220 m down, slopes ~30 %; 2 interventions

What this means: the distance the robot could cover, and how much of the decision-making it did itself, grew with each chapter.

The Zermatt interventions were both at construction sites: once the segmentation missed a fence the VLM had flagged, and once it could not tell a new detour path from rough construction ground. At Seealpsee the operator stepped in once as a precaution near the lake, once after a slip on steep gravel, and the robot had to be carried over a turnstile gate.

Thesis Fig. 5.4. The two autonomous hikes: Zermatt (10.1 km, from above 3,000 m down to Furi) and around Seealpsee (2.38 km), with photos of gravel, lake shore, rail tracks and cows.
Thesis Fig. 5.1. The ROAM system across mountain trails, forest, campus, roads, indoor stairs, smoke and rail tracks, including a sign it read and acted on.

Beyond these tests, the thesis reports that the Chapter 3 controller was also used in the ESA Lunar Resource Challenge and at Cybathlon, and became the default controller for the lab’s ANYmal robots, supporting research from navigation to mobile manipulation.

Discussion

What made it work, and what’s next

In short: the thesis credits two choices for surviving the jump from simulation to reality, and names the limits of those same choices.

What made it work. A policy trained in simulation behaves well in the real world when real conditions fall inside what it saw in training. So the training randomized terrain, the robot’s mass and size, and pushes. And the robot never saw raw camera pixels: it saw simple, cleaned-up inputs (heights around each foot, a small voxel grid, heights plus risk values), which are much easier to make realistic in simulation than full images.

Where that stops. Choosing the right simplification is hard for more complex tasks such as dexterous manipulation. An in-between map adds parts to the system and can throw away subtle cues, and it can stop the policy from learning clever behaviour of its own, such as looking around to reduce uncertainty. In Chapter 5, answers are only as good as the VLM and the segmentation (a flagged fence was missed), the model gives slow answers, and the system still relies on GPS and a planned route.

Next steps named in the thesis. Agility closer to animals (leaping gaps, climbing steep slopes); less hand-tuning of rewards and training worlds; learning from richer, rawer senses as simulators get better; learning or fine-tuning on the real robot instead of going back to simulation after each failure; and foundation models with a better grasp of space and physics.