Over four projects, this thesis taught a four-legged robot to use its eyes without being fooled by them, and then to use common sense, until it could hike ten kilometres down a Swiss mountain largely on its own.
Why this matters
Robots with legs can go where wheels can’t and where people shouldn’t: collapsed buildings, caves, steep trails, maybe other planets. In those places the radio link is weak or slow, so nobody can steer the robot step by step. It has to walk and choose its way by itself.
What makes it hard
The robot’s “eyes” (laser scanners and depth cameras) measure shape, and in the wild they are often wrong. Tall grass looks like a wall. Snow looks solid until the legs sink in. A shiny floor looks like a hole. And shape alone can’t tell a lake from a lawn, or read a sign that says “keep off”.
What people did before
Classic controllers planned every footstep on a map and assumed the map was right. That works in a lab. Newer controllers trained by trial and error in simulation were very robust, but blind: they felt the ground with their feet, like walking in the dark, so they were slow and tripped on high steps.
What this thesis does
It builds one system in four steps: a map that turns sensor points into terrain fast (Ch 2); a walking controller that mixes sight with touch and learns when not to trust sight (Ch 3); a 3-D view so the robot knows about ceilings and crouches (Ch 4); and a vision-language model that adds common sense about what is risky (Ch 5).
What it showed
The robot hiked a 2.2 km mountain loop in 78 minutes, about the time a hiking planner gives people, without a fall. The same controller helped the team that won the DARPA Subterranean Challenge. The robot crouched through a mock collapsed building, read road-work signs, and hiked 10.1 km above Zermatt with two human interventions.
Why it’s a step forward
The walking controller became the lab’s default for its ANYmal robots. The main lesson: give the learning robot simple, cleaned-up inputs (heights, cubes, risk values) so that what it learns in simulation holds up outside. The open questions are richer raw senses, learning on the real robot, and AI models that understand physics better.
- Touch (proprioception)
- the robot’s sense of its own joints and body motion
- Sight (exteroception)
- measuring the world outside: lasers, depth and color cameras
- Elevation map
- a grid around the robot with one ground height per cell
- Reinforcement learning
- training a controller by trial and error, here in simulation
- Voxels
- small 3-D cubes marking where something is
- Vision-language model
- an AI trained on internet images and text that can describe a scene